Director of Infrastructure Engineering
runpod · Remote
- Posted
- 45 days ago
- Last confirmed live
- Today
What this role involves
Runpod is seeking a Director of Infrastructure Engineering to lead and scale its core cloud and bare-metal environments. The role oversees SRE, global networking, HPC networks, and distributed storage, with a focus on high availability and performance for AI workloads. The position requires 7+ years of leadership experience and 8+ years in large-scale distributed systems, and is remote-first.
Skills this posting asks for
- site reliability engineering
- networking
- high-performance computing
- distributed storage
- infiniband
- rdma over converged ethernet
- observability
- automated remediation
- infrastructure as code
- bare-metal provisioning
- virtualization
- network fabrics
- storage clusters
- sla/slo
- incident response
- mtbf
- mttr
- gpu computing
- deep learning
- capacity planning
- technical roadmap
- program management
- product engineering
- gtm leadership
Requirements
- 8 years of experience
- Level: director
- Remote policy: remote
From the employer’s posting
Runpod is the AI Developer Cloud. More than one million developers, from indie researchers to teams running frontier models in production, use Runpod to experiment, train, fine-tune, deploy, and scale AI on one platform. The platform has processed more than 20 billion inference requests. We closed…
Read the full description on runpod’s careers pageApply without filling the form
Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.
Other roles at runpod
- Senior Data Engineer Remote
- Senior Developer Advocate Remote
- Developer Relations Engineer Remote
- Software Engineer (Full-Stack) Remote
- Senior Product Manager Remote
- Site Reliability EngineerRemote