Director of Infrastructure Engineering

runpod · Remote

Posted
45 days ago
Last confirmed live
Today

What this role involves

Runpod is seeking a Director of Infrastructure Engineering to lead and scale its core cloud and bare-metal environments. The role oversees SRE, global networking, HPC networks, and distributed storage, with a focus on high availability and performance for AI workloads. The position requires 7+ years of leadership experience and 8+ years in large-scale distributed systems, and is remote-first.

Skills this posting asks for

  • site reliability engineering
  • networking
  • high-performance computing
  • distributed storage
  • infiniband
  • rdma over converged ethernet
  • observability
  • automated remediation
  • infrastructure as code
  • bare-metal provisioning
  • virtualization
  • network fabrics
  • storage clusters
  • sla/slo
  • incident response
  • mtbf
  • mttr
  • gpu computing
  • deep learning
  • capacity planning
  • technical roadmap
  • program management
  • product engineering
  • gtm leadership

Requirements

  • 8 years of experience
  • Level: director
  • Remote policy: remote

From the employer’s posting

Runpod is the AI Developer Cloud. More than one million developers, from indie researchers to teams running frontier models in production, use Runpod to experiment, train, fine-tune, deploy, and scale AI on one platform. The platform has processed more than 20 billion inference requests. We closed…

Read the full description on runpod’s careers page

Apply without filling the form

Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.

Other roles at runpod

All 8 roles at runpod