Skip to main content
Runpod

Director of Infrastructure Engineering

RemoteUnited States only
Published
Role
Engineering Management
Experience
Lead
Employment
Full-time
Company size
Startup
$225k–$325k/yr
Check eligibility

Open to US only. Set where you work from to check your eligibility.

Core skills

InfiniBand/RoCEKubernetes

Required skills

BGPCeph/Lustre/Weka/NVMe-oFTerraformAnsible

Optional skills

NVLink

What you'll do

  • Own Core Infrastructure & SRE: Lead multiple engineering teams responsible for Site Reliability Engineering, networking, and storage. Establish rigorous SRE practices, driving SLA/SLO definitions, incident response, observability, and automated remediation.
  • Architect HPC & Global Networking: Oversee the design, scaling, and operation of Runpod’s global network backbone, as well as ultra-low-latency HPC cluster networks. Drive the implementation and optimization of InfiniBand and RDMA over Converged Ethernet (RoCE) to support massive, multi-node GPU training workloads.
  • Drive Storage Engine Innovation: Direct the architecture and performance tuning of highly scalable, distributed storage systems. Ensure our storage engines can deliver the massive IOPS and throughput required to keep high-end GPUs fed with data during deep learning tasks.
  • Build a High-Output Org: Hire, mentor, and grow highly technical engineering managers and senior ICs (network architects, systems engineers, SREs). Create a culture of ownership, operational excellence, and craft in a remote-first environment.
  • Translate Scale into Strategy: Partner with Program Management and Product to forecast capacity requirements, shape technical roadmaps, and convert massive scale challenges into clear technical scopes, milestones, and measurable outcomes.

What they require

  • Engineering Leadership Experience: 7+ years leading software, infrastructure, SRE, or networking teams, including managing managers and multiple squads, with a proven record of scaling high-availability cloud environments.
  • Deep Infrastructure Expertise: 8+ years building and operating large-scale distributed systems, bare-metal infrastructure, or public/private cloud platforms.
  • HPC & Advanced Networking: Proven hands-on background or strong architectural understanding of ultra-low latency networking. Deep familiarity with InfiniBand and/or RoCE, spine-leaf architectures, and global WAN routing protocols (BGP).
  • Storage Systems Knowledge: Experience building, operating, or tuning high-performance distributed storage systems and parallel file systems (e.g., Ceph, Lustre, Weka, NVMe-oF) capable of handling heavy AI/ML I/O loads.
  • SRE / DevOps Culture: Strong foundation in reliability engineering, infrastructure-as-code (Terraform, Ansible), container orchestration (Kubernetes), and modern observability stacks.

Benefits

  • Meaningful equity in a fast-growing company- everyone on the team receives stock options — your impact drives our growth, and you share in the upside.
  • Generous medical, dental & vision plans.
  • Flexible PTO- take the time you need to recharge.
  • $1,200 Home Office & Equipment Stipend- We set you up for success from day one with gear and support to create your ideal workspace.
  • Join a passionate team on the cutting edge of AI infrastructure — where culture, learning, and ownership are at the heart of how we scale.

Runpod is the AI Developer Cloud. More than one million developers, from indie researchers to teams running frontier models in production, use Runpod to experiment, train, fine-tune, deploy, and scale AI on one platform.

AI InfrastructureStartup

Details

Visa sponsorshipNo
$225k–$325k/yr