Перейти к основному содержимому
Runpod

Director of Infrastructure Engineering

УдалённоUnited States только
Опубликовано
Роль
Инженерный менеджмент
Опыт
Лид
Занятость
Полная занятость
Размер компании
Стартап
$225k–$325k/yr
Проверьте доступность

Доступно для: US only. Укажите, откуда вы работаете, чтобы проверить доступность.

Ключевые навыки

InfiniBand/RoCEKubernetes

Обязательные навыки

BGPCeph/Lustre/Weka/NVMe-oFTerraformAnsible

Желательные навыки

NVLink

Чем предстоит заниматься

  • Own Core Infrastructure & SRE: Lead multiple engineering teams responsible for Site Reliability Engineering, networking, and storage. Establish rigorous SRE practices, driving SLA/SLO definitions, incident response, observability, and automated remediation.
  • Architect HPC & Global Networking: Oversee the design, scaling, and operation of Runpod’s global network backbone, as well as ultra-low-latency HPC cluster networks. Drive the implementation and optimization of InfiniBand and RDMA over Converged Ethernet (RoCE) to support massive, multi-node GPU training workloads.
  • Drive Storage Engine Innovation: Direct the architecture and performance tuning of highly scalable, distributed storage systems. Ensure our storage engines can deliver the massive IOPS and throughput required to keep high-end GPUs fed with data during deep learning tasks.
  • Build a High-Output Org: Hire, mentor, and grow highly technical engineering managers and senior ICs (network architects, systems engineers, SREs). Create a culture of ownership, operational excellence, and craft in a remote-first environment.
  • Translate Scale into Strategy: Partner with Program Management and Product to forecast capacity requirements, shape technical roadmaps, and convert massive scale challenges into clear technical scopes, milestones, and measurable outcomes.

Что требуется

  • Engineering Leadership Experience: 7+ years leading software, infrastructure, SRE, or networking teams, including managing managers and multiple squads, with a proven record of scaling high-availability cloud environments.
  • Deep Infrastructure Expertise: 8+ years building and operating large-scale distributed systems, bare-metal infrastructure, or public/private cloud platforms.
  • HPC & Advanced Networking: Proven hands-on background or strong architectural understanding of ultra-low latency networking. Deep familiarity with InfiniBand and/or RoCE, spine-leaf architectures, and global WAN routing protocols (BGP).
  • Storage Systems Knowledge: Experience building, operating, or tuning high-performance distributed storage systems and parallel file systems (e.g., Ceph, Lustre, Weka, NVMe-oF) capable of handling heavy AI/ML I/O loads.
  • SRE / DevOps Culture: Strong foundation in reliability engineering, infrastructure-as-code (Terraform, Ansible), container orchestration (Kubernetes), and modern observability stacks.

Преимущества

  • Meaningful equity in a fast-growing company- everyone on the team receives stock options — your impact drives our growth, and you share in the upside.
  • Generous medical, dental & vision plans.
  • Flexible PTO- take the time you need to recharge.
  • $1,200 Home Office & Equipment Stipend- We set you up for success from day one with gear and support to create your ideal workspace.
  • Join a passionate team on the cutting edge of AI infrastructure — where culture, learning, and ownership are at the heart of how we scale.

Runpod is the AI Developer Cloud. More than one million developers, from indie researchers to teams running frontier models in production, use Runpod to experiment, train, fine-tune, deploy, and scale AI on one platform.

AI InfrastructureСтартап

Детали

Спонсорство визыНет
$225k–$325k/yr