Skip to main content
OpenAI
OpenAI

Software Engineer, Compute Infrastructure

RemoteUnited States, United Kingdom only
Published
Role
DevOps
Employment
Full-time
Company size
Enterprise
$230k–$405k/yr
Check eligibility

Open to US, GB only. Set where you work from to check your eligibility.

No BS summary

Software engineer for large-scale AI compute infrastructure. Needs strong production infrastructure experience and one or more areas like distributed systems, Kubernetes, networking/RDMA/NCCL, storage, HPC, GPU infrastructure, observability, reliability, or tooling. Hiring is limited to the United States and United Kingdom.

Core skills

Distributed systems/Operating systems/Networking protocols/RDMA/NCCL/Storage/Kubernetes/Scheduling/Observability/High-performance computing/GPU infrastructure/CaaS/Benchmarking/Infrastructure tooling

What you'll do

  • Build and optimize reliable system software for large-scale compute systems running demanding AI workloads
  • Design and operate infrastructure across accelerators, CPUs, NICs, switches, networking protocols, storage, data centers, orchestration, scheduling, and fleet health
  • Profile, benchmark, and optimize training workloads across compute, memory, storage, networking, NCCL, collective communication, and cluster scheduling bottlenecks
  • Create hardware-aware automation for provisioning, firmware and driver upgrades, incident response, and operations
  • Build CaaS, agent infrastructure, profiling, observability, benchmarking, and platform tools for researchers, product engineers, and operators
  • Turn operational lessons into better systems, stronger abstractions, and clearer ownership boundaries
  • Collaborate across research, engineering, security, networking, hardware, and data center teams

What they require

  • Strong software engineering skills and experience building, operating, or improving production infrastructure systems
  • Experience in one or more relevant areas such as distributed systems, operating systems, networking protocols, RDMA, NCCL or collective communication, storage, Kubernetes, scheduling, observability, reliability engineering, high-performance computing, GPU infrastructure, CaaS, agent infrastructure, hardware-aware performance optimization, benchmarking, developer experience, or infrastructure tooling
  • Ability to debug complex system behavior across software, hardware, networking, and workload layers and turn findings into robust improvements
  • Comfort with ambiguity, strong ownership, and a bias toward practical, durable solutions
  • Interest in working on infrastructure that directly enables frontier AI research and product impact

Benefits

  • Offers equity
  • Reasonable accommodations for applicants with disabilities

American artificial intelligence research organization

🇺🇸 United StatesAIEnterpriseopenai.com

What people say about this company

4.6/ 5

Also posted in 1 other channel
$230k–$405k/yr