Skip to main content
Runpod

Site Reliability Engineer

RemoteUnited States only
Published
Role
SRE
Experience
Senior
Employment
Full-time
Company size
Startup
$150k–$200k/yr
Check eligibility

Open to US only. Set where you work from to check your eligibility.

No BS summary

The Reliability team owns the availability, performance, and operational excellence of Runpod’s global platform. While infrastructure teams build the systems, the Reliability team ensures those systems remain resilient, observable, and scalable under real-world production conditions.

Required skills

LinuxNetworkingPrometheusGrafanaPythonGoBash

Optional skills

GPU infrastructureAI/ML platformsGPU observability toolingInfrastructure as Code

What you'll do

  • Define and implement SLIs/SLOs for critical services
  • Lead incident response and coordinate cross-team mitigation efforts
  • Design and improve monitoring, alerting, and dashboards (Prometheus, Grafana, etc.)
  • Automate recurring operational workflows
  • Partner with engineering teams to improve system resilience

What they require

  • 5+ years of experience in SRE, Reliability Engineering, or Production Engineering
  • Strong Linux systems and Networking expertise
  • Experience managing containerized production systems
  • Strong understanding of distributed systems and failure modes
  • Experience defining and managing SLIs/SLOs

Benefits

  • Competitive base pay
  • Meaningful equity in a fast-growing company
  • Generous medical, dental & vision plans
  • Flexible PTO
  • Remote work first with an inclusive, collaborative teams

Runpod is the AI Developer Cloud. More than one million developers, from indie researchers to teams running frontier models in production, use Runpod to experiment, train, fine-tune, deploy, and scale AI on one platform.

AI InfrastructureStartup

Details

Visa sponsorshipNo
$150k–$200k/yr