Skip to main content
Lambda

Senior Incident Manager

RemoteUnited States only
Published
Role
SRE
Experience
Lead
Employment
Full-time
Company size
Mid-size
$125k–$195k/yr
Check eligibility

Open to US only. Set where you work from to check your eligibility.

No BS summary

Senior incident manager with 8+ years in incident management, SRE, or infrastructure operations. Must have large-scale distributed infrastructure experience across AI/data center operations, GPU clusters, networking, storage, and cloud/hybrid platforms. US remote role with on-call incident leadership.

Core skills

PagerDutyServiceNowDatadog

Required skills

JiraPrometheusGrafana

Optional skills

NVIDIA clustersInfiniBandautomation toolingincident response toolingIncident Command System

What you'll do

  • Lead the response to critical (SEV-1 / SEV-2) incidents impacting AI infrastructure, GPU clusters, networking, storage, and data center operations.
  • Serve as the Incident Commander during major outages, coordinating engineering, networking, facilities, and vendor teams.
  • Act as the liaison between leadership and external teams during incidents / post-incidents to provide updates and status summaries.
  • Establish clear incident timelines, triage actions, and resolution plans.
  • Own the incident response lifecycle including assisting technical triage, escalation, coordination, resolution, and post-incident review.
  • Ensure timely and accurate communication with internal stakeholders and leadership.
  • Maintain incident response documentation and operational playbooks.
  • Conduct analysis on incidents and identify patterns / trends for improvement in response and systems reliability.
  • Work in an On-Call Rotation to respond to, lead, and coordinate incidents.
  • Work closely with data center operations, infrastructure engineering & operations, network engineering, platform reliability engineering, security operations, hardware and facility vendors.
  • Drive alignment during outages involving multiple infrastructure layers.
  • Lead post-incident reviews (PIRs) and root cause analysis.
  • Identify systemic reliability gaps and implement corrective actions.
  • Track incident metrics including MTTR, MTTD, and incident recurrence rates.
  • Improve incident response processes, escalation paths, and tooling by working with technical support and engineering teams.
  • Contribute to runbooks, operational standards, and reliability frameworks.
  • Support implementation of automation and observability improvements.
  • Provide executive-level incident summaries and reports.
  • Deliver clear, concise updates during active incidents.
  • Maintain incident dashboards and operational health reporting.

What they require

  • 8+ years experience in incident management, site reliability engineering, or infrastructure operations.
  • Experience managing incidents in large-scale distributed infrastructure environments.
  • Strong understanding of data center operations.
  • Strong understanding of GPU compute clusters.
  • Strong understanding of networking and storage infrastructure.
  • Strong understanding of cloud or hybrid infrastructure platforms.
  • Proven ability to lead high-pressure incident response situations.
  • Experience with incident management frameworks (ITIL, SRE, or equivalent).
  • Excellent communication and stakeholder management skills.
  • Experience with incident tracking and monitoring tools such as PagerDuty, ServiceNow, Jira, Datadog, Prometheus / Grafana.
  • Preferred: Experience operating AI or HPC infrastructure.
  • Preferred: Background in SRE, infrastructure engineering, or data center operations.
  • Preferred: Familiarity with high-density GPU environments (NVIDIA clusters, InfiniBand networks).
  • Preferred: Experience with hyperscale or colocation data center environments.
  • Preferred: Knowledge of automation and incident response tooling.
  • Preferred: Knowledge of and experience with Incident command system (ICS).
  • Preferred: Experience in leading and developing incident command from stractch.
  • Incident Command & Leadership.
  • Operational Decision Making.
  • Cross-Team Coordination.
  • Root Cause Analysis.
  • Crisis Communication.
  • Infrastructure Reliability.

Benefits

  • Generous cash & equity compensation.
  • Health, dental, and vision coverage for you and your dependents.
  • Wellness and commuter stipends for select roles.
  • 401k Plan with 2% company match (USA employees).
  • Flexible paid time off plan that we all actually use.

award for published works which celebrate or explore LGBT themes

AI CloudMid-size
$125k–$195k/yr