Skip to main content
Buildkite

Incident Management & Resilience Lead

RemoteANZUS-Pacific· UTC-8…UTC+13
Published
Role
SRE
Experience
Lead
Company size
Mid-size
Salary not disclosed
Check eligibility

Open to Anywhere in ANZ & US-Pacific · UTC-8…UTC+13. Set where you work from to check your eligibility.

No BS summary

Incident management lead for production systems at real scale, owning incident response, postmortems, resilience standards, and cross-team process. Must be able to investigate code and telemetry hands-on, work with observability tooling, AWS, Terraform, and distributed systems. Hiring only in ANZ and US-Pacific time zones; no sponsorship.

Core skills

AWSIncident ManagementObservability

Required skills

DatadogHoneycombTerraform

Optional skills

Ruby on RailsPostgreSQLKafkaRedis

What you'll do

  • Own incident management across engineering — the definition, process, tooling, and adherence — and make it stick.
  • Be in the room on high-impact incidents, driving clarity, coordination and a clean resolution.
  • Run postmortems that go somewhere.
  • Make sure followups get done every time.
  • Turn what each postmortem taught into changes that make the next incident smaller.
  • Work across every product team plus Security and Support, holding one consistent resilience bar and bringing engineering managers with you.
  • Coach engineers on how to run an incident, communicate through it, and own the outcome.
  • Set the standard for how the whole engineering org detects, responds to, learns from, and designs out the incidents that matter — end to end.

What they require

  • You've built or materially lifted incident management capability somewhere before, and you've got a clear, earned point of view on what good looks like.
  • You can read the code, dig into the telemetry, and lead a real investigation — not just chair the call.
  • You're calm when it's sustained and messy: you make incidents smaller, not louder.
  • You can hold a firm bar across teams that don't report to you without softening the standard or bruising the relationship — the diplomacy to find common ground and the spine to keep the line.
  • You'll be at home with observability tooling (Datadog, Honeycomb), AWS, Terraform, and the failure modes distributed systems throw up at scale.
  • You've owned incident response for production systems at real scale for a number of years, meaning you've experienced a breadth of different challenges and scenarios.
  • Would rather draw the map on a blank page than inherit someone else's.
  • Get more out of the incident that never happened than the heroic save.
  • Can hold a standard across teams you don't manage, and keep the relationship.
  • Read "flat, high-autonomy, little scaffolding" as room to move, not risk.
  • Buildkite is fully remote, but doesn't hire everywhere: if you're applying from outside ANZ and US-Pacific regions, they're not currently in a position to hire you.
  • Buildkite can't offer sponsorship.

Benefits

  • First of its kind here. No one has owned this before you.
  • You define what incident management at Buildkite means — the mandate is real and yours to shape.
  • Scale that's genuinely rare. Billions of builds, in the critical path for some of the strongest engineering teams on the planet — systems that run continuously and can't be casually restarted.
  • Ownership, not tickets. Flat structure, high autonomy. You're trusted to make the calls that matter and own the result.
  • Remote, properly. We've worked this way since 2013 — async and built for deep focus, with genuine overlap across our ANZ and US-Pacific teams.
  • Small enough that it counts. ~150 people. What you change is visible across the org.
  • Every application gets a response.

Buildkite runs production builds for teams like OpenAI, Anthropic, Uber, Shopify, Airbnb, Canva and Pinterest — software delivery in the critical path for over a billion daily users.

🇺🇸 United StatesDevOpsMid-sizebuildkite.com/

Details

Visa sponsorshipNo
Salary not disclosed