Skip to main content
Heidi

Engineering Lead - Reliability & Cloud

RemoteAustralia only
Published
Role
SRE
Experience
Lead
Employment
Full-time
Company size
Startup
Salary not disclosed
Check eligibility

Open to AU only. Set where you work from to check your eligibility.

No BS summary

Lead the team that ensures Heidi's cloud platform and reliability practices are robust, supporting millions of patient visits weekly. Own the multi-region cloud footprint, set reliability standards, manage incident response, and partner with engineering to improve deployment processes. Requires significant cloud/SRE leadership experience, hands-on cloud provider, Kubernetes, and IaC skills, and a proven track record in high-availability design and incident command.

Core skills

KubernetesSRECloud Infrastructure

Required skills

TerraformIaC

What you'll do

  • Lead and grow the platform team covering cloud infrastructure and reliability.
  • Set direction, run hiring, coach engineers through hard technical calls and make sure the on-call load is sustainable rather than heroic.
  • Own Heidi's multi-region cloud footprint end to end: compute, networking, data stores, Kubernetes, identity and the infrastructure-as-code that defines all of it.
  • Clinical data lives in-region, and your architecture is what keeps that promise across the US, UK, Canada, Australia and Europe.
  • Set the reliability bar and hold it.
  • Define SLIs and SLOs with the product teams that own each service, run error budgets that actually change roadmap decisions, and report honestly when we spend one.
  • Run incident response and make it better every time.
  • Own severity definitions, escalation paths, comms to health system customers during an outage, and blameless post-incident reviews where the follow-ups get finished.
  • Partner with the Release Manager and engineering leads to sharpen how Heidi ships: feature flags, canary and staged rollouts, and rollback paths that get exercised before they're needed.
  • Move deployment frequency and change lead time without moving change failure rate.

What they require

  • Significant experience running cloud infrastructure or SRE teams at scale, including time leading engineers.
  • A regulated environment (healthcare, fintech, clinical technology) is a strong plus.
  • Deep, current, hands-on skill with a major cloud provider, Kubernetes, Terraform or equivalent IaC, and modern observability stacks.
  • You still read the dashboards yourself.
  • A real reliability practice behind you: SLOs you defined, error budgets you enforced, incidents you commanded and post-incident reviews that changed how a system was built.
  • Track record designing for high availability across regions, with the failure modes, data residency constraints and DR testing that come with it.
  • A service mindset toward product engineers.
  • You measure the platform by whether it makes other teams faster, not by how elegant it is.
  • You'll be in the on-call rotation you design.
  • Leaders here carry the pager.

Benefits

  • A $1,000 annual learning and development budget
  • a $150/month health and wellness allowance
  • a $500 home office budget
  • 26 weeks paid primary parental leave
  • 18 weeks paid secondary parental leave
  • fertility support up to $10,000
  • four weeks of work from anywhere per year
  • serious equity

We're building the future of healthcare by giving every clinician the earth's finest AI Care Partner. In just 18 months, our clinical AI products have absorbed the administrative chaos of 73 million patient visits. Today, we support over 2.5 million patient sessions a week across 190+ countries.

🇺🇸 United StatesHealthcareStartup
Salary not disclosed