Skip to main content
Supabase

Site Reliability Engineer

RemoteWorldwide
Published
Role
SRE
Experience
Senior
Employment
Full-time
Company size
Mid-size
Salary not disclosed
Check eligibility

Open to Worldwide. Set where you work from to check your eligibility.

No BS summary

SRE with 7+ years in reliability or production engineering, able to shape SRE practices across engineering teams. Must know SLOs/SLIs, incident response, large-scale multi-tenant systems, cloud infrastructure, and infrastructure-as-code. Global remote role in an async distributed team; AWS and Pulumi are preferred.

Required skills

SLOsSLIsAWSPulumi/Terraform/CDK

Optional skills

KubernetesOpenTelemetryVictoriaMetricsGrafanaPostgreSQL

What you'll do

  • Partner with service teams to define meaningful SLIs and SLOs grounded in customer experience
  • Build error budget policies that turn SLOs into engineering decisions
  • Own and evolve the Operational Readiness Review process for new services and major changes
  • Review observability, alerting, runbooks, capacity, and graceful degradation readiness
  • Connect postmortem findings to operational readiness gaps and drive systemic fixes
  • Support architecture reviews, failure mode analysis, dependency mapping, and resilience design
  • Identify and quantify operational toil and build or advocate for automation to eliminate it
  • Help teams design sustainable on-call practices, escalation paths, runbook coverage, and noise reduction
  • Track and report on org-wide operational maturity and drive remediation

What they require

  • 7+ years of experience in SRE, production engineering, or reliability-focused roles
  • Experience shaping SRE practices and driving adoption across engineering teams
  • Software engineering mindset; writes code and builds tools, not just configures them
  • Hands-on experience defining and operationalizing SLOs/SLIs at scale
  • Experience with error budget policies influencing engineering decisions
  • Deep experience with incident response and postmortem facilitation
  • Experience turning incident learnings into systemic improvements
  • Experience with large-scale multi-tenant systems
  • Proficiency with cloud infrastructure; AWS preferred
  • Proficiency with infrastructure-as-code; Pulumi preferred, Terraform/CDK acceptable
  • Clear and persuasive communication with ability to influence without authority across a distributed organization
  • Experience in async or globally distributed teams
  • Energized by making other teams more effective rather than personally fixing everything

Benefits

  • Fully remote work from anywhere
  • WeWork membership or co-working allowance usable anywhere in the world
  • ESOP equity ownership
  • Tech allowance for laptop, monitor, headphones, or work setup
  • 100% employee health insurance coverage
  • 80% dependent health insurance coverage
  • Annual company off-sites
  • Flexible asynchronous work
  • Annual education allowance for courses, books, conferences, or other learning

open source backend platform for app development

TechnologyMid-sizesupabase.com

Details

Apply routeDom
Salary not disclosed