Skip to main content

Senior Site Reliability Engineer

RemoteCanada, Mexico, United States only· UTC-8…UTC-8
Published
Role
SRE
Experience
Senior
$180k–$205k/yr
Check eligibility

Open to CA, MX, US only · UTC-8…UTC-8. Set where you work from to check your eligibility.

No BS summary

Senior SRE with 5+ years operating production services at scale. Must have strong production Kubernetes, AWS cloud infrastructure, Terraform, Prometheus/Grafana, Python/Bash automation, incident response, 24/7 on-call, and strong English. Remote from anywhere in PST timezone.

Core skills

KubernetesAWSTerraform

Required skills

EKSRDSS3EC2PrometheusGrafanaPythonBash

Optional skills

DevelocityJavaKotlin

Required languages

English Strong written and verbal communication

What you'll do

  • Operate and maintain all Develocity instances and supporting services.
  • Participate in a follow-the-sun on-call rotation, owning incident response and troubleshooting issues across the stack.
  • Drive automation across application deployment, upgrades, monitoring, self-healing, and recovery.
  • Build and maintain observability for all managed services (logging, metrics, tracing, and alerting).
  • Work with engineering teams to build reliability into features from the start.
  • Run incident response and retrospectives, and make sure we learn from them.
  • Own disaster recovery, backups, and business continuity.
  • Communicate with customers during incidents and maintenance windows.
  • Optimize performance, resource usage, and costs.
  • Help evolve our SaaS operations as we grow.

What they require

  • 5+ years in SRE, DevOps, or equivalent role operating production services at scale.
  • Strong Kubernetes experience in production environments.
  • Cloud infrastructure expertise, preferably AWS (EKS, RDS, S3, EC2).
  • Proficiency with observability tools (Prometheus, Grafana) and Infrastructure as Code (Terraform).
  • Track record of incident management and response.
  • Knowledge of SRE best practices (SLAs, SLOs).
  • Scripting proficiency (Python, Bash) for automation.
  • Experience with 24/7 on-call rotations.
  • Strong written and verbal English communication.
  • Preferred: Experience operating SaaS platforms at scale.
  • Preferred: Disaster recovery planning and execution experience.
  • Preferred: Customer-facing incident communication skills.
  • Preferred: Experience establishing SRE practices in new or growing teams.
  • Strong self-direction and clear communication across time zones are essential.

Benefits

  • A ground-floor role in a new SRE team—you'll shape how we do things, not inherit someone else's decisions.
  • Real ownership of production systems used by engineers at companies you've heard of.
  • Direct interaction with customers when things go wrong (and when they go right).
  • A culture that values automation over heroics.
  • In-person meetings, such as our annual company offsite and team meetings.
  • Work from home in a remote-first environment.
  • Competitive salaries and equity grants.

AI-native company building Develocity, a toolchain observability and intelligence platform used by software organizations for delivery excellence, build and test acceleration, and AI-powered intelligence across the toolchain.

Developer Tools

Details

Apply routeGreenhouse
$180k–$205k/yr