Skip to main content
CertifyOS

Senior Site Reliability Engineer

RemoteUnited States only
Published
Role
SRE
Experience
Senior
Employment
Full-time
Salary not disclosed
Check eligibility

Open to US only. Set where you work from to check your eligibility.

No BS summary

Senior SRE with 5+ years of experience in production systems at scale. Deep hands-on expertise with GCP (GKE, Cloud Run), Infrastructure as Code (Terraform/Pulumi), and observability (Datadog, Prometheus). Must have a track record of preventing incidents through automation and strong Linux systems administration skills.

Core skills

GKECloud RunTerraform

Required skills

LinuxPulumiGitHub ActionsPythonBashGoGoogle Cloud MonitoringDatadogGrafanaPrometheusSentrySnykSonarQubeJiraSlackDockerKubernetesCloud BuildGit

Optional skills

Experience operating large-scale distributed systems or microservices architecturesFamiliarity with healthcare, credentialing, or health-tech environmentsExperience leveraging AI-assisted observability or incident response toolingFamiliarity with NodeJS, TypeScript, Java, or React application stacks

Required languages

English native

What you'll do

  • Design for reliability, ship automation, and stand behind it in production
  • Own the operational lifecycle end-to-end and influence platform architecture, reliability standards, and deployment workflows
  • Maintain uptime, reduce alert fatigue, and build actionable observability across GKE and Cloud Run
  • Improve autoscaling behavior, resource utilization, and workload efficiency across cloud-native distributed systems
  • Own incident response processes, root cause analysis, escalation workflows, and runbooks
  • Build and maintain Infrastructure as Code, CI/CD pipelines, and operational tooling that reduce manual work
  • Instrument data freshness and infrastructure health, not just service uptime

What they require

  • 5+ years in SRE, DevOps, Platform Engineering, or Infrastructure Engineering
  • Track record of improving reliability end-to-end: debugging hard production problems and building alerting
  • Strong Linux systems administration, incident response, and root cause analysis skills
  • Comfort influencing operational standards and mentoring teams on reliability practices
  • Deep hands-on experience with GCP — GKE, Cloud Run, and containerized workloads at scale
  • Experience building and maintaining Infrastructure as Code with Terraform and/or Pulumi
  • Fluency across deployment patterns and the judgment to know when each fits: rolling deployments, blue/green, canary
  • Experience with autoscaling, resource optimization, and infrastructure efficiency for distributed systems
  • Experience managing infrastructure security, secrets, and access controls in regulated or security-conscious environments
  • Strong understanding of Golden Signals monitoring — latency, traffic, errors, saturation
  • Experience designing SLIs, SLOs, error budgets, alerting strategies, dashboards, and escalation workflows
  • Hands-on experience with observability platforms: Google Cloud Monitoring, Datadog, Grafana, Prometheus, or similar
  • Strong sense of data platform health: lineage, freshness, and correctness matter as much to you as throughput
  • Experience building and maintaining CI/CD pipelines using GitHub Actions or similar
  • Scripting or programming fluency in Python, Bash, Go, or similar
  • Experience working with Git workflows and modern software delivery practices
  • Strong written and verbal communication
  • Experience operating systems handling sensitive data or PII in regulated or compliance-adjacent environments

Benefits

  • 100% coverage of health, dental, and vision insurance premiums for employees
  • Unlimited PTO (at least two weeks off each year)
  • Health insurance, statutory leave benefits, and additional wellness (menstrual) leave for women (India employees)

CertifyOS is building the data infrastructure that powers modern healthcare. Its API-first platform automates provider licensing, enrollment, credentialing, and network monitoring by connecting directly to hundreds of primary data sources.

HealthcareStartup
Salary not disclosed