Skip to main content
NextGen Healthcare

Sr. Cloud Operations Reliability Engineer (SRE)

RemoteUnited States only
Published
Role
SRE
Experience
Senior
Employment
Full-time
Company size
Enterprise
Salary not disclosed
Check eligibility

Open to US only. Set where you work from to check your eligibility.

No BS summary

Senior SRE/Cloud Operations engineer with 10+ years experience. Must have hands-on experience with GCP and/or AWS, Kubernetes, IaC (Terraform/CloudFormation), observability and incident response. Remote hiring for Georgia (GA) / US timezone implied.

Core skills

KubernetesObservabilityGCP

Required skills

Google Cloud Platform (GCP)AWSTerraformCloudFormationDeployment ManagerPrometheusGrafanaDatadogNew RelicDistributed tracingInfrastructure as CodePythonBashGoCI/CDIncident responseMonitoringLoggingSLOsSLIsVersion control

Optional skills

Google Cloud Associate Cloud EngineerGoogle Cloud Professional Cloud ArchitectGoogle Cloud Professional Cloud Operations EngineerGoogle Cloud Professional Data EngineerAWS SysOps AdministratorKubernetes certificationsTerraform certificationsAPM platforms (Datadog, New Relic)

What you'll do

  • Drive operational excellence and strengthen reliability posture of cloud-based services and supported platforms.
  • Own service reliability and operational health—establish and maintain SLOs/SLIs, design monitoring and alerting strategies.
  • Lead incident response coordination and post-incident processes, including troubleshooting complex production issues and conducting root cause analysis.
  • Design and implement reliability-focused automation, operational tooling, and runbooks; apply Infrastructure as Code practices to support recovery and consistency.
  • Build observability solutions through comprehensive monitoring, logging, and alerting strategies; establish event correlation and escalation procedures.
  • Conduct performance and capacity analysis, evaluate utilization trends, and provide recommendations for scaling and capacity planning.
  • Partner with development and engineering teams to evaluate deployment readiness and support deployment reliability improvements.
  • Contribute to disaster recovery and business continuity planning; conduct operational readiness exercises and maintain recovery documentation.
  • Mentor team members and establish reliability standards and practices; create and maintain operational documentation and knowledge base.
  • Support operational adherence to cloud governance, compliance, access control, tagging, logging, and audit readiness.

What they require

  • Bachelor's degree in Computer Science, Engineering, Information Systems, or related field, or equivalent combination of education and experience.
  • 10+ years of professional experience in Cloud Operations, Site Reliability Engineering, DevOps, Infrastructure Operations, or related discipline with demonstrated ownership of production systems.
  • Extensive hands-on experience supporting production cloud environments using GCP, AWS, or equivalent cloud service providers.
  • Proven expertise in monitoring, observability platforms, alerting strategies, incident response, root cause analysis, and production support in distributed or cloud-native architectures.
  • Demonstrated experience with Infrastructure as Code (Terraform, Deployment Manager, CloudFormation, etc.) and version control best practices.
  • Strong background in incident management and post-incident review processes; experience driving corrective actions and establishing reliability improvements.
  • Experience with Kubernetes operations, containerization, and orchestration platforms.
  • Experience with application performance monitoring (APM) and distributed tracing.
  • Experience mentoring junior engineers or leading operational improvement initiatives.
  • Required/desired certifications: Google Cloud certifications (Associate Cloud Engineer, Professional Cloud Architect, Professional Cloud Operations Engineer, or Professional Data Engineer) and AWS SysOps Administrator or equivalent. Advanced certifications in Kubernetes, Terraform, observability platforms, DevOps, SRE, or ITIL.

NextGen Healthcare is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.

HealthcareEnterprisenextgen.com/
Salary not disclosed