Перейти к основному содержимому
NextGen Healthcare

Sr. Cloud Operations Reliability Engineer (SRE)

УдалённоUnited States только
Опубликовано
Роль
SRE
Опыт
Синьор
Занятость
Полная занятость
Размер компании
Крупная
Зарплата не указана
Проверьте доступность

Доступно для: US only. Укажите, откуда вы работаете, чтобы проверить доступность.

Коротко по делу

Senior SRE/Cloud Operations engineer with 10+ years experience. Must have hands-on experience with GCP and/or AWS, Kubernetes, IaC (Terraform/CloudFormation), observability and incident response. Remote hiring for Georgia (GA) / US timezone implied.

Ключевые навыки

KubernetesObservabilityGCP

Обязательные навыки

Google Cloud Platform (GCP)AWSTerraformCloudFormationDeployment ManagerPrometheusGrafanaDatadogNew RelicDistributed tracingInfrastructure as CodePythonBashGoCI/CDIncident responseMonitoringLoggingSLOsSLIsVersion control

Желательные навыки

Google Cloud Associate Cloud EngineerGoogle Cloud Professional Cloud ArchitectGoogle Cloud Professional Cloud Operations EngineerGoogle Cloud Professional Data EngineerAWS SysOps AdministratorKubernetes certificationsTerraform certificationsAPM platforms (Datadog, New Relic)

Чем предстоит заниматься

  • Drive operational excellence and strengthen reliability posture of cloud-based services and supported platforms.
  • Own service reliability and operational health—establish and maintain SLOs/SLIs, design monitoring and alerting strategies.
  • Lead incident response coordination and post-incident processes, including troubleshooting complex production issues and conducting root cause analysis.
  • Design and implement reliability-focused automation, operational tooling, and runbooks; apply Infrastructure as Code practices to support recovery and consistency.
  • Build observability solutions through comprehensive monitoring, logging, and alerting strategies; establish event correlation and escalation procedures.
  • Conduct performance and capacity analysis, evaluate utilization trends, and provide recommendations for scaling and capacity planning.
  • Partner with development and engineering teams to evaluate deployment readiness and support deployment reliability improvements.
  • Contribute to disaster recovery and business continuity planning; conduct operational readiness exercises and maintain recovery documentation.
  • Mentor team members and establish reliability standards and practices; create and maintain operational documentation and knowledge base.
  • Support operational adherence to cloud governance, compliance, access control, tagging, logging, and audit readiness.

Что требуется

  • Bachelor's degree in Computer Science, Engineering, Information Systems, or related field, or equivalent combination of education and experience.
  • 10+ years of professional experience in Cloud Operations, Site Reliability Engineering, DevOps, Infrastructure Operations, or related discipline with demonstrated ownership of production systems.
  • Extensive hands-on experience supporting production cloud environments using GCP, AWS, or equivalent cloud service providers.
  • Proven expertise in monitoring, observability platforms, alerting strategies, incident response, root cause analysis, and production support in distributed or cloud-native architectures.
  • Demonstrated experience with Infrastructure as Code (Terraform, Deployment Manager, CloudFormation, etc.) and version control best practices.
  • Strong background in incident management and post-incident review processes; experience driving corrective actions and establishing reliability improvements.
  • Experience with Kubernetes operations, containerization, and orchestration platforms.
  • Experience with application performance monitoring (APM) and distributed tracing.
  • Experience mentoring junior engineers or leading operational improvement initiatives.
  • Required/desired certifications: Google Cloud certifications (Associate Cloud Engineer, Professional Cloud Architect, Professional Cloud Operations Engineer, or Professional Data Engineer) and AWS SysOps Administrator or equivalent. Advanced certifications in Kubernetes, Terraform, observability platforms, DevOps, SRE, or ITIL.

NextGen Healthcare is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.

HealthcareКрупнаяnextgen.com/

Что говорят о компании

3.1/ 5

  • Flexible work hours and remote work options are appreciated by many employees.
  • Some employees mention a supportive team environment and good colleagues.
  • There are concerns about management and leadership effectiveness.
Зарплата не указана