Skip to main content
Claritas Rx

Senior Manager, Site Reliability Engineering

RemoteUnited States only
Published
Role
SRE
Experience
Lead
Employment
Full-time
Company size
Startup
$155k–$170k/yr
Check eligibility

Open to US only. Set where you work from to check your eligibility.

No BS summary

Experienced and hands-on Sr. Manager of Site Reliability Engineering needed to lead SRE/DevOps for an AWS-hosted SaaS platform. Must have 7+ years in SRE/DevOps with 3+ years in management, deep AWS expertise, strong IaC skills, and experience with SLOs, incident management, and compliance (HIPAA, SOC 2, HITRUST). Player-coach role requiring direct infrastructure contribution and team leadership.

Core skills

AWSInfrastructure as CodeSRE

Required skills

AWS CDKTerraformGitHub ActionsECSEC2Aurora RDSDynamoDBLambdaS3SQSEventBridgeCognitoSecrets ManagerCloudFrontCloudWatchSentryHIPAASOC 2HITRUST

Optional skills

OpenTofuNestJSTypeScriptReactPostgreSQLTurborepopnpmAWS Glue

Required languages

English

What you'll do

  • Own the reliability, availability, performance, and scalability of Claritas Rx's AWS-hosted SaaS platform, with accountability for SLA/SLO commitments made to customers.
  • Define and maintain Service Level Objectives (SLOs) and Service Level Indicators (SLIs) across all production services; use error budgets to drive engineering prioritization conversations.
  • Lead incident response and on-call operations: triage, coordinate resolution, communicate to stakeholders, and conduct thorough post-incident reviews with actionable corrective actions.
  • Drive a proactive reliability culture — identifying risks before they become incidents through load testing, chaos engineering, and systematic failure mode analysis.
  • Architect, build, and maintain AWS cloud-native infrastructure using infrastructure-as-code (AWS CDK, Terraform, or equivalent), ensuring environments are reproducible, auditable, and secure.
  • Oversee and continuously improve CI/CD pipelines (GitHub Actions) to enable rapid, safe, and consistent delivery of application and infrastructure changes across environments.
  • Manage and optimize core AWS services including ECS, EC2, Aurora RDS, DynamoDB, Lambda, S3, SQS, EventBridge, Cognito, Secrets Manager, and CloudFront.
  • Ensure robust observability across the stack — centralizing logs, metrics, traces, and alerts using CloudWatch, Sentry, and related tooling — so the team can detect and respond to issues quickly.
  • Manage platform capacity planning, cost optimization, and cloud spend governance.
  • Ensure all infrastructure design and operational practices meet HIPAA, SOC 2, and HITRUST requirements, given the PHI our platform processes.
  • Partner with the Security function on vulnerability management, infrastructure hardening, secrets management, and access control.
  • Maintain and regularly test disaster recovery (DR) and business continuity plans, including defined RTO/RPO targets for all production systems.
  • Support audit readiness and evidence collection for compliance certifications.
  • Lead, mentor, and grow a blended team of full-time SREs/DevOps engineers and offshore contractors, fostering a culture of ownership, continuous improvement, and operational excellence.
  • Manage distributed team dynamics effectively — establishing clear communication rhythms, documentation standards, and handoff protocols to ensure offshore resources are productive and well-integrated.
  • Conduct regular 1:1s, set clear goals and development plans for direct reports, and advocate for your team's growth and recognition.
  • Build and maintain a healthy on-call rotation with appropriate tooling, runbooks, and escalation paths to protect team sustainability.
  • Collaborate with Software Engineering teams to embed reliability practices into the SDLC — including production readiness reviews, deployment standards, and shared observability tooling.
  • Work with Product Management and Engineering leadership to balance feature delivery velocity against operational risk and technical debt.
  • Contribute to architecture decisions across the platform, providing an operational and reliability perspective on new services and major technical changes.
  • Champion automation-first thinking: eliminate toil through tooling, scripting, and process improvement wherever possible.

What they require

  • 7+ years of experience in SRE, DevOps, or infrastructure engineering, with at least 3 years in a people management or team lead capacity.
  • Deep, hands-on AWS expertise — you understand how to architect, operate, and optimize cloud-native workloads at the service level, not just at a conceptual level.
  • Strong infrastructure-as-code skills (AWS CDK, Terraform, or equivalent); you treat infrastructure like software.
  • Demonstrated experience owning SLOs, incident management, and on-call operations in a commercial SaaS environment.
  • Experience managing CI/CD pipelines and developer productivity tooling, with a strong understanding of deployment safety (canary releases, feature flags, rollback strategies).
  • Solid working knowledge of security and compliance requirements relevant to regulated data environments (HIPAA, SOC 2, HITRUST).
  • Proven ability to lead and develop a team, including working effectively with offshore or distributed contractors across time zones.
  • Strong written and verbal communication skills — you can explain infrastructure risk and trade-offs clearly to both technical and non-technical audiences.
  • Comfort operating in a fast-paced, high-growth startup environment where priorities evolve and initiative is expected.
  • Role is primarily remote with occasional travel requirements.
  • Experience in a healthcare technology or digital health environment with direct exposure to HIPAA-regulated PHI.
  • Familiarity with the application stack in use at Claritas Rx (TypeScript, NestJS, PostgreSQL, React).
  • Experience with chaos engineering practices and tools.
  • Experience leveraging AI tools (including Claude) to improve operational workflows, accelerate runbook development, or automate incident triage.
  • B.S. in Computer Science, Engineering, or a related discipline, or equivalent practical experience.

Benefits

  • unlimited PTO
  • stock options
  • competitive salary of $155,000 to $170,000
  • benefits package
  • opportunity to make a significant impact on a first-in-industry digital health solution
  • shared learning and collaboration
  • a respectful and fun work environment
  • employee empowerment through the effective use of technology and tools
  • opportunities to connect in person
  • regional town hall gatherings approximately every other month for employees within a reasonable driving distance of each other

Claritas Rx uses AI and predictive modeling to help rare disease and specialty brands remove barriers that keep patients from accessing and staying on treatments, combining advanced analytics, real-world data, AI, and CRM capabilities for digital health.

HealthcareMid-size
$155k–$170k/yr