Skip to main content

Senior Site Reliability Engineer

RemoteUnited States only
Published
Role
SRE
Experience
Senior
Employment
Full-time
$104.1k–$140k/yr
Check eligibility

Open to US only. Set where you work from to check your eligibility.

No BS summary

Senior SRE needed for cloud engineering team to manage AWS production environment, databases, backups, and monitoring for a serverless platform. Requires 5+ years in SRE/DevOps, 3+ years AWS (serverless focus), strong PostgreSQL admin, and scripting proficiency (TypeScript, Python, bash).

Core skills

AWSPostgreSQLTypeScript

Required skills

LambdaSQSEventBridgeCloudWatchS3PythonbashLinuxDNSTLSDockerGitHub ActionsSSTPulumiTerraform

What you'll do

  • Own day-to-day administration across AWS services, accounts, and access, as well as database administration across PostgreSQL and our other data stores.
  • Own backup posture across databases, S3 buckets, and queues; verify restores regularly and maintain a tested disaster recovery plan.
  • Proactively monitor production — CloudWatch dashboards, metric alarms, log-based metrics, and Slack alerting — addressing operational issues before they impact users.
  • Lead production debugging and incident response: build and maintain runbooks, participate in the on-call rotation, and resolve queue and dead-letter-queue failures through retry, redrive, and recovery.
  • Continuously refine our infrastructure to ensure it is easily deployable and scalable: keep infrastructure as code (SST/Pulumi) accurate, retire unused infrastructure, and keep cost visible and justified.
  • Share your knowledge of production operations with the team, fostering a culture of learning and growth.

What they require

  • Bachelor's degree and 4-6 years of related experience or equivalent work experience.
  • 5+ years of experience in DevOps, site reliability, or platform operations, with significant responsibility for production systems.
  • 3+ years of hands-on experience with AWS, with an emphasis on serverless services (Lambda, SQS, EventBridge, CloudWatch, S3).
  • Strong database administration experience: PostgreSQL operations, backup and recovery, and query performance; comfort administering other data stores.
  • Proficiency in scripting languages such as TypeScript, Python, and bash for production automation and operational tooling.
  • Strong understanding of Linux, DNS, TLS, Docker, GitHub Actions, and infrastructure as code (SST, Pulumi, or Terraform).
  • Experience with production monitoring and alerting, incident response, and on-call ownership.
$104.1k–$140k/yr