Skip to main content
Stord

Senior Site Reliability Engineer

RemoteUnited States only
Published
Role
SRE
Experience
Senior
Employment
Full-time
Company size
Mid-size
Salary not disclosed
Check eligibility

Open to US only. Set where you work from to check your eligibility.

No BS summary

Senior SRE with 5+ years building and operating production infrastructure on GCP (GKE, Cloud Run, AlloyDB). Strong Terraform, Kubernetes, container, and observability (Datadog/Prometheus) skills; able to write automation in TypeScript/Python/Go. Remote — United States only.

Core skills

GKETerraformDatadog

Required skills

GCPCloud RunAlloyDBDockerKubernetesTypeScriptPythonGoGit

Optional skills

PostgreSQLRedisClickHouseKafkaRedpandaCloudflareGCP certificationsmulti-cloud

What you'll do

  • Own architecture and implementation of scalable, reliable infrastructure on GCP, including GKE, Cloud Run, AlloyDB, and networking.
  • Own Infrastructure as Code in Terraform: modules, org policies, and the patterns the team builds on.
  • Manage containerized workloads on Kubernetes, including performance tuning, capacity planning, and resource optimization.
  • Drive down cost and toil through better defaults, right-sizing, and automation.
  • Build monitoring, alerting, and observability in Datadog (APM, logs, RUM).
  • Define reliability signals and hold the line on them.
  • Develop and maintain disaster recovery and business-continuity strategies and validate them.
  • Design and maintain CI/CD pipelines in GitHub Actions, including runner strategy and deployment safety.
  • Automate operational workflows and infrastructure provisioning.
  • Build custom tooling and scripts to remove recurring operational pain.
  • Partner with data and development teams to improve deployment practices and application reliability.
  • Provide escalation support for production incidents, lead post-incident reviews, and turn findings into durable fixes.
  • Participate in technical design reviews and offer architectural input across teams.
  • Participate in on-call for critical systems and help improve SRE and infrastructure best practices.

What they require

  • 5+ years in SRE, platform, or infrastructure engineering.
  • Strong hands-on experience with GCP core services (GKE, Cloud Run, AlloyDB, networking, IAM).
  • Fluent in Docker and Kubernetes and experience debugging, tuning, and scaling workloads.
  • Deep Terraform experience: write reusable modules and manage state.
  • Productive in a programming language such as TypeScript, Python, or Go to build tooling and automation.
  • Experience building actionable monitoring and alerting (Datadog or equivalents).
  • Knowledge of distributed systems fundamentals: failure modes and consistency.
  • Experience with Git and collaborative development workflows and code review.
  • Experience running incidents and post-mortems and calm incident handling.
  • Ownership, strong communication, collaborative approach, production mindset, and learning agility.
  • Directed AI-assisted development: familiarity with AI coding tools while maintaining quality and judgment.

Stord is The Consumer Experience Company, powering seamless checkout through delivery for today's leading brands. Stord manages over $10 billion of commerce annually through its fulfillment, warehousing, transportation, and operator-built software suite including OMS, Pre- and Post-Purchase, and WMS platforms. Stord is headquartered in Atlanta with facilities across the United States, Canada, and Europe. Stord is backed by top-tier investors including Kleiner Perkins, Franklin Templeton, Founders Fund, Strike Capital, Baillie Gifford, and Salesforce Ventures.

🇺🇸 United StatesLogisticsEnterprisestord.kommune.no/
Salary not disclosed