Skip to main content
Alpaca
Alpaca

Senior DevOps Engineer

RemoteAmericas
Published
Role
Unknown
Experience
Senior
Salary not disclosed
Check eligibility

Open to Anywhere in Americas. Set where you work from to check your eligibility.

No BS summary

As a Senior DevOps Engineer at Alpaca, you will design, build, and operate the infrastructure that enables global scaling and trading-critical systems. You'll work with cloud architecture on GCP, Infrastructure-as-Code with Terraform, CI/CD pipelines, observability stacks, GKE clusters, and participate in on-call rotations.

Core skills

GCP/Terraform/GitOps/CI/CD/Kubernetes/GKE/Helm/Prometheus/Thanos/Grafana/Loki/Tempo/Alertmanager/PostgreSQL/RabbitMQ/IBM MQ/RedPanda

Optional skills

OPA/ConftestCheckovtflintAtlantisBackstageTiltAlloy collectorRootly

Required languages

English professional

What you'll do

  • Design and evolve our cloud architecture on GCP - networking, interconnects, IAM and high-availability topology - and express it entirely as code with Terraform, following GitOps as a first principle.
  • Build and own the CI/CD pipelines that plan, review, test and safely apply IaC changes - Policy-as-Code guardrails, drift detection and progressive rollout so infrastructure changes ship as confidently as application code.
  • Advance Platform-as-a-Product: build self-serve capabilities and paved paths so engineers can provision what they need, through a golden path rather than a hand-off.
  • Strengthen our observability stack - metrics, logs, traces and alerting across Prometheus, Thanos, Grafana, Loki, Tempo and Alertmanager - so the platform is easy to run and reason about.
  • Operate our GKE clusters and the infrastructure services that run on them - Helm-packaged workloads, message brokers (RabbitMQ, IBM MQ) and data stores.
  • Participate in our Follow-The-Sun on-call model: watch and triage alerts, join and declare incidents, lead structured debugging and escalation, and drive blameless post-mortems and the post-actions that actually close the loop.
  • Embed SRE practices - SLIs/SLOs and error budgets, capacity planning - into how Core Infrastructure builds and operates, working closely with our SRE function.

What they require

  • 5+ years in a DevOps, Platform/Infrastructure, or SRE role, with a proven track record operating large-scale, high-availability, high-performance systems in production.
  • Deep hands-on experience designing cloud architecture on Google Cloud Platform (GCP) as the primary cloud - landing zones, networking, IAM and high-availability topology.
  • Strong Infrastructure-as-Code skills with Terraform, structuring large codebases across multiple environments, with GitOps as a first principle and least-privilege as a default mindset.
  • Proven experience building CI/CD pipelines for IaC - automated plan/apply, code review, Policy-as-Code, drift detection and safe rollout.
  • Significant production experience with Kubernetes (ideally GKE) and packaging/deploying workloads with Helm.
  • Solid cloud and L3/L4-L7 networking fundamentals (VPCs, routing, load balancing, DNS, TLS, interconnects) and comfort debugging cross-service connectivity.
  • Hands-on experience with a modern observability stack - Prometheus, Thanos, Grafana, Loki, Tempo and Alertmanager - across metrics, logs, traces and alerting.
  • Operator-level familiarity with data stores such as PostgreSQL and Message Brokers (e.g. RabbitMQ, RedPanda) - able to run and troubleshoot them in production.
  • A good understanding of SRE practices - SLOs/error budgets, capacity planning - and a Platform-as-a-Product mindset.
  • Strong grasp of incident management end to end: joining and declaring incidents, structured debugging under pressure, escalation, clear documentation, and post-mortems that drive real change.
  • Able and willing to take part in a Follow-The-Sun on-call rotation from APAC hours, and to work effectively in a distributed, async-first team with strong written communication.

Benefits

  • Competitive Salary & Stock Options
  • Health Benefits
  • New Hire Home-Office Setup: One-time USD $500
  • Monthly Stipend: USD $150 per month via a Brex Card

Alpaca is a modern platform for trading. Alpaca's API is the interface for your trading algorithms, bots, or applications to communicate with Alpaca''s brokerage and other services.

🇺🇸 United StatesFinanceMid-sizealpaca.markets
Salary not disclosed