Skip to main content
SimScale GmbH

Senior SRE / Platform Engineer (m/f/d)

RemoteUTC-1…UTC+9
Published
Role
SRE
Experience
Senior
Salary not disclosed
Check eligibility

Open to UTC-1…UTC+9. Set where you work from to check your eligibility.

No BS summary

SimScale is looking for a Senior SRE / Platform Engineer to own and improve their cloud infrastructure, spanning AWS, EKS, observability, disaster recovery, security, multi-region architecture, elastic GPU/HPC capacity, and internal developer tooling. This is a hands-on senior individual contributor role within a small infrastructure team supporting 50+ engineers.

Core skills

AWS/GCPKubernetesOpenTelemetry

Required skills

Python/Go/Rust/JavaLinuxTerraformArgoCDPrometheus

What you'll do

  • Evolve our Kubernetes platform: Evaluate and adopt technologies such as Kubernetes Gateway API and service mesh patterns, and coordinate platform evolution across 10+ engineering teams.
  • Take observability to the next level: Drive organization-wide adoption of OpenTelemetry for distributed tracing and metrics, and help teams define meaningful SLOs.
  • Shape multi-region architecture and data residency: Support our move from an EU-centered footprint toward a global, multi-cloud architecture that satisfies disaster-recovery and data-residency requirements.
  • Own cloud cost and efficiency at scale: Keep petabyte-scale infrastructure cost-efficient, secure, and well-instrumented.
  • Improve tooling: Build self-service AWS account provisioning, guardrails and AI-assisted automations that help engineering teams manage infrastructure safely and efficiently at scale.

What they require

  • 5+ years of professional experience in SRE, platform, or infrastructure engineering.
  • Software development experience: Your background is rooted in software development, and you moved into SRE from there. You write production-quality software in at least one of Python, Go, Rust, or Java.
  • Strong systems foundation: You understand Linux internals and distributed systems well enough to debug complex production behavior.
  • Hands-on cloud and infrastructure experience: AWS (or GCP), declarative infrastructure (Terraform), gitops-workflow (ArgoCD) and container orchestration (Kubernetes).
  • Observability and reliability experience: You have worked with OpenTelemetry, Prometheus, distributed tracing, monitoring, and meaningful SLOs/SLIs.

Benefits

  • Join a dedicated, supportive team with unlimited growth opportunities and leadership potential
  • Make an impact quickly by sharing ideas and contributing to creative, goal-oriented projects
  • Work in a diverse, inclusive environment with colleagues from over 35 countries
  • Enjoy flexible hours and the freedom to work remotely from anywhere in the world
  • Access comprehensive health coverage, retirement plans, paid time off, and wellness support
🇩🇪 GermanyTechnology, Information And InternetMid-size
Salary not disclosed