Skip to main content
TOPdesk

Senior Site Reliability Engineer (m/f/d)

RemoteGermany only
Published
Role
SRE
Experience
Senior
Employment
Full-time
Company size
Mid-size
Salary not disclosed
Check eligibility

Open to DE only. Set where you work from to check your eligibility.

No BS summary

Senior SRE with 5+ years running Azure production cloud environments at scale. Must be strong with SLOs/error budgets, Kubernetes, Terraform, Python automation, Linux, observability, and incident response. Role is tied to Kaiserslautern, Germany, with remote work possible.

Core skills

AzureKubernetesTerraform

Required skills

Grafana/Prometheus/VictoriaMetricsHelmIngress controllersRBACPersistent volumesPythonCI/CDLinuxPuppetAnsible

Optional skills

Feature flagsAzure Cloud Adoption FrameworkClaude Code

What you'll do

  • Define and own service-level objectives across the Azure and potentially multi-cloud estate.
  • Use error budgets to steer the balance between shipping change and protecting reliability.
  • Identify toil, classify it, and engineer it out.
  • Feed self-healing automation and findings into the reliability roadmap.
  • Stand up an agent-based support layer that owns recurring toil and continuously feeds improvements back into reliability.
  • Standardise metrics, alerting, and tracing across all datacenters.
  • Close coverage gaps on cloud workloads.
  • Measurably reduce the alert-to-incident ratio from baseline.
  • Lead incidents to resolution.
  • Run blameless postmortems.
  • Turn every learning into a durable fix or an automation candidate.
  • Harden CI/CD and progressive delivery with canaries, safe rollouts, and automated rollback.
  • Model capacity and load-test critical paths.
  • Keep the platform within its performance envelope as it scales across regions.
  • Bring agents and bounded automation with observability, approvals, containment, and rollback into detection, diagnosis, and remediation.
  • Ensure every alert links to a runbook and every runbook links to an automation candidate.
  • Own capacity and cost planning across the multi-cloud estate.
  • Model usage and growth trends.
  • Forecast short and long term infrastructure needs so spend and scaling decisions stay ahead of demand.
  • Automate repeated manual work.
  • Measure before optimising using SLOs, baselines, and dashboards.
  • Design for failure and make recovery automatic and observable.
  • Pair with product engineering teams and transfer knowledge.
  • Treat cost and reliability as joint objectives.
  • Collaborate proactively with product teams during early product development and design.
  • Help teams make optimal choices and introduce appropriate SRE practices.

What they require

  • Proven hands-on experience (5+ years) as a Site Reliability, DevOps, or Infrastructure Engineer running a production cloud environment at scale (Azure).
  • Fluent with SLOs, error budgets, and reliability engineering practice — you have set them, not just read about them.
  • Strong observability skills at scale — Grafana, Prometheus, VictoriaMetrics, or equivalent — including alerting and tracing.
  • Kubernetes at operator level: Helm, namespace management, ingress controllers, RBAC, persistent volumes.
  • Coding for automation (Python or equivalent) and Terraform delivered via CI/CD.
  • Linux system administration — you understand what Puppet or Ansible is doing, not just whether it ran green.
  • Comfortable leading incidents in an on-call rotation with real SLA obligations, and the maturity to know when to escalate.
  • Strong written communication — your postmortems, runbooks, and architecture notes are unambiguous.
  • Preferred: Experience with progressive delivery — canaries, feature flags, automated rollback.
  • Preferred: Experience working within or migrating toward an Azure Cloud Adoption Framework or enterprise landing-zone structure.
  • Preferred: Current, personal practice of AI-native software delivery.
  • Preferred: Experience with EU data residency / sovereign cloud requirements.

Benefits

  • Permanent employment contract and 30 days of annual vacation
  • Pleasant working atmosphere with flat hierarchies
  • Open working atmosphere in an international environment
  • Flexible working hours within a modern working environment
  • Possibility to work remotely
  • Well-founded onboarding by a buddy
  • Time for individual training opportunities to further develop your personal strengths
  • Joint employee events and team building measures
  • Employee subsidy for gym membership
  • Company health measures such as health days or fresh fruit
  • Free drinks (coffee, tea, water)
  • Gifts on special occasions, e.g. anniversary
  • Monthly tax-free payment in the form of a Mastercard
  • Possibility of time off (sabbatical)
  • Quality time together (table football, table tennis table, massage chair)
  • Corporate benefits
  • Vacation bonus

TOPdesk builds service management software used across education, healthcare, government, and manufacturing.

ITSMEnterprise
Salary not disclosed