Skip to main content
Pliant

Engineering Manager - Site Reliability

RemoteGermany only
Published
Role
Engineering Management
Experience
Senior
Employment
Full-time
Salary not disclosed
Check eligibility

Open to DE only. Set where you work from to check your eligibility.

No BS summary

Engineering Manager for a new Site Reliability function, with 7–10 years of engineering experience and at least 3 years managing engineers. Must have hands-on production/reliability background, AWS, Terraform, platform observability, on-call and incident management experience. Remote role with applicant location requirement in Germany, despite the title mentioning EU/UK remote.

Core skills

AWSTerraformDatadog

Required skills

Claude CodeCursor

What you'll do

  • Define the framework other teams use to set their own SLOs and error budgets, educating and supporting product teams along the way.
  • Own blameless post-mortems and root-cause fixes.
  • Implement production readiness reviews so nothing new ships without one.
  • Close gaps in Datadog observability coverage, including missing alerts, dashboard blind spots, and noisy pages that erode trust in on-call.
  • Hire and build the team from the ground up, setting the technical and cultural bar for every engineer who joins after you.
  • Build the on-call rotation and incident process from scratch in the first few months.
  • Hire the first engineers onto the team.
  • Define SLOs for the most important services by mid-year.
  • Establish a real incident review process by mid-year.
  • Build Site Reliability into a function other teams route to by year one.
  • Drive repeat incidents down through root-cause fixes.

What they require

  • 7-10 years of engineering experience, including at least 3 years directly managing engineers.
  • Track record of hiring and developing engineers, with specific people levelled or promoted.
  • Hands-on production or reliability engineering background.
  • This is not a first management role.
  • Experience carrying a pager.
  • Strong AWS and Terraform experience, comfortable working inside a managed infrastructure-as-code pipeline.
  • Experience building or running an on-call rotation and incident management process, not just participating in one.
  • Strong platform observability experience.
  • Clear communication for a technical, cross-team audience.
  • Track record of pushing reliability practices upstream into product engineering teams, not just reacting to incidents after the fact.
  • Proficiency with AI-assisted development using Claude Code and Cursor.
  • Comfortable reviewing AI-written PRs as rigorously as any other.
  • Experience with PCI DSS, SOC 2, and ISO 27001 context is relevant to the stack/environment.
  • Kubernetes will increasingly shape reliability work in the Platform Core migration.

Benefits

  • Opportunity to work in a growing team with big responsibilities that thrives on a strong exchange of knowledge and excellence.
  • Attractive remuneration.
  • Choice of preferred OS, Windows or Mac.
  • Flat hierarchy and transparent communication in a relaxed, professional atmosphere.
  • Opportunity to develop your talent in a dynamic team with ambitious goals.
  • Flexibility and possibility to work remotely.
  • Pliant Card with monthly credit to explore the product and enjoy food with colleagues.

Pliant is a European fintech specializing in B2B payment solutions. Its modular, API-first platform helps businesses streamline spending, improve cash flow, and integrate payments into financial workflows.

🇩🇪 GermanyFintechStartupfullpliant.org/

Details

Apply routeDom
Salary not disclosed