Skip to main content
Grafana Labs

Staff Software Engineer - Databases SRE

RemoteUnited Kingdom, Sweden, Spain +1 more only
Published
Role
SRE
Experience
Staff
Employment
Full-time
Company size
Enterprise
€94k–€112.8k/yr
Check eligibility

Open to GB, SE, ES, DE only. Set where you work from to check your eligibility.

No BS summary

Staff-level SRE/software engineer with 8+ years engineering experience and 4+ years in SRE/CRE/production engineering. Must have strong Kubernetes experience on AWS, GCP, or Azure, infrastructure-as-code familiarity, SLO/incident response experience, and production multi-tenant systems experience. Hiring remotely from the UK, Sweden, Spain, or Germany.

Required skills

KubernetesAWSGCPAzureHelmTerraformJsonnetGoPythonJavaLinuxSLOs

What you'll do

  • Partner closely with product engineering squads in an embedded model
  • Own production reliability for high-SLA and complex customer environments
  • Design and implement automation to scale reliability practices
  • Ensure customers meet SLO targets
  • Define and evolve per-tenant SLOs and reliability models
  • Proactively reduce SLO burn to prevent repeat incidents
  • Serve as a primary escalation point and participate in on-call for incidents
  • Lead customer-impacting incident response and post-incident reviews
  • Contribute to design documents and code reviews
  • Influence feature design for production scalability and operability
  • Build automation to eliminate toil
  • Improve alert quality and reduce noisy escalations
  • Improve observability of customer environments
  • Design and implement reliability and scalability solutions
  • Develop fault-tolerant design patterns
  • Collaborate with engineering leaders on product strategy, roadmaps, and technical designs
  • Teach Site Reliability Engineering best practices
  • Participate in incident response through resolution, PIR, and customer communication

What they require

  • 8+ years engineering experience, including 4+ years in SRE, CRE, or production engineering
  • Strong preference for formal customer reliability engineering experience
  • Strong Kubernetes experience in AWS, GCP, or Azure
  • Familiarity with infrastructure-as-code tooling
  • Strong experience with technical leadership, leading projects, mentoring engineers, and acting as a force multiplier
  • Experience operating multi-tenant systems in production
  • Strong experience designing and implementing SLOs
  • Experience with one or more programming languages
  • Experience with Linux operating system internals
  • Knowledge of networking, cloud storage, and scaling
  • Excellent problem-solving and troubleshooting skills
  • Experience participating in blame-free incident response, following up on actions, and writing high-quality post-incident reviews
  • Ability to reason about performance, scaling, and failure modes
  • Comfortable working with autonomy and self-direction
  • Ability to partner deeply with product engineering teams

Benefits

  • Equity
  • Bonus if applicable
  • 100% remote global culture
  • Company-funded budget for AI coding assistants
  • Access to frontier AI models such as GPT-Codex 5/3, Claude Opus 4.6, and Gemini 3 Pro
  • Career growth pathways
  • In-person onboarding
  • Global annual leave policy of 30 days per year
  • Grafana Shutdown Days

Grafana Labs is the company behind Grafana Cloud, a fully managed observability platform built on open source, open standards, open ecosystems, and open culture. More than 35 million users and 7,000+ customers use Grafana Labs to ensure reliability of applications and systems, resolve incidents, and optimize telemetry.

🇸🇪 SwedenSoftwareEnterprise

What people say about this company

3.7/ 5

  • Employees appreciate the flexibility and remote work options.
  • The company culture is often described as open and collaborative.
  • Some employees mention a lack of clear career progression.

Details

Apply routeGreenhouse
€94k–€112.8k/yr