Skip to main content
Grafana Labs

Staff Software Engineer - Databases SRE

RemoteUnited Kingdom, Sweden, Spain +2 more only
Published
Role
SRE
Experience
Staff
Company size
Enterprise
€117.6k–€141.1k/yr
Check eligibility

Open to GB, SE, ES, DE, IE only. Set where you work from to check your eligibility.

No BS summary

Staff-level SRE/software engineer with 8+ years engineering experience and 4+ years in SRE/CRE/production engineering. Must have Kubernetes on AWS/GCP/Azure, infrastructure-as-code, SLOs, multi-tenant production systems, Linux, and programming experience. Remote role hiring from the UK, Sweden, Spain, Germany, or Ireland.

Required skills

KubernetesAWSGCPAzureHelmTerraformJsonnetGoPythonJavaLinuxSLOs

What you'll do

  • Partner closely with product engineering squads in an embedded model
  • Own production reliability for high-SLA and complex customer environments
  • Design and implement automation to scale reliability practices
  • Ensure customers meet SLO targets
  • Define and evolve per-tenant SLOs and reliability models
  • Reduce SLO burn to prevent repeat incidents
  • Serve as a primary escalation point and on-call for relevant incidents
  • Lead customer-impacting incident response and post-incident reviews
  • Contribute to design docs and code reviews
  • Influence feature design for production scalability and operability
  • Build automation to eliminate toil
  • Improve alert quality and reduce noisy escalations
  • Review and create SLOs
  • Improve observability of customers within their environments
  • Design and implement reliability and scalability solutions
  • Develop fault-tolerant design patterns
  • Collaborate with engineering leaders on product strategy, roadmaps, and technical designs
  • Teach others Site Reliability Engineering best practices
  • Participate in incident response through resolution, PIR, and customer communication

What they require

  • 8+ years engineering experience, including 4+ years in SRE, CRE, or production engineering
  • Strong preference for formal customer reliability engineering experience
  • Strong experience with technical leadership, leading projects, mentoring engineers, and acting as a force multiplier
  • Experience operating multi-tenant systems in production
  • Strong experience designing and implementing SLOs
  • Experience with one or more programming languages
  • Experience with Linux operating system internals
  • Knowledge of networking, cloud storage, and scaling
  • Excellent problem-solving and troubleshooting skills
  • Experience participating in blame-free incident response and writing high-quality post-incident reviews
  • Ability to reason about performance, scaling, and failure modes
  • Comfortable working in an autonomous engineering team
  • Ability to partner deeply with product engineering teams

Benefits

  • Equity
  • Bonus if applicable
  • 100% remote global culture
  • Company-funded AI coding assistant usage budget
  • Access to frontier AI models such as GPT-Codex 5/3, Claude Opus 4.6, and Gemini 3 Pro
  • Career growth pathways
  • In-person onboarding
  • Global annual leave policy of 30 days per annum
  • 3 Grafana Shutdown Days

Grafana Labs is the company behind Grafana Cloud, a fully managed observability platform built on open source, open standards, open ecosystems, and open culture. More than 35 million users and 7,000+ customers use Grafana Labs to ensure reliability of applications and systems, resolve incidents, and optimize telemetry.

🇸🇪 SwedenSoftwareEnterprise

What people say about this company

3.7/ 5

  • Employees appreciate the flexibility and remote work options.
  • The company culture is often described as open and collaborative.
  • Some employees mention a lack of clear career progression.

Details

Apply routeGreenhouse
€117.6k–€141.1k/yr