Skip to main content
ClickHouse

Senior Site Reliability Engineer

RemoteEMEA
Published
Role
SRE
Experience
Senior
Company size
Startup
Salary not disclosed
Check eligibility

Open to Anywhere in EMEA. Set where you work from to check your eligibility.

No BS summary

Senior SRE with 8+ years in site reliability or related work. Must have production ClickHouse experience, Go or Python, cloud platforms, SQL, container orchestration, and automation/config management tools. Requires valid work authorization in the UK, Netherlands, Sweden, Germany, France, Spain, Italy, Portugal, Poland, or Czech Republic.

Core skills

ClickHouseGo/Python

Required skills

AWS/Azure/GCPSQLKubernetes/Docker SwarmAnsible/Terraform/Puppet

What you'll do

  • Build and lead processes to ensure reliability, availability, scalability, and performance of the cloud infrastructure running ClickHouse databases.
  • Collaborate with Control Plane, Dataplane, Core, Security, Support, and Operations teams to design and implement scalable, secure, highly available, and fault-tolerant distributed systems.
  • Own incident management and response.
  • Own post-mortem analysis, including running blameless postmortems.
  • Drive continuous improvement of ClickHouse services.
  • Develop software platforms and tools to optimize operational and engineering efficiency of ClickHouse Cloud.
  • Collaborate with engineering teams to design and implement scalable, secure, and highly available systems for ClickHouse.
  • Establish and manage service level objectives and service level agreements for ClickHouse Cloud.
  • Ensure ClickHouse Cloud infrastructure components, including Dataplane, Control Plane, and ClickHouse Core, have monitoring and alerting for timely incident detection and resolution.
  • Enhance and refine incident response processes and post-mortem analysis for ClickHouse Cloud outages, including working with support to communicate with impacted customers.
  • Continuously improve reliability and performance of ClickHouse services.
  • Plan, enable, and drive Chaos initiatives across Engineering teams based on internal priorities.
  • Manage on-call processes for performance and reliability issues.
  • Establish best practices for coordinating escalation to resolve issues and minimize downtime.

What they require

  • Bachelor’s or Master’s degree in Computer Science or a related field.
  • At least 8 years of experience in Site Reliability Engineering or a related field.
  • Previous experience using ClickHouse in production.
  • Hands-on experience with Go and/or Python.
  • Strong knowledge of cloud computing platforms such as AWS, Azure, or Google Cloud Platform.
  • Excellent understanding of distributed databases and SQL.
  • Hands-on experience with container orchestration tools such as Kubernetes or Docker Swarm.
  • Strong experience with automation and configuration management tools such as Ansible, Terraform, or Puppet.
  • Strong problem-solving ability and solid production debugging skills.
  • Passion for efficiency, availability, scalability, and data governance.
  • Ability to thrive in a fast-paced environment and partner with the business to move it forward.
  • High level of responsibility, ownership, and accountability.
  • Excellent communication and interpersonal skills.
  • Valid work authorization status in the UK, Netherlands, Sweden, Germany, France, Spain, Italy, Portugal, Poland or Czech Republic.

Benefits

  • Flexible work environment; ClickHouse is globally distributed and remote-friendly.
  • Healthcare employer contributions.
  • Stock options for every new team member.
  • Flexible time off in the US and generous entitlement in other countries.
  • $500 home office setup for remote employees.
  • Opportunities to engage with colleagues at company-wide offsites.

open-source Column-oriented DBMS (columnar database management system) for online analytical processing (OLAP)

Cloud ComputingStartupclickhouse.com

Details

Apply routeGreenhouse
Salary not disclosed