Skip to main content
Thoughtworks

Senior Service Reliability Engineer

RemoteNot specified. Estimate: worldwide · 83% confidenceArchived
Published
Role
SRE
Experience
Senior
Employment
Full-time
Salary not disclosed
Check eligibility

The listing doesn't say where it hires from. It may be open worldwide (83% confidence). This is an estimate, not an eligibility rule; verify before applying.Signals: company workplace model and company offices.

No BS summary

Senior SRE with 5+ years of experience in reliability, resilience, and system performance. Must have hands-on experience with Python, Go, or Bash, and be excellent with Terraform. Good understanding of Azure, network, and security is required. Experience with observability tools and Kubernetes is essential. Must be proficient in English and willing to be part of a 24x7 on-call rotation.

Core skills

TerraformKubernetesObservability

Required skills

PythonGoBashAzureAWS EKSDocker SwarmNomadmicroservicesserverless functionsNoSQLRESTful APIsDevOpsGitOps

Optional skills

GrafanaDatadogNewRelicELK StackDynatrace

Required languages

English

What you'll do

  • Improve site reliability by building mechanisms/architectures that enable fault tolerance and faster median time to respond and median time to detect.
  • Drive the integration of observability automation into the CI/CD pipeline.
  • Handle production incidents, manage incident communication with clients and draft root cause analysis documents.
  • Monitor performance of production systems and improve their scaling to ensure business goals are met within expected SLA and SLO metrics.
  • Work closely with application development teams as advisors on improving system reliability and assisting in implementation for reliability improvements.
  • Improve system observability across multiple facets such as logging and metrics, reducing false alarms to eliminate unnecessary toil and improving process efficiency.
  • Implement chaos engineering practices as necessary to test system reliability, setting up processes for such testing to be done regularly.
  • Achieve application availability with minimum/no disruption (99.999%) if necessary for business.

What they require

  • You have hands-on experience in programming and scripting languages such as Python, Go or Bash.
  • Excellent in Infra as a Code using Terraform
  • You have a good understanding of Azure
  • Experience on Network and security
  • You have had exposure to observability tools such as Grafana, Datadog, NewRelic, ELK Stack, Dynatrace or equivalent and you are proficient in using data from these tools to dissect and identify root causes of system and infrastructure issues .
  • You are familiar with DevOps and GitOps practices .
  • You have a good knowledge of container-based architecture and orchestration tools such as Kubernetes, AWS EKS, Docker Swarm, Nomad, etc.
  • You understand technical architecture and modern design patterns, including microservices, serverless functions, NoSQL and RESTful APIs, with experience in fixing bugs, analyzing logs, building metrics and operational dashboards.
  • You are familiar with creating infrastructure resources for improving reliability of system that follows Cloud’s Well Architected Framework principles: Reliability, security, cost optimization, performance efficiency and operational.
  • You have strong communication and articulation skills, and are proficient in English.
  • You have good people skills with an emphasis on negotiation and close collaboration with multiple cross-functional teams from the client side and/or Thoughtworks.
  • You solve challenging problems and difficult to debug issues with a never give up attitude.
  • You have the ability to work under pressure and with composure during production incidents.
  • You can confidently recommend improvements backed by strong technical arguments to client stakeholders or application development teams.
  • You are able to understand requirements provided by the client on both technical and business aspects and break them down for successful implementation.
  • You have a strong drive and ownership mentality, with a willingness to sign up for and deliver work when called upon, without being too concerned about role boundaries.
  • You’re willing to be part of a rotation- and need-based 24x7 available team.

Benefits

  • There is no one-size-fits-all career path at Thoughtworks: however you want to develop your career is entirely up to you.
  • But we also balance autonomy with the strength of our cultivation culture.
  • This means your career is supported by interactive tools, numerous development programs and teammates who want to help you grow.
  • We see value in helping each other be our best and that extends to empowering our employees in their career journeys.

Thoughtworks is a dynamic and inclusive community of bright and supportive colleagues who are revolutionizing tech. As a leading technology consultancy, we’re pushing boundaries through our purposeful and impactful work. For 30+ years, we’ve delivered extraordinary impact together with our clients by helping them solve complex business problems with technology as the differentiator.

🇺🇸 United StatesTechnology ConsultingEnterprisethoughtworks.com/
Salary not disclosed