Перейти к основному содержимому
Tyk

Site Reliability Engineer

УдалённоВесь мир
Опубликовано
Роль
Фулстек
Занятость
Contract
Зарплата не указана
Проверьте доступность

Доступно для: Worldwide. Укажите, откуда вы работаете, чтобы проверить доступность.

Коротко по делу

We’re looking for a Site Reliability Engineer to manage, maintain, improve and provide support on our platform. You will be curious by nature, always looking for ways to improve, as we will look to you for new ideas, solutions and metrics on how we can improve the platform. You will also be our first line of incident management to our clients and will help define our response going forward.

Ключевые навыки

LinuxKubernetes & containersAWS / EKS

Обязательные навыки

Kubernetes/containersAWS/EKSTerraform/IaCHelmGoMongoDB/Redisprometheus/grafana/thanosDNS/TCP/IP/HTTP/TLS/UDP

Желательные навыки

GCPAzureBare metal infrastructure engineeringAPI management experienceLarge scale distributed storage managementFamiliarity with RancherCKA/CKAD/CKS certificatesCreating and delivering production software in Go language

Чем предстоит заниматься

  • Maintaining global Tyk Cloud within SL(A/I/O)s you will help to define
  • Identifying reliability issues and working together with your squad to solve them
  • Identifying and introducing new metrics and building relevant dashboards
  • Participating in the on-call rotation
  • Working with your squad to expand multi-region and multi-cloud reach of the platform
  • Documenting operational knowledge
  • Conducting post-incident analysis
  • Automating common tasks
  • Being a key shaper and contributor to our continuous improvement agenda - be it the clarity of our user stories, how we estimate, communicate with other teams or customers – we expect this role to be advocate of continuous improvement
  • Reliability of our new global Tyk Cloud platform
  • Automation of operations and support
  • Writing and maintaining documentation on SRE processes and policies
  • Recommending and implementing ways of driving operational efficiency and driving down our cost to run, without impacting service
  • Assisting in penetration testing for Cloud through liaising with our provider, providing technical details, and environment setup
  • Incident management

Что требуется

  • Strong collaboration skills
  • Launching and operating production scale kubernetes clusters
  • Designing and operating infrastructure on AWS and other providers
  • Operating MongoDB (or other document database) clusters
  • Operating Redis (or other key-value storage) clusters
  • Administering Linux servers
  • Maintaining distributed software
  • Operating Prometheus and Grafana
  • Operating logging collection and analysis systems
  • Participating in the on-call rotation(16:00pm – 4:00am UTC)

Преимущества

  • Everyone has unlimited paid holiday.
  • We have total flexibility in hours, as we believe creativity flows better when our people are given freedom to decide when they are most productive. Everyone is unique after all.
  • Employee share scheme
  • Generous maternity and paternity leave
  • Company retreats

Tyk

The Tyk API Management platform helps organisations connect systems and services and powers products and services across industries. Tyk is on a mission to connect every system in the world, starting with an API Management platform.

🇬🇧 ВеликобританияAPI ManagementСтартап
Зарплата не указана