Перейти к основному содержимому
Backblaze

Site Reliability Engineer ll (DBA)

УдалённоArgentina, Colombia, Costa Rica +1 more только
Опубликовано
Роль
Не указано
Опыт
Мидл
Зарплата не указана
Проверьте доступность

Доступно для: AR, CO, CR, MX only. Укажите, откуда вы работаете, чтобы проверить доступность.

Коротко по делу

We are seeking a Site Reliability Engineer (SRE) with a DBA focus to help ensure the stability, scalability, and reliability of our production database systems - primarily Vitess (distributed MySQL) and Cassandra - alongside the rest of our services and infrastructure.

Ключевые навыки

Kubernetes/Docker

Обязательные навыки

LinuxMySQLSQLNoSQLPythonBashGoTerraform/Ansible/JenkinsPrometheus/Grafana/Catchpoint/ELK/FireHydrant

Желательные навыки

Experience with Vitess or other distributed/sharded MySQL systemsFamiliarity with ITIL/OSS practices and SLO/SLA'sExperience with cloud platforms (AWS, GCP, or Azure)

Чем предстоит заниматься

  • Operating and maintaining high-availability database systems—primarily Vitess (distributed MySQL) and Cassandra—against established architecture and runbooks.
  • Optimizing database performance through query tuning, indexing strategies, and schema design.
  • Executing documented backup, recovery, and replication procedures to ensure data durability; escalating architecture-level changes to senior DBA SREs.
  • Ensuring database security compliance and access control.
  • Support the availability and durability of critical services across production environments.
  • Monitor service health using SLIs, SLOs, and error budgets, and escalate issues when thresholds are at risk.
  • Participate in on-call rotations, incident response, and post-incident reviews to drive service improvements.
  • Follow established ITIL/OSS processes (incident, change, problem, and capacity management).
  • Develop automation for common operational tasks, reducing manual intervention and toil.
  • Contribute to monitoring, logging, and alerting frameworks (e.g., Prometheus, Grafana, Catchpoint, ELK), and help integrate runbooks with FireHydra.
  • Work with CI/CD pipelines, configuration management, and infrastructure as code tools (Terraform, Ansible, Jenkins).
  • Write scripts (Bash, Python, Go, etc.) to improve system reliability and efficiency.
  • Partner with engineering, product, and operations teams to support resilient system design and operations.
  • Assist in capacity planning and disaster recovery exercises.
  • Work with vendors and service providers to troubleshoot service issues and track SLA performance.
  • Document systems, share learnings, and help grow a reliability-minded engineering culture.
  • Contribute to playbooks, runbooks, and operational documentation.
  • Identify recurring issues and propose long-term improvements.
  • Promote reliability-focused practices within development and operations teams.

Что требуется

  • Bachelor’s degree in Computer Science, Engineering, or related field (or equivalent experience)
  • 2–4 years of experience in site reliability, systems engineering, or operations centered around database systems
  • Exposure to large-scale, production-grade systems
  • Solid Linux systems administration and troubleshooting skills
  • Familiarity with service reliability concepts - monitoring, alerting, incident response, and root cause analysis
  • Proficiency in at least one scripting language (Python, Bash, or Go)
  • Understanding of containers (Kubernetes, Docker) and microservices concepts
  • Knowledge of incident response and operational best practices
  • Hands-on experience with MySQL performance tuning, replication, and disaster recovery
  • Proficiency in SQL and NoSQL database management
  • Experience in a SaaS, service provider, or distributed systems environment
  • Familiarity with ITIL/OSS practices and SLO/SLA's
  • Strong problem-solving skills and willingness to learn new technologies
  • Experience with cloud platforms (AWS, GCP, or Azure)
  • Ability to work independently, take ownership, and drive projects from problem discovery through resolution

Преимущества

  • Remote work
  • Opportunity to work with cutting-edge storage technology
  • Collaborative, mission-driven culture
  • Commitment to diversity and inclusion
  • Golden: Employee stock ownership? (Note: The posting does not mention specific benefits, so I left empty)

Backblaze is the object storage leader in the open cloud movement, fueling customer success with cloud storage built purposefully to unlock budgets, unburden administrators, and unleash innovators.

Cloud StorageСредняяbackblaze.com
Зарплата не указана