Skip to main content
Jalasoft

Site Reliability Engineer - Azure, Observability and Scripting

RemoteColombia, Brazil, Peru +2 more only
Published
Role
SRE
Experience
Senior
Employment
Full-time
Salary not disclosed
Check eligibility

Open to CO, BR, PE, BO, AR only. Set where you work from to check your eligibility.

No BS summary

SRE with 6+ years overall and 3+ years running Kubernetes in production. Must know Azure observability, Prometheus/Grafana, SLOs, incident response, IaC, Azure DevOps pipelines, scripting, and professional working English.

Core skills

KubernetesAzure MonitorPrometheus

Required skills

Log AnalyticsKQLGrafanaTerraform/BicepAzure DevOpsPython/PowerShell/Bash

Optional skills

Azure Managed PrometheusAzure Managed GrafanaOpenTelemetryAzure BackupAzure Site RecoveryOracleAKSCKA

Required languages

English Professional working English required=true.

What you'll do

  • Help ensure the reliability, scalability, and performance of cloud-native platforms running on Microsoft Azure and Kubernetes.
  • Improve system availability, monitoring, incident response, and infrastructure automation in production environments.

What they require

  • 6+ years of experience
  • 3+ years of experience operating Kubernetes in production
  • Site reliability engineering or production operations for Kubernetes workloads at scale.
  • Azure Monitor, Log Analytics and KQL, including workspace design, data collection rules and retention strategy.
  • Prometheus and Grafana: metrics and exporters, recording and alerting rules, and dashboard design.
  • Definition and implementation of service level indicators, objectives and error budgets.
  • Alerting and incident response design, including runbook authoring and on-call practice.
  • Backup, restore and disaster recovery design and testing, including validation against RPO and RTO targets.
  • Kubernetes operations: workload troubleshooting, resource management and cluster upgrades.
  • Ability to read and modify infrastructure as code (Terraform or Bicep) and Azure DevOps pipelines.
  • Scripting in Python, PowerShell or Bash.
  • Professional working English.
  • Preferred: Chaos engineering or structured game day practice.
  • Preferred: Cost and capacity management for AKS estates.
  • Preferred: Incident management tooling and postmortem practice.
  • Preferred: Certification: CKA, AZ-400 or equivalent.

Benefits

  • Remote work
  • 13 floating holiday
  • 15 vacation days per year completed
  • Good working environment

Jalasoft is a world-class technology company delivering software solutions across LATAM, with more than 1,000 engineers and a strong focus on training, education, and intellectual property development.

SoftwareEnterprise
Salary not disclosed