Skip to main content
Mozn

Senior Site Reliability Engineer

RemoteEgypt only
Published
Role
SRE
Experience
Senior
Employment
Full-time
Salary not disclosed
Check eligibility

Open to EG only. Set where you work from to check your eligibility.

No BS summary

Senior SRE with strong Python and hands-on Kubernetes/cloud experience who will do on-call incident response and build LLM-based agents to automate repetitive SRE toil. Requires 3+ years building production software with LLMs and real on-call/SRE experience. Must be able to operate in production, write fixes in application repos, and meet Saudi data residency/regulatory requirements.

Core skills

KubernetesLLM-based agents

Required skills

PythonAWSGCPOCIAzurePrometheusGrafanaDatadogELKLLM agent workflowsRAGtool/function callingincident responseroot cause analysis

Optional skills

TerraformAnsibleDockerVM/on-prem setupsML engineeringLLMOpsplatform engineering

What you'll do

  • Carry a normal on-call rotation and act as a hands-on responder: investigate, fix, and document incidents.
  • Read and debug service code, understand business logic, and ship fixes or PRs directly into application repositories.
  • Design, build, and ship LLM-based agents that integrate with Kubernetes, cloud APIs, Prometheus/Grafana/ELK/Datadog, PagerDuty/Slack.
  • Define tool interfaces for agents and build safe wrappers around them.
  • Set guardrails for agents: define autonomous actions vs human approval and bias toward human-in-the-loop until trust is earned.
  • Own agent evaluation: define correctness/safety and build test/backtest suites against historical incidents.
  • Continuously tune prompts, context, and tool schemas as agent scope grows.
  • Partner with SRE/platform team to identify repetitive, well-scoped, auditable workflows for agenting.
  • Report on agent impact (MTTD/MTTR/MTTX, false positive/negative rates, engineer-hours removed).
  • Maintain security- and compliance-first posture: audit trails, least-privilege access, and alignment with Saudi data residency/regulatory requirements.

What they require

  • 3+ years building production software with LLMs — agentic workflows, tool/function calling, multi-step planning, RAG.
  • Hands-on experience shipping real work with an agentic coding tool such as Claude Code, OpenAI Codex, or Kimi K2/K3.
  • Strong Python (or similar) for building agent tooling, API wrappers, and orchestration.
  • Real, hands-on SRE experience: comfortable being a primary on-call responder, running incident response, and doing root cause analysis under pressure.
  • Application-level debugging skill: able to read a service's codebase, trace failures to the actual line/logic, and ship fixes.
  • Solid hands-on Kubernetes and cloud provider experience (AWS/GCP/OCI/Azure) and fluency with observability tools (Prometheus, Grafana, Datadog, ELK).
  • Understands guardrails for autonomous systems: permissioning, approval gates, rollback paths, auditability.
  • Can build trust with technical stakeholders; start agents with limited autonomy and grow trust over time.

Benefits

  • Competitive compensation
  • Top-tier health insurance
  • High responsibility and trust
  • Dynamic workplace alongside leading AI talent
  • Opportunity to be at the forefront of AI in the Middle East

MOZN is a leading Enterprise AI company enabling organizations to make informed decisions in two critical domains: Financial Crime Prevention and Enterprise Knowledge Intelligence.

AI
Salary not disclosed