Senior Site Reliability Engineer
- Role
- SRE
- Experience
- Senior
- Employment
- Full-time
Open to EG only. Set where you work from to check your eligibility.
No BS summary
Senior SRE with strong Python and hands-on Kubernetes/cloud experience who will do on-call incident response and build LLM-based agents to automate repetitive SRE toil. Requires 3+ years building production software with LLMs and real on-call/SRE experience. Must be able to operate in production, write fixes in application repos, and meet Saudi data residency/regulatory requirements.
Core skills
Required skills
Optional skills
About Mozn MOZN is a leading Enterprise AI company enabling organizations to make informed decisions in two critical domains: Financial Crime Prevention and Enterprise Knowledge Intelligence. We’re a diverse, collaborative team of innovators united by a shared purpose: to build AI that delivers tangible business value, builds trust, and empowers people and organizations with augmented intelligence. Our culture is built on the relentless pursuit of excellence and meaningful impact. If you’re passionate about working alongside exceptional talent on world-class AI, and you want the autonomy and runway to do the best work of your career, join us in shaping the future of intelligent enterprises. About the role We're hiring an AI Senior SRE: someone who carries a normal SRE workload — on-call rotation, incident response, root cause analysis, hands-on Kubernetes/cloud work — and builds the agentic layer that automates that workload over time. You are not exempt from operating production. You're the person best placed to know what should be automated, because you're the one doing it. This role exists because most reliability toil (triage, root-causing, remediation, SLO tracking, onboarding checks) is repetitive and well-defined enough to hand to an LLM-based agent with the right guardrails. Your job is to do the on-call/ops work like any SRE, then turn what you learn into agents — identify the workflow, build the agent, and earn the trust to let it act with increasing autonomy. What you'll do Carry a normal on-call rotation and act as a hands-on responder: investigate, fix, and document incidents yourself, exactly like any SRE on the team — especially for workflows that don' t have an agent yet. Go deep on application-level reliability, not just infrastructure: read and debug service code, understand business logic well enough to find the real root cause, and ship fixes or PRs directly into application repos when the fix belongs in the app, not the platform. Design, build, and ship LLM-based agents (using tools like Claude Code, OpenAI Codex, or Kimi K2/K3) that plug into our existing stack: Kubernetes, cloud APIs, Prometheus/Grafana/ELK/Datadog, PagerDuty/Slack. Define the "tool" interface each agent needs — the specific APIs, scripts, and read/write actions — and build safe wrappers around them. Set clear guardrails for every agent: what it may do autonomously vs. what it must propose for human approval, with a bias toward human-in-the-loop until an agent has earned trust. Own agent evaluation — define what "correct" and "safe" look like per agent, and build test/backtest suites against real historical incidents. Continuously tune prompts, context, and tool schemas as an agent' s scope grows. Partner with the SRE/platform team to find good agent candidates: repetitive, well-scoped, auditable workflows. Report on agent impact — MTTD/MTTR/MTTX movement, false positive/negative rates, and engineer-hours of toil removed. Keep a security- and compliance-first posture: audit trails for every autonomous action, least-privilege access to production, and alignment with Saudi data residency/regulatory requirements. Requirements 3+ years building production software with LLMs — agentic workflows, tool/function calling, multi-step planning, RAG — not just personal use of a chat assistant. Hands-on experience shipping real work with an agentic coding tool such as Claude Code, OpenAI Codex, or Kimi K2/K3. Strong Python (or similar) for building agent tooling, API wrappers, and orchestration. Real, hands-on SRE experience: comfortable being a primary on-call responder, running incident response, and doing root cause analysis under pressure — not just familiar with the concepts. Application-level debugging skill, not just infra: able to read a service' s codebase, trace a failure back to the actual line/logic causing it, and ship a fix yourself — SRE work here isn' t limited to restarting pods or scaling nodes. Solid hands-on Kubernetes and cloud provider experience (AWS/GCP/OCI/Azure) and fluency with observability tools (Prometheus, Grafana, Datadog, ELK) — both as an operator and as integration points for agents. Understands guardrails for autonomous systems: permissioning, approval gates, rollback paths, auditability. Can build trust with technical stakeholders — every agent starts with limited autonomy and has to earn more. Nice to have Experience in Saudi Arabia / MENA, ideally consulting across public or private sector clients. Familiarity with Terraform/Ansible, Docker, VM/on-prem setups — useful context for the agents you' ll build, not the core job. Experience building eval/backtest harnesses for LLM agents against historical incident data. Background in ML engineering, LLMOps, or platform engineering. Benefits You will be at the forefront of an exciting time for the Middle East, joining a high-growth rocket-ship in an exciting space You will be given a lot of responsibility and trust. We believe that the best results come when the people responsible for a function are given the freedom to do what they think is best The fundamentals will be taken care of: competitive compensation, top-tier health insurance, and an enabling culture so that you can focus on what you do best You will enjoy a fun and dynamic workplace working alongside some of the greatest minds in AI We believe strength lies in difference, embracing all for who they are and empowered to be the best version of themselves
What you'll do
- Carry a normal on-call rotation and act as a hands-on responder: investigate, fix, and document incidents.
- Read and debug service code, understand business logic, and ship fixes or PRs directly into application repositories.
- Design, build, and ship LLM-based agents that integrate with Kubernetes, cloud APIs, Prometheus/Grafana/ELK/Datadog, PagerDuty/Slack.
- Define tool interfaces for agents and build safe wrappers around them.
- Set guardrails for agents: define autonomous actions vs human approval and bias toward human-in-the-loop until trust is earned.
- Own agent evaluation: define correctness/safety and build test/backtest suites against historical incidents.
- Continuously tune prompts, context, and tool schemas as agent scope grows.
- Partner with SRE/platform team to identify repetitive, well-scoped, auditable workflows for agenting.
- Report on agent impact (MTTD/MTTR/MTTX, false positive/negative rates, engineer-hours removed).
- Maintain security- and compliance-first posture: audit trails, least-privilege access, and alignment with Saudi data residency/regulatory requirements.
What they require
- 3+ years building production software with LLMs — agentic workflows, tool/function calling, multi-step planning, RAG.
- Hands-on experience shipping real work with an agentic coding tool such as Claude Code, OpenAI Codex, or Kimi K2/K3.
- Strong Python (or similar) for building agent tooling, API wrappers, and orchestration.
- Real, hands-on SRE experience: comfortable being a primary on-call responder, running incident response, and doing root cause analysis under pressure.
- Application-level debugging skill: able to read a service's codebase, trace failures to the actual line/logic, and ship fixes.
- Solid hands-on Kubernetes and cloud provider experience (AWS/GCP/OCI/Azure) and fluency with observability tools (Prometheus, Grafana, Datadog, ELK).
- Understands guardrails for autonomous systems: permissioning, approval gates, rollback paths, auditability.
- Can build trust with technical stakeholders; start agents with limited autonomy and grow trust over time.
Benefits
- Competitive compensation
- Top-tier health insurance
- High responsibility and trust
- Dynamic workplace alongside leading AI talent
- Opportunity to be at the forefront of AI in the Middle East
MOZN is a leading Enterprise AI company enabling organizations to make informed decisions in two critical domains: Financial Crime Prevention and Enterprise Knowledge Intelligence.