Staff DevOps Engineer
- Role
- DevOps
- Experience
- Staff
- Employment
- Full-time
Open to VN, VD only. Set where you work from to check your eligibility.
No BS summary
Staff DevOps engineer in Vietnam with deep AWS, Kubernetes, Terraform, GitOps, CI/CD, observability, security, and scripting experience. Must communicate well in spoken and written English and work asynchronously with UK colleagues.
Core skills
Required skills
Required languages
Our mission is to become the global standard for addressing. Street addresses weren’t designed for 2026. They aren’t accurate enough to specify building entrances, and they don’t exist for parks, rural areas and many parts of the world. This makes it hard to find places and causes problems and inefficiencies on a global scale.
That’s why we created what3words. We divided the world into 3m squares and gave each square a unique combination of three words. It’s the easiest way to find and share precise locations.
Over the last year, what3words has been used in 193 countries, and our monthly active users continue to grow at an impressive pace. Our tech is used by emergency services, delivery companies, eCommerce businesses, ride-hailing apps and NGOs, and is integrated into the navigation systems of millions of cars around the world.
The role:
At what3words, we're proud of our tech stack. We have built a massively scalable, microservices-based architecture, based around Kubernetes on AWS. You'll be exposed to technologies such as EKS, Terraform, Terragrunt, Istio, Karpenter, Flux, GitHub Actions, Prometheus, Grafana, Go, Python, Lambda and CloudFront — with a little GCP alongside.
Our DevOps team is deliberately broad: the team owns platform engineering, SRE, observability and security. The remit is wide and the ownership is end-to-end, so you'll carry work from the design call through to how it behaves in production.
You'll work closely with colleagues in our UK office, so you'll need to be effective asynchronously and comfortable making your case in writing.
Responsibilities:
- Own our infrastructure as code as a governed product rather than a pile of configuration — a versioned, tested module library with clear ownership, change control and automated quality gates.
- Build the self-service paths that let product teams provision what they need without waiting on us, then measure adoption and iterate until they're genuinely used.
- Make reliability a property of the delivery path rather than something inspected afterwards, with service level objectives, dashboards and alerting wired in by default.
- Shift Left Security approach for workflows — policy as code, least privilege and secure-by-default templates, so that our teams can build on a foundation of good decisions.
- Treat cloud cost as an engineering responsibility, with commitment and allocation decisions made deliberately.
- Take open-ended platform problems through to durable outcomes against multi-quarter goals, and write the design docs behind decisions that outlast any single project.
- Share an on-call rotation with the team, and improve how we alert and how we respond.
Essential Skills:
- Multi-account AWS: Substantial production experience running AWS at scale across multiple accounts, with the identity and networking discipline that goes with them.
- Kubernetes in Production: Deep, hands-on experience owning cluster lifecycle and version upgrades, tuning auto scaling under real load, and debugging the failures that follow.
- Advanced Terraform: You've authored and versioned modules other people depend on, made deliberate state-management decisions, and run an orchestration or automation layer on top.
- GitOps: Practical experience with declarative delivery, and a clear grasp of what reconciliation gives you and what it takes away when something is wrong.
- CI/CD as a Platform: A track record of owning the pipelines other engineers ship through, including self-hosted runners or build agents and the build times and failure modes that come with them.
- Observability & Incident Response: Experience managing alert systems that the team can trust and be well prepared to respond to and resolve.
- Automation & Security: Strong scripting skills (Shell, plus Python or Go), with least privilege, secrets handling and supply chain awareness built into platform defaults.
- AI tools: Experience and enthusiasm in leveraging Agentic AI workflows to improve developer experience.
- Soft Skills: Excellent communication in spoken and written English.
Diversity & Inclusion
Our mission is to help everyone talk about everywhere, and we believe diverse perspectives make for a better company and better products too. We strongly encourage applications from underrepresented groups and are committed to equality and inclusivity in our hiring processes and company culture.
Benefits
We offer the following benefits to all permanent employees of what3words:
- Competitive salary
- Flexible working
- 6 week remote working (work from anywhere) policy
- 25 days holiday: plus the option to buy more!
- Share options
- Private health insurance
- Wellbeing Days
- Generous parental leave policies
- Family friendly policies
- Employee Assistance Programme (EAP)
- Lunch & learn sessions
- Team social budget
What you'll do
- Own infrastructure as code as a governed product with a versioned, tested module library, clear ownership, change control, and automated quality gates.
- Build self-service paths that let product teams provision what they need without waiting on DevOps, then measure adoption and iterate until they are genuinely used.
- Make reliability a property of the delivery path with service level objectives, dashboards, and alerting wired in by default.
- Apply a Shift Left Security approach to workflows using policy as code, least privilege, and secure-by-default templates.
- Treat cloud cost as an engineering responsibility, with commitment and allocation decisions made deliberately.
- Take open-ended platform problems through to durable outcomes against multi-quarter goals and write design docs behind long-lasting decisions.
- Share an on-call rotation with the team and improve alerting and incident response.
What they require
- Substantial production experience running multi-account AWS at scale, including identity and networking discipline.
- Deep hands-on experience with Kubernetes in production, including owning cluster lifecycle and version upgrades, tuning autoscaling under real load, and debugging failures.
- Advanced Terraform experience, including authoring and versioning modules other people depend on, deliberate state-management decisions, and running an orchestration or automation layer on top.
- Practical experience with GitOps and declarative delivery, including understanding what reconciliation gives and takes away when something is wrong.
- Track record of owning CI/CD pipelines other engineers ship through, including self-hosted runners or build agents and their build times and failure modes.
- Experience managing alert systems that the team can trust and be prepared to respond to and resolve.
- Strong scripting skills with Shell plus Python or Go, with least privilege, secrets handling, and supply chain awareness built into platform defaults.
- Experience and enthusiasm in leveraging Agentic AI workflows to improve developer experience.
- Excellent communication in spoken and written English.
- Able to work effectively asynchronously and make your case in writing with UK office colleagues.
Benefits
- Competitive salary
- Flexible working
- 6 week remote working work-from-anywhere policy
- 25 days holiday, plus the option to buy more
- Share options
- Private health insurance
- Wellbeing Days
- Generous parental leave policies
- Family friendly policies
- Employee Assistance Programme (EAP)
- Lunch & learn sessions
- Team social budget
what3words created a system that divides the world into 3m squares and gives each square a unique combination of three words to make precise locations easy to find and share.