Principal Engineer - AI Platform & Operations
- Role
- AI / ML
- Experience
- Principal
- Employment
- Full-time
Open to US only. Set where you work from to check your eligibility.
No BS summary
Principal-level engineer for ML infrastructure and AI platforms, with 12+ years engineering experience and 4+ years at Principal or Distinguished level. Must know model serving, Kubernetes, cloud ecosystems, MLOps tools, Python, Docker/Helm/containerization, and LLM inference at scale. Remote role based in Washington or California State.
Core skills
Required skills
Serko is a cutting-edge tech platform in global business travel & expense technology. When you join Serko, you become part of a team of passionate travellers and technologists bringing people together, using the world’s leading business travel marketplace. We are proud to be an equal opportunity employer, and we embrace the richness of diversity, showing up authentically to create a positive impact. There's an exciting road ahead of us, where travel needs real, impactful change. With offices in New Zealand, Australia, North America, and China, we are thrilled to be expanding our global footprint, landing our new hub in Bengaluru, India. With a rapid growth plan in place, we’re hiring people from different backgrounds, experiences, abilities, and perspectives to help us build a world-class team and product. We are looking for a Principal Engineer to serve as the technical expert for our AI Platform & Operations. This isn’t just a role about maintaining infrastructure; it’s about owning the technical vision for the foundational systems that will power every AI product team across the company. Remote Role - can be based in either Washington or California State. Requirements By building robust, self-serve internal developer platforms, you’ll enable our engineers to deploy, monitor, and scale AI models with unprecedented efficiency and safety. Your work ensures that our "Agentic" future is reliable, cost-effective, and cutting-edge. What You’ll Be Doing Architect the Future: Define the long-term technical roadmap for our AI platform, covering everything from model serving and feature stores to experiment tracking and CI/CD for ML. Set the Gold Standard: Establish engineering benchmarks for deployment, versioning, A/B testing, and automated rollbacks. Optimize & Scale: Lead strategies for GPU/compute efficiency and cost optimization while managing the complexities of LLM inference (quantization, batching, and latency) at scale. Enhance Observability: Design sophisticated monitoring and alerting systems specifically tailored for AI workloads in production. Champion Reliability: Drive platform stability and partner with application teams to ensure our infrastructure meets evolving product needs. Technical Leadership: Mentor Senior engineers, lead architecture reviews, and evaluate the next generation of cloud services and ML frameworks. What You’ll Bring We are looking for a heavy-hitter in the ML infrastructure space who thrives on solving complex, distributed systems problems. Senior Expertise: 12+ years of engineering experience, with at least 4 years at a Principal or Distinguished level. ML Infrastructure Mastery: Expert-level knowledge of model serving (e.g., Triton, vLLM, Ray Serve) and deep experience with Kubernetes and cloud ecosystems (AWS/GCP/Azure). The MLOps Toolkit: Proven experience with tools like MLflow, Weights & Biases, or Kubeflow. Coding & Systems: High proficiency in Python and a "systems-thinking" approach to Docker, Helm, and containerization. AI Specialization: Hands-on experience operating LLM inference at scale and a deep understanding of the trade-offs between throughput and latency. Platform Mindset: A track record of building internal platforms that treat other engineers as the primary customer, drastically improving engineering velocity. Benefits At Serko, we aim to create a place where people can come and do their best work. This means you’ll be operating in an environment with great tools and support to enable you to perform at the highest level of your abilities, producing high-quality and delivering innovative and efficient results. Our people are fully engaged, continuously improving, and encouraged to make an impact. Some of the benefits of working at Serko are: A competitive base pay Medical Benefits Discretionary incentive plan based on individual and company performance Focus on development: Access to a learning & development platform and opportunity for you to own your career pathways Flexible work policy The pay range is between $168,000 - $230,000 USD as a base salary offering annually.
What you'll do
- Define the long-term technical roadmap for the AI platform, covering model serving, feature stores, experiment tracking, and CI/CD for ML.
- Establish engineering benchmarks for deployment, versioning, A/B testing, and automated rollbacks.
- Lead strategies for GPU/compute efficiency and cost optimization while managing LLM inference complexities including quantization, batching, and latency at scale.
- Design sophisticated monitoring and alerting systems tailored for AI workloads in production.
- Drive platform stability and partner with application teams to ensure infrastructure meets evolving product needs.
- Mentor Senior engineers, lead architecture reviews, and evaluate the next generation of cloud services and ML frameworks.
- Build robust, self-serve internal developer platforms enabling engineers to deploy, monitor, and scale AI models efficiently and safely.
What they require
- 12+ years of engineering experience, with at least 4 years at a Principal or Distinguished level.
- Expert-level knowledge of model serving and deep experience with Kubernetes and cloud ecosystems.
- Proven experience with tools like MLflow, Weights & Biases, or Kubeflow.
- High proficiency in Python and a systems-thinking approach to Docker, Helm, and containerization.
- Hands-on experience operating LLM inference at scale and a deep understanding of the trade-offs between throughput and latency.
- A track record of building internal platforms that treat other engineers as the primary customer, drastically improving engineering velocity.
- Thrives on solving complex, distributed systems problems.
- Remote Role - can be based in either Washington or California State.
Benefits
- A competitive base pay
- Medical Benefits
- Discretionary incentive plan based on individual and company performance
- Access to a learning & development platform and opportunity for you to own your career pathways
- Flexible work policy
- Great tools and support to enable you to perform at the highest level of your abilities
Serko is a tech platform in global business travel & expense technology, offering a business travel marketplace.