Skip to main content
Zillow

Principal Software Engineer, Applied AI Services

RemoteMexico only
Published
Role
Backend
Experience
Principal
Salary not disclosed
Check eligibility

Open to MX only. Set where you work from to check your eligibility.

No BS summary

Principal Software Engineer for Applied AI Services. Requires 10+ years of experience in large-scale distributed backend systems, microservices, event-driven architectures, and cloud infrastructure (Kubernetes, AWS). Must have experience with data pipelines, ML workflows, and productionizing AI/ML/LLM systems. Role involves leading architecture and technical initiatives across multiple teams.

Core skills

Applied AILLMML

Required skills

microservicesevent-driven architecturescloud infrastructureKubernetesDatabricksSparkKafkaAirflowRESTGraphQL

Optional skills

personalizationrecommendationsuser intent modelingprompt designevaluation strategiessafety and guardrail patterns

Required languages

English

What you'll do

  • Architect end‑to‑end applied AI services that connect offline data ingestion, AI/ML/LLM workflows, and online services and APIs, defining shared patterns for batch and streaming data pipelines (e.g., Databricks, Spark, Kafka or equivalents), feature and signal stores, and evaluation and guardrail frameworks for AI‑powered capabilities.
  • Create reusable building blocks—such as libraries, templates, and reference implementations—that make it straightforward for product teams to integrate AI into their services and ship AI‑powered features faster.
  • Lead conventional backend and platform excellence by architecting and guiding high‑scale microservices in a Kubernetes environment, driving patterns for event‑driven architectures (including event schemas, contracts, and consumption patterns), and setting standards for databases, caching, and data access that support both HDP backend and AI use cases.
  • Drive AI/ML and LLM‑powered systems from prototype to production, including ingestion and transformation of training and inference data, integration of models and LLMs into online decision flows and APIs, and the definition of evaluation methodologies, metrics, and regression gates (e.g., LLM‑as‑judge, offline/online evaluation, human‑in‑the‑loop review loops).
  • Partner with AI/ML, Agentic AI, and data platform teams to clarify ownership boundaries and interfaces (for example, around cross‑cutting evaluation capabilities such as Evaluate MCP), and to ensure AI systems remain measurable, debuggable, and reproducible as they scale.
  • Lead multi‑team technical initiatives that span SJS, AI/ML teams, HDP, and other backend groups, defining and rolling out system‑wide standards and abstractions for APIs and contracts (REST/GraphQL, events, DRDCs), data schemas and lineage across offline and online paths, and observability, evaluation, and operational runbooks for AI‑powered services.
  • Mentor senior engineers, run deep design reviews, and champion the use of AI as a force multiplier for engineering—such as background agents for KTLO (library updates, security posture, config drift) and AI‑assisted design, implementation, and testing patterns—helping define which workflows should be agent‑assisted versus human‑led in safe, observable, and cost‑effective ways.

What they require

  • 10+ years of software engineering experience with a strong track record of delivering and scaling complex, distributed backend systems in large engineering organizations.
  • You have built large‑scale microservices and event‑driven architectures in cloud environments (AWS or equivalent), ideally including Kubernetes, and you bring strong expertise in databases, caching, and data‑intensive services, including schema design, performance optimization, and reliability.
  • You are experienced with data pipelines and ML workflows (e.g., Databricks, Spark, Kafka, Airflow or equivalents) and how they connect to online systems, and you have designed system‑wide abstractions, frameworks, or platforms that are used by multiple teams.
  • You have hands‑on experience building or scaling AI/ML or LLM‑powered systems in production, including integrating models or LLMs into production services, owning or co‑owning data ingestion, feature pipelines, or model‑serving paths, and defining or implementing evaluation and guardrail mechanisms.
  • You have led cross‑team technical initiatives as an IC, influencing architecture and standards beyond a single team, and you communicate clearly with engineers, product partners, and leadership to drive clarity and alignment in ambiguous problem spaces.
  • You are familiar with personalization, recommendations, or user intent modeling domains, and you have experience or interest in LLM‑based workflows, including prompt design, evaluation strategies, and safety and guardrail patterns.
  • You have experience with modern API and integration layers, such as GraphQL or similar patterns that sit between backend services and user‑facing clients, and you have helped modernize legacy services and data paths into cohesive, platform‑aligned systems.
  • You care deeply about reliability, observability, and operational excellence, and you design systems for long‑term maintainability, debuggability, and measurability from the start.

Benefits

  • competitive base salary and benefits
  • equity awards based on factors such as experience, performance and location

At Zillow, we’re reimagining how people move—through the real estate market and through their careers. As the most-visited real estate platform in the U.S., we help customers navigate buying, selling, financing and renting with greater ease and confidence.

Real EstateEnterprisezillow.com/

What people say about this company

3.8/ 5

Salary not disclosed