Skip to main content
Unity

Senior Machine Learning Engineer, ML Infrastructure- Online

RemoteUnited States only
Published
Role
AI / ML
Experience
Senior
Employment
Full-time
Company size
Enterprise
$187.2k–$243.3k/yr
Check eligibility

Open to US only. Set where you work from to check your eligibility.

No BS summary

Senior ML infrastructure engineer to build and operate low-latency, production online model inference platform. Must have deep experience with PyTorch-based serving, Triton/other model servers, Kubernetes and distributed serving frameworks. US-remote candidates only (no visa sponsorship); Seattle/Washington area preferred.

Core skills

PyTorchTriton Inference ServerKubernetes

Required skills

TorchServeRay ServeTensorFlow ServingGKERayPythonmodel compilationquantizationdynamic batchingGPU accelerationautoscalingobservability/monitoringworkflow orchestration (Airflow or Flyte)

Optional skills

Ray TrainRay DataFlyteAirflowGPU kernel optimizationmodel packaging/validation tools

Required languages

English Professional

What you'll do

  • Design and operate large-scale online inference infrastructure that serves production ML models with low latency and high reliability.
  • Develop infrastructure supporting distributed training workflows using Pytorch, Ray Data, and Ray Train.
  • Integrate ML pipelines with workflow orchestration systems (e.g., Flyte, Airflow) to enable reliable multi-stage training workflows.
  • Optimize model performance through model compilation, GPU/CPU utilization improvements, request scheduling, kernel fusion, and runtime-level tuning.
  • Improve observability of ML systems through latency, throughput, error-rate, cost, saturation, and model-health monitoring.
  • Partner closely with ML engineers to support faster model iteration while maintaining production safety, scalability, and cost efficiency.
  • Improve reliability and reproducibility of model serving workflows, including model packaging, artifact validation, compatibility testing, and deployment automation.
  • Lead architectural improvements to make the online ML platform more robust, user-friendly, scalable, and cost-efficient.

What they require

  • Experience building and operating production-grade online ML inference systems such as NVIDIA Triton Inference Server, TorchServe, Ray Serve, TensorFlow Serving, or similar.
  • Experience optimizing inference workloads using dynamic batching, model compilation, quantization, GPU acceleration, GPU kernel optimization, caching, or runtime tuning.
  • Strong experience with distributed systems, Kubernetes, autoscaling, service reliability, and production observability.
  • Strong programming skills in Python with practical experience on production ML systems and high-scale services.
  • Experience with PyTorch and modern model deployment workflows, including model packaging, validation, and serving lifecycle management.
  • Experience designing infrastructure for safe model rollout, canary testing, A/B experimentation, and automated rollback.
  • Strong systems thinking with ability to reason about latency, throughput, reliability, scalability, and cost tradeoffs.
  • Proven ability to lead technical direction and influence architectural decisions across teams.
  • Relocation support not available.
  • Work visa/immigration sponsorship not available.
  • Sufficient knowledge of English to have professional verbal and written exchanges.

Benefits

  • Comprehensive health, life, and disability insurance
  • Commute subsidy
  • Employee stock ownership
  • Competitive retirement/pension plans
  • Generous vacation and personal days
  • Support for new parents through leave and family-care programs
  • Office food snacks
  • Mental Health and Wellbeing programs and support
  • Employee Resource Groups
  • Global Employee Assistance Program
  • Training and development programs
  • Volunteering and donation matching program

Unity [NYSE: U] is the world’s leading game engine, powering play for more than 3 billion consumers each month. Unity enables teams across industries like automotive, manufacturing, and healthcare to design, simulate, and collaborate in 3D.

GamingEnterpriseunity.com

Details

Visa sponsorshipNo
$187.2k–$243.3k/yr