Skip to main content

Principal Machine Learning Engineer, Artificial Intelligence (AI) Required, Work From Home

RemoteUnited States only
Published
Role
AI / ML
Experience
Principal
Employment
Full-time
Salary not disclosed
Check eligibility

Open to US only. Set where you work from to check your eligibility.

No BS summary

Principal ML engineer for hands-on architecture of large-scale AI/ML systems across training, inference, evaluation, deployment, and GPU infrastructure. Requires AI experience, deep learning/transformers, production large-scale ML models, ML frameworks, distributed training/inference, and GPU optimization. Remote role tied to San Francisco, CA / United States signals.

Core skills

Machine LearningDeep LearningGPU optimization

Required skills

Artificial IntelligencePyTorch/JAXDeepSpeed/FSDP/Megatron/ZeRO/Ray

Optional skills

vLLMTensorRT-LLMFasterTransformerRLHFPPODPOORPOApache Arrow

What you'll do

  • Architect and build large-scale ML systems spanning data, training, evaluation, inference, and deployment.
  • Design reproducible, high-performance training pipelines across GPU infrastructure.
  • Architect inference systems that balance latency, throughput, cost, and reliability at scale.
  • Design and maintain data systems for high-quality synthetic and real-world training data.
  • Implement evaluation pipelines covering performance, robustness, safety, and bias, in partnership with research leadership.
  • Own production deployment, including GPU optimization, memory efficiency, latency reduction, and scaling policies.
  • Collaborate closely with application engineering to integrate ML systems cleanly into backend, mobile, and desktop products.
  • Make pragmatic trade-offs and ship improvements quickly, learning from real usage.
  • Work under real production constraints: latency, cost, reliability, and safety

What they require

  • Strong background in deep learning and transformer-based architectures.
  • Artificial Intelligence (AI) experience required.
  • Hands-on experience training, fine-tuning, or deploying large-scale ML models in production.
  • Proficiency with at least one modern ML framework (e.g. PyTorch, JAX), and ability to learn others quickly.
  • Experience with distributed training and inference frameworks (e.g. DeepSpeed, FSDP, Megatron, ZeRO, Ray).
  • Strong software engineering fundamentals; you write robust, maintainable, production-grade systems.
  • Experience with GPU optimization, including memory efficiency, quantization, and mixed precision.
  • Comfort owning ambiguous, zero-to-one ML systems end-to-end.
  • A bias toward shipping, learning fast, and improving systems through iteration.
  • Preferred: Experience with LLM inference frameworks such as vLLM, TensorRT-LLM, or FasterTransformer.
  • Preferred: Contributions to open-source ML or systems libraries.
  • Preferred: Background in scientific computing, compilers, or GPU kernels.
  • Preferred: Experience with RLHF pipelines (PPO, DPO, ORPO).
  • Preferred: Experience training or deploying multimodal or diffusion models.
  • Preferred: Experience with large-scale data processing (Apache Arrow, Spark, Ray).

Benefits

  • medical insurance
  • Dental
  • Vision
  • Savings Plan Options
  • PTO

IT recruiting agencies and staffing companies helping companies hire IT talent.

Staffing
Salary not disclosed