Skip to main content
Boundless

Applied AI/ML Engineer

RemoteUnited States only
Published
Role
AI / ML
Experience
Mid
Employment
Full-time
$175k–$250k/yr
Check eligibility

Open to US only. Set where you work from to check your eligibility.

No BS summary

Applied AI/ML engineer with 3+ years shipping production ML/AI systems. Must know Python, PyTorch, LLM inference serving with vLLM/SGLang/TensorRT-LLM, GPU execution basics, and RL/post-training methods. Public GitHub with at least 1 year of activity is mandatory.

Core skills

PyTorchLLM inferenceReinforcement learning

Required skills

vLLM/SGLang/TensorRT-LLMGRPO/PPO/DPO/SFTPythonCUDAGitHub

Optional skills

slimeprime-rlverifiers libraryMegatron-LMFSDPTP parallelismPP parallelismDP parallelism

What you'll do

  • Own AI features and products from prototype through production — model selection, serving, evaluation, and iteration — shipping working software rather than research artifacts.
  • Deploy and optimize LLM inference across the fleet using vLLM and SGLang.
  • Tune continuous batching, KV-cache management, quantization, speculative decoding, and multi-model routing to maximize throughput and minimize latency and cost per token.
  • Build and operate reinforcement-learning and post-training pipelines using slime (Megatron-LM + SGLang) and Prime Intellect (prime-rl + the Environments Hub / verifiers).
  • Design rewards and verifiers.
  • Orchestrate rollouts.
  • Handle weight synchronization.
  • Keep long-running training stable.
  • Build eval harnesses and benchmarks that measure quality, throughput, and cost together.
  • Use evals and benchmarks to drive fast, data-informed iteration.
  • Partner with Infrastructure on GPU scheduling and fleet utilization.
  • Partner with Product on what to build next and why.

What they require

  • 3+ years shipping ML/AI systems to production
  • Hands-on experience serving LLM inference with vLLM, SGLang, or TensorRT-LLM
  • Experience with RL / post-training methods (GRPO, PPO, DPO, or SFT), or strong adjacent experience and a clear desire to go deep here
  • Strong Python and PyTorch
  • Working understanding of GPU execution: batching, memory, and basic CUDA concepts
  • Comfort operating in ambiguity with a strong bias for action
  • Candidates must include a public GitHub profile in their application.
  • The GitHub profile should demonstrate a minimum of 1 year of activity/history.
  • Applications that do not include a GitHub profile, or show insufficient activity, will not be considered.

Benefits

  • Competitive salary (proposed band b/t US$175k and $250k annually) + equity allocation
  • Health, dental, vision (for U.S. employees; region-adjusted globally)
  • Flexible PTO
  • Professional development and conference travel budget
  • Remote-first with regular off-sites and a high-trust, high-velocity team environment

Boundless is coordinating GPU compute at scale and building toward becoming a leader in AI.

AI Infrastructure
$175k–$250k/yr