Skip to main content
Thinking Machines

Product Manager - Deployment

RemoteUnited States only
Published
Role
Unknown
Experience
Lead
Employment
Full-time
$300k–$450k/yr
Check eligibility

Open to US only. Set where you work from to check your eligibility.

No BS summary

As Product Manager for Deployment, you will own how Thinking Machines' models and fine-tuned checkpoints go from training into production use.

Core skills

ML serving/infrastructure/deployment

Required skills

autoscaling/rollback/incident responsemodel serving frameworks/GPU scheduling/batching/quantization

Optional skills

ML inference infrastructuremodel serving frameworksGPU schedulingbatchingquantization

What you'll do

  • Own deployment strategy, roadmap, and success metrics for taking models and Tinker-trained checkpoints into production, in close partnership with infrastructure, research, engineering, and GTM
  • Define priority deployment paths and workflows across model serving, autoscaling, versioning, rollback, monitoring, and incident response
  • Work at engineering depth on serving architecture, latency and cost tradeoffs, reliability targets, capacity planning, and API/SDK surfaces for deployment
  • Build direct feedback loops with users deploying models in production, and turn scattered signals into a clear view of what's broken, what's missing, and what to prioritize next
  • Drive ambiguous workstreams end to end: technical scoping, dependency resolution, launch readiness, on-call/escalation design, and post-incident learning
  • Connect deployment decisions to the model and infrastructure roadmap, making visible the tradeoffs between flexibility, reliability, and operational cost
  • Shape SLAs, pricing/packaging inputs for hosted inference, and the operating model for a deployment platform expected to scale quickly
  • Do whatever work makes deployment succeed — reviewing a serving config, joining an incident retro, inspecting latency data, or writing the rollout plan for a new model

What they require

  • Experience owning a production ML serving, infrastructure, or deployment product, with direct involvement in reliability, scaling, or performance decisions
  • Track record working at engineering depth with production systems — comfortable discussing latency, throughput, autoscaling, rollback, or incident response in specifics
  • Experience taking a technical product from early usage through to reliable, scaled production use

Benefits

  • Generous health, dental, and vision benefits
  • Unlimited PTO
  • Paid parental leave
  • Relocation support as needed

American artificial intelligence company

$300k–$450k/yr