Skip to main content
Beam

Applied AI Research Engineer

RemoteUnited States only
Published
Role
AI / ML
Experience
Mid
Employment
Full-time
$140k–$200k/yr
Check eligibility

Open to US only. Set where you work from to check your eligibility.

No BS summary

Applied AI Research Engineer for an AI inference platform. Needs 3+ years experience in LLM inference, systems/research background, and deep understanding of LLM serving (kernel to scheduler). Must be a US citizen or hold a US visa.

Core skills

PyTorchGPU ProgrammingLLM Inference

Required skills

TorchReinforcement LearningLLM ServingKernel ProgrammingScheduler

Optional skills

Cloud Native TechnologiesOpen Source SoftwareDeveloper Tools

What you'll do

  • Own inference research hands-on
  • Find ways to lower cost per token and latency on customer workloads
  • Low-level inference optimization (speculative decoding, quantization, KV-cache, memory management)
  • Work directly with customers to optimize production workloads
  • Apply learnings to platform as product improvements
  • Guide the future of the inference platform based on research

What they require

  • 3+ years experience
  • Systems or research background in LLM inference
  • Deep understanding of LLM serving from kernel to scheduler
  • History of shipping products or research used in production-like scenarios
  • Excited to collaborate closely with customers
  • Enthusiasm for developer tools, cloud native technologies, and open source software
  • US citizen or visa holder

Benefits

  • Competitive salary
  • Meaningful equity
  • Health, dental, and vision benefits (90% coverage for employee, 50% for dependents)
  • Join a fast-growing pre-series A company
  • Opportunities to participate in cloud native community events
  • Fitness stipend
  • Learning budget

AI-Native Cloud Platform. Ultrafast AI inference platform with a serverless runtime that launches GPU-backed containers in less than 1 second.

🇺🇸 United StatesCloud ComputingStartup

Details

Visa sponsorshipNo
$140k–$200k/yr