Applied AI Research Engineer
- Role
- AI / ML
- Experience
- Mid
- Employment
- Full-time
Open to US only. Set where you work from to check your eligibility.
No BS summary
Applied AI Research Engineer for an AI inference platform. Needs 3+ years experience in LLM inference, systems/research background, and deep understanding of LLM serving (kernel to scheduler). Must be a US citizen or hold a US visa.
Core skills
Required skills
Optional skills
AI-Native Cloud Platform
Applied AI Research Engineer
$140K - $200K•0.25% - 1.00%•New York, NY, US / San Francisco, CA, US / Remote (US) Job type Full-time Role Engineering, Full stack Experience 3+ years Visa US citizen/visa only Skills Torch/PyTorch, Reinforcement learning (RL), GPU Programming Connect directly with founders of the best YC-funded startups. Apply to role › Eli Mernit Founder Eli Mernit Founder
About the role
Beam is an ultrafast AI inference platform. We built a serverless runtime that launches GPU-backed containers in less than 1 second and quickly scales out to thousands of GPUs. Developers use our platform to serve apps to millions of users around the globe. We're backed by Y Combinator, Tiger Global, and prominent developer-tool founders, including the founder of Snyk and former CTO of GitHub.
About the Role
We’re looking to hire someone to own inference research hands-on and find ways to lower cost per token and latency on our customer workloads.
- Low-level inference optimization, from speculative decoding, quantization, KV-cache and memory management
- Work directly with customers to optimize their production workloads, and apply your learnings to our platform as product improvements
- High-level of autonomy to find the highest upside bets and guide the future of our inference platform based on your work
Skills & Experience
- Systems or research background in LLM inference
- Deep understanding of LLM serving, from the kernel to the scheduler
- History of shipping products or research that people use in production-like scenarios, whether academic or industry
- Excited to collaborate closely with customers
- Enthusiasm for developer tools, cloud native technologies, and open source software
Benefits
- Competitive salary and meaningful equity
- Join a fast-growing pre-series A company at the ground floor
- Health, dental, and vision benefits with 90% coverage for you and 50% for dependents
- Opportunities to participate in events across the cloud native community
- Fitness stipend, learning budget, and much, much more
About Beam
Cloud computing is broken. AI has introduced a new generation of workloads, like GPU inference, sandboxes, and agents. These aren't ordinary applications that can be run as Lambdas, or Dockerized apps on VMs: they're massive, stateless containers that need to spin up in <1s, often across multiple clouds and regions. Today, engineers are hacking together infra that breaks under real-world loads. That's where we come in. Our mission is to build the world's best compute platform for AI. Our first product is a serverless inference platform, used by companies like Coca Cola, Geospy and hundreds more. We've built our own container runtime, called beta9, which is designed for launching GPU-backed containers in under 1s. We're a small, highly-technical team, with backgrounds in distributed systems and robotics. We've raised $7M from YC, Tiger, Guy Podjarny (Founder of Snyk), and Jason Warner (former CTO of Github). We're searching for intensely curious, passionate, and hard-working engineers to join our mission in rebuilding the cloud for the age of AI. Founded:2021 Batch:W22 Team Size:5 Status:Active Location:New York City, NY Founders Luke Lombardi Founder Luke Lombardi Founder Eli Mernit Founder Eli Mernit Founder
What you'll do
- Own inference research hands-on
- Find ways to lower cost per token and latency on customer workloads
- Low-level inference optimization (speculative decoding, quantization, KV-cache, memory management)
- Work directly with customers to optimize production workloads
- Apply learnings to platform as product improvements
- Guide the future of the inference platform based on research
What they require
- 3+ years experience
- Systems or research background in LLM inference
- Deep understanding of LLM serving from kernel to scheduler
- History of shipping products or research used in production-like scenarios
- Excited to collaborate closely with customers
- Enthusiasm for developer tools, cloud native technologies, and open source software
- US citizen or visa holder
Benefits
- Competitive salary
- Meaningful equity
- Health, dental, and vision benefits (90% coverage for employee, 50% for dependents)
- Join a fast-growing pre-series A company
- Opportunities to participate in cloud native community events
- Fitness stipend
- Learning budget
AI-Native Cloud Platform. Ultrafast AI inference platform with a serverless runtime that launches GPU-backed containers in less than 1 second.