Skip to main content
Prime Intellect

Member of Technical Staff - Inference

RemoteUnited States only
Published
Role
Backend
Experience
Senior
Employment
Full-time
Company size
Startup
$150k–$300k/yr
Check eligibility

Open to US only. Set where you work from to check your eligibility.

No BS summary

Seeking an engineer with 3+ years of experience building and running large-scale ML/LLM services with clear latency/availability SLOs. Must have hands-on experience with LLM inference frameworks like vLLM, SGLang, or TensorRT-LLM, and a deep understanding of inference internals. Experience with distributed serving infrastructure and full-stack debugging is required. Python proficiency is essential.

Core skills

LLM ServingInference OptimizationRL Systems

Required skills

PythonPyTorchAWSGCPKubernetesCUDANCCLInfiniBandvLLMSGLangTensorRT-LLMNVIDIA Dynamo

Optional skills

CUDA kernel developmentTriton kernel developmentNsight Systems profilingNsight Compute profilingRustC++KafkaPubSub

Required languages

English

What you'll do

  • Build a multi-tenant LLM serving platform that operates across our cloud GPU fleets.
  • Design placement and scheduling algorithms for heterogeneous accelerators.
  • Implement multi-region/zone failover and traffic shifting for resilience and cost control.
  • Build autoscaling, routing, and load balancing to meet throughput/latency SLOs.
  • Optimize model distribution and cold-start times across clusters.
  • Integrate and contribute to LLM inference frameworks such as vLLM, SGLang, TensorRT-LLM.
  • Optimize configurations for tensor/pipeline/expert parallelism, prefix caching, memory management and other axes for maximum performance.
  • Profile kernels, memory bandwidth and transport; apply techniques such as quantization and speculative decoding.
  • Develop reproducible performance suites (latency, throughput, context length, batch size, precision).
  • Embed and optimize distributed inference within our RL stack.
  • Establish CI/CD with artifact promotion, performance gates, and reproducible builds.
  • Build metrics, logs, tracing; structured incident response and SLO management.
  • Document architectures, playbooks, and API contracts; mentor and collaborate cross-functionally.

What they require

  • 3+ years building and running large-scale ML/LLM services with clear latency/availability SLOs.
  • Hands-on with at least one of vLLM, SGLang, TensorRT-LLM.
  • Familiarity with distributed and disaggregated serving infrastructure such as NVIDIA Dynamo.
  • Deep understanding of prefill vs. decode, KV-cache behavior, batching, sampling, speculative decoding, parallelism strategies.
  • Comfortable debugging CUDA/NCCL, drivers/kernels, containers, service mesh/networking, and storage, owning incidents end-to-end.
  • Systems tooling and backend services.
  • LLM Inference engine development and integration, deployment readiness.
  • AWS/GCP service experience, cloud deployment patterns.
  • Running infrastructure at scale with containers on Kubernetes.
  • Architecture, CUDA runtime, NCCL, InfiniBand; GPU-aware bin-packing and scheduling across heterogeneous fleets.

Benefits

  • Cash Compensation Range of $150-300k with significant equity incentives
  • Flexible work arrangement (remote or San Francisco office)
  • Full visa sponsorship and relocation support
  • Professional development budget
  • Regular team off-sites and conference attendance
  • Opportunity to shape decentralized AI and RL at Prime Intellect
🇺🇸 United StatesAI InfrastructureStartupprimeintellect.ai

Details

Visa sponsorshipYes
$150k–$300k/yr