Skip to main content
Venice

Engineering Lead, Inference Optimization

RemoteUnited States only
Published
Role
Unknown
$270k–$330k/yr
Check eligibility

Open to US only. Set where you work from to check your eligibility.

Core skills

LLM inference optimizationcontinuous batchingKV cache management plausible

Required skills

PythonRustGovLLMSGLangCUDACUDA GraphsPagedAttentiontorch.compileNsight SystemsNsight ComputePyTorch Profiler

Optional skills

C++CUDA

Required languages

English required

What you'll do

  • Own Venice’s technical strategy for inference performance
  • Recruit and lead the Inference Optimization Team at Venice
  • Optimize Venice's GPU infrastructure across a range of architectures (e.g. H200s, B300s)
  • Improve latency, throughput, and cost per token for LLM inference workloads
  • Build reproducible benchmarking harnesses across inference engines (e.g. vLLM, SGLang) to identify the optimal engine, quantization scheme, and parallelism strategy per workload and GPU SKU
  • Work with our inference routing system to optimize multivariate inference load-balancing algorithms
  • Evaluate emerging inference optimization techniques (custom CUDA/Triton kernels), novel attention variants, new quantization schemes, and compilation stack improvements
  • Evaluate emerging inference hardware expansion (FPGAs, ASICs, custom silicon) for viability in Venice's stack.

What they require

  • 8+ years in performance optimization or HPC, with deep GPU architecture and parallel programming knowledge
  • 5+ years experience leading engineering teams
  • Proficiency in Python, Rust, or Go
  • Hands-on experience with at least one production LLM inference engine running at high volume in production
  • Demonstrated experience with LLM inference techniques: continuous batching, state, quantization, CUDA graphs, and torch.compile
  • Fluency with quantization tradeoffs, both qualitative and quantitative
  • Experience with distributed strategies (tensor parallelism, pipeline parallelism, MoE parallelism) in multi-GPU and multi-node environments
  • Fluency with GPU profiling and bias toward measuring before optimizing.

Venice is the world's leading consumer AI company built on principles of privacy, free speech, and user sovereignty. We're building the Port City of AI, in which millions of individuals, third party apps, and AI agents gather, interact, and access sophisticated AI resources.

AIStartup
$270k–$330k/yr