Skip to main content
Cerence

Sr. Principal Software Engineer

RemoteUnited States only
Published
Role
AI / ML
Experience
Principal
$185k–$280k/yr
Check eligibility

Open to US only. Set where you work from to check your eligibility.

No BS summary

Senior Principal Software Engineer with proven experience optimizing ML inference performance in production. Requires deep understanding of GPU architecture and hands-on experience with CUDA and low-level performance tuning. Must be able to deploy models beyond research environments.

Core skills

CUDALLM InferenceGPU Architecture

Required skills

vLLMTensorRT-LLMllama.cppQAIRTINT8INT4FP4FP8AWQGPTQbatchingcontinuous batchingspeculative decoding

Required languages

English

What you'll do

  • Optimize and deploy high-performance LLM inference pipelines
  • Own inference runtimes across data center, edge, and embedded platforms
  • Push model performance through quantization, kernel fusion, and cache optimization
  • Drive latency and throughput improvements that directly impact production products
  • Enable efficient, reliable deployment without external vendor dependency
  • Build deep expertise and ownership of: vLLM, TensorRT-LLM, llama.cpp, QAIRT
  • Extend and tune inference engines using custom CUDA kernels
  • Adapt runtimes for constrained and embedded deployment environments
  • Implement and evaluate quantisation strategies: INT8, INT4, FP4, FP8, mixed precision AWQ GPTQ
  • Balance accuracy, latency, memory footprint, and throughput
  • Optimize key-value cache performance through: Paging, Prefix caching, Cache-aware memory layout design
  • Reduce memory pressure while sustaining high throughput
  • Design and tune: Batching strategies, Continuous batching, Speculative decoding
  • Optimize tail latency and tokens/sec under real production traffic patterns

What they require

  • Proven experience optimizing ML inference performance in production
  • Deep understanding of GPU architecture and memory hierarchies
  • Hands-on experience with CUDA and low-level performance tuning
  • Experience deploying models beyond research environments
  • Models deploy efficiently on edge and embedded devices, not just servers
  • Tokens/sec significantly outperform baseline implementations
  • End-to-end latency is minimized and predictable
  • Inference cost per request is materially reduced
  • Deploy efficiently on edge or embedded targets
  • Achieve competitive tokens/sec
  • Reduce and stabilize inference latency
  • You will be responsible for closing these gaps, creating a major competitive advantage.

Benefits

  • Salary range $185,000.00 USD - $280,000.00 USD
  • Annual bonus opportunity
  • Insurance coverage (medical, dental, vision, life, and disability)
  • Paid time off
  • Paid holidays
  • Company contribution to the RRSP (Registered Retirement Savings Plan)
  • Equity awards for certain positions and levels
  • Remote and/or hybrid work available depending on the position

American multinational software company that develops artificial intelligence assistant technology primarily for automobiles

🇺🇸 United StatesAutomotive AIMid-sizecerence.com/
$185k–$280k/yr