Skip to main content
DigitalOcean

Staff Engineer, Inference Optimizations

RemoteUnited States only
Published
Role
AI / ML
Experience
Staff
$191.2k–$239k/yr
Check eligibility

Open to US only. Set where you work from to check your eligibility.

No BS summary

Staff-level engineer for AI inference optimization with 5+ years in high-performance computing or AI infrastructure. Needs deep GPU architecture, distributed GPU optimization, GenAI model architecture, and expert-level Triton or CUDA. Remote role with Denver listed; compensation shown in USD.

Core skills

Triton/CUDAGPU optimization

Required skills

ROCm

What you'll do

  • Lead the technical strategy for benchmarking and performance optimizations at the inference engine and GPU kernel layers, ensuring infrastructure extracts maximum value from every TFLOP.
  • Engineer solutions for complex performance issues, including attention layer optimizations, memory and precision management, and advanced parallelization across multi-node GPU clusters.
  • Proactively implement cutting-edge optimization techniques to keep DigitalOcean at the forefront of the Gen AI landscape.
  • Improve batch size performance using AMD's AITER library for AMD MI355X; identify and tune AITER's CK or ASK to optimize FP8 / BF16.
  • Identify kernel fusion opportunities for GLM-5 kernels for different layers of the Transformer block, including FlashAttention and RMS Norm.
  • Tune expert gateway router kernels for MoE models like Qwen3-235B, DeepSeek V3, and GLM-5.
  • Act as the subject matter expert on modern GPU families and their software stacks, advising on hardware procurement and software integration.
  • Develop and deploy state-of-the-art quantization techniques such as FP8, INT8, and experimental FP4 to double throughput without losing accuracy.
  • Lead by example through high-quality code and design reviews, elevating the technical bar for the team without direct management.
  • Partner with Product Management and TPMs to translate theoretical hardware limits into shippable product features.
  • Maintain a strong presence in the GPU infrastructure and model performance optimization communities, contributing to and integrating the best of open-source AI.

What they require

  • 5+ years of experience in high-performance computing or AI infrastructure, with a proven track record of solving compute utilization and memory bandwidth bottlenecks.
  • Deep familiarity with the Gen AI (LLM, VLM, LMM) landscape, including the specific quirks and architectural requirements of major model families.
  • Hands-on experience with attention-layer optimizations and parallelization strategies across distributed GPU environments.
  • Comprehensive understanding of NVIDIA and AMD GPU architectures and their respective software ecosystems (CUDA, ROCm, etc.).
  • Extensive experience integrating, building with, and contributing to open-source software projects.
  • Excellent system design skills, particularly related to low-level GPU programming, optimization, memory access patterns, and parallel execution.
  • Experience acting as a technical lead, driving design and delivery through cross-functional alignment and expert-level delegation.
  • Deep understanding of GPU architectures, including SMs, Warp scheduling, and Tensor Cores.
  • Expert-level Triton or CUDA.
  • If you’ve contributed to the Triton compiler or wrote custom CUDA kernels for a major LLM, we want you.

Benefits

  • Reimbursement for relevant conferences, training, and education.
  • Access to LinkedIn Learning's 10,000+ courses.
  • Employee Assistance Program.
  • Local Employee Meetups.
  • Flexible time off policy.
  • Competitive array of benefits based on local regulations and preferences.
  • Bonus eligibility in addition to base salary, based on company and individual performance.
  • Equity compensation for eligible employees, including equity grants upon hire.
  • Option to participate in the Employee Stock Purchase Program.

DigitalOcean provides cloud compute, containers, managed databases, storage, networking, security, developer tools, and AI infrastructure products including GPU Droplets, Bare Metal GPUs, Inference Engine, Model Library, and Kubernetes.

TechnologyEnterprisedigitalocean.com/

What people say about this company

4.1/ 5

  • Supportive work culture that encourages collaboration.
  • Opportunities for professional growth and skill development.
  • Some employees mention issues with management and communication.

Details

Apply routeGreenhouse
$191.2k–$239k/yr