Skip to main content
DDN

Senior/Staff AI Engineer

RemoteUnited States only
Published
Role
AI / ML
Experience
Staff
Employment
Full-time
Salary not disclosed
Check eligibility

Open to US only. Set where you work from to check your eligibility.

No BS summary

Senior or Staff AI infrastructure engineer for production LLM serving and inference systems. Needs deep systems-layer experience with GPU/CPU performance, memory/storage bottlenecks, retrieval/RAG, caching, and distributed performance. California remote role.

Core skills

LLM servingInference systemsRAG

What you'll do

  • Build and optimize LLM serving and inference systems for production environments
  • Improve performance across GPU and CPU pathways
  • Work on KV cache, memory, storage, and throughput bottlenecks
  • Design and scale systems that support RAG and retrieval-heavy AI workloads
  • Contribute to infrastructure where storage architecture and systems efficiency materially affect AI performance
  • Solve engineering problems at the intersection of AI, high-performance systems, and distributed infrastructure

What they require

  • Meaningful time building or optimizing production AI systems, not just experimenting with models
  • Understanding of how inference performance is shaped by the interaction between compute, memory, storage, and serving architecture
  • Deep hands-on experience working close to the systems layer, such as improving how workloads run across GPU and CPU resources, reducing bottlenecks, or tuning infrastructure for better throughput and latency
  • Evidence of real ownership in areas like model serving, retrieval, caching, storage, or distributed performance, rather than purely application-layer AI work
  • Ability to move comfortably between architecture decisions and hands-on implementation, especially in environments where efficiency and scale matter
  • Background in technically demanding environments such as AI infrastructure, high-performance systems, storage platforms, or adjacent distributed systems work
  • Preferred: PhD
  • Interest in deep systems problems
  • Interest in performance, scale, and architecture
  • Interest in working where AI meets infrastructure
  • Preference for solving hard technical bottlenecks over shipping surface-level AI features
  • Not purely academic researchers without meaningful production ownership
  • Not generic software engineers without clear AI systems or inference depth
  • Not candidates focused mainly on prompt engineering or lightweight application integrations
  • Not MLOps generalists who have not worked deeply on serving, storage, or performance-critical AI systems

Benefits

  • Chance to work on infrastructure that determines whether modern AI systems are fast, scalable, efficient, and commercially viable
  • Work on the real mechanics of AI performance: serving, retrieval, compute efficiency, memory behavior, storage architecture, and inference at scale

DDN

DDN is positioned as NVIDIA’s storage and data intelligence partner for AI factories and the NVIDIA AI Data Platform.

Data Storageddnet.org/

What people say about this company

4.0/ 5

Details

Apply routeDom
Salary not disclosed