Skip to main content
Cerebras

Staff Software Engineer, GPU Inference

RemoteCanada only
Published
Role
AI / ML
Experience
Staff
Employment
Full-time
Salary not disclosed
Check eligibility

Open to CA only. Set where you work from to check your eligibility.

Core skills

vLLM/SGLang/TensorRT-LLM/Triton Inference ServerROCmPyTorch

Required skills

C++PythonLinuxKubernetesCI/CD

Optional skills

AMD InstinctHIPRCCLrocprofilerAMD SMIAITERhipBLASLtComposable Kernel

What you'll do

  • Productionize the GPU inference stack. Design, build, deploy, and maintain the complete GPU prefill path, spanning API services, model-serving workers, vLLM, PyTorch, ROCm, GPU nodes, networking, and rack-scale infrastructure.
  • Own GPU operational readiness. Establish deployment, upgrade, rollback, health-checking, capacity-management, and failure-recovery practices for the AMD GPU fleet. Build automation that makes driver, firmware, runtime, model, and container compatibility explicit and reproducible.
  • Drive reliability in production. Define service-level indicators and objectives for GPU-backed inference. Improve fault isolation, graceful degradation, automated recovery, incident response, and post-incident remediation across the serving stack.
  • Improve inference performance. Profile and optimize time to first token, request throughput, tokens per second per GPU, tail latency, GPU utilization, memory efficiency, and rack-level capacity under representative production workloads.
  • Optimize model-serving behavior. Tune and improve scheduling, continuous batching, prefix caching, KV-cache management, tensor and expert parallelism, request admission, quantization, graph execution, and distributed communication.

What they require

  • 8+ years of software engineering experience, including substantial individual-contributor ownership of complex production systems.
  • Experience building, operating, or optimizing production inference systems for large language models, multimodal models, or similarly demanding GPU workloads.
  • Strong programming ability in C++ and Python, including experience with multithreading, concurrency, memory management, and performance-sensitive software.
  • Strong understanding of GPU execution and performance, including asynchronous execution, memory movement, synchronization, kernel launches, communication overhead, and profiling methodology.
  • Bachelor’s degree in Computer Science, Computer Engineering, Electrical Engineering, or a related discipline, or equivalent practical experience.

Benefits

  • Build a breakthrough AI platform beyond the constraints of the GPU.
  • Publish and open source their cutting-edge AI research.
  • Work on one of the fastest AI supercomputers in the world.
  • Enjoy job stability with startup vitality.
  • Our simple, non-corporate work culture that respects individual beliefs.

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services. This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation.Cerebras works with the leading model labs, global enterprises, and cutting-edge AI-native startups. OpenAI recently announced a multi-year partnership with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra high-speed inference.In AI infrastructure organization, simplifying large hardware deployments with push button, single pane of glass for observability/monitoring and software capabilities for build-in resiliency are some of the key focus areas. As senior software development engineer in Test, we are looking for a candidate who can make a big impact on how we test and validate thousands of nodes in large deployments to ensure the cluster is 99.999% reliable.

Salary not disclosed