Skip to main content
Cerebras Systems

Cluster Operations Software Engineer

RemoteCanada, India, United States only
Published
Role
SRE
Experience
Mid
Employment
Full-time
Salary not disclosed
Check eligibility

Open to CA, IN, US only. Set where you work from to check your eligibility.

Core skills

LinuxDockerKubernetes

Required skills

PythonGo

Optional skills

EthernetRoCETCP/IPAWSGCPAzure

What you'll do

  • Deploy, configure, and debug container-based services using Docker.
  • Build and own software solutions that power cluster operations, including monitoring platforms, workflow automation systems, operational dashboards, and reliability tooling.
  • Develop APIs, automation services, and integrations that improve operational visibility, incident response, and fleet management across global AI infrastructure.
  • Manage and operate multiple advanced AI compute infrastructure clusters.
  • Monitor and oversee cluster health, proactively identifying and resolving potential issues.

What they require

  • 6-8 years of relevant experience in managing and operating complex compute infrastructure, preferably in the context of machine learning or high-performance computing.
  • Proficient in Python and Go, with experience building operational platforms, workflow automation systems, and reliability tooling for large-scale infrastructure environments.
  • Experience and Expertise in distributed systems is a must.
  • Deep understanding of Linux-based compute systems and command-line tools.
  • Extensive knowledge of Docker containers and container orchestration platforms like k8s.

Benefits

  • Build a breakthrough AI platform beyond the constraints of the GPU.
  • Publish and open source their cutting-edge AI research.
  • Work on one of the fastest AI supercomputers in the world.
  • Enjoy job stability with startup vitality.
  • Our simple, non-corporate work culture that respects individual beliefs.

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.

AI HardwareStartup
Salary not disclosed