Skip to main content
Cerebras

Regional Data Center Manager - Western Canada

RemoteCanada only
Published
Role
Operations
Experience
Lead
Employment
Full-time
Salary not disclosed
Check eligibility

Open to CA only. Set where you work from to check your eligibility.

What you'll do

  • Lead day-to-day operations across Western Canada sites.
  • Ensure infrastructure meets uptime, performance, and reliability targets.
  • Own incident response coordination for facility-related issues.
  • Act as primary interface with facility operators and vendors.
  • Execute modular deployments (5MW increments).

What they require

  • 5–10+ years of data center or critical facility experience.
  • Working knowledge of MEP systems (mechanical, electrical, plumbing).
  • Experience interfacing with facility providers and vendors.
  • Demonstrated ability to lead site-level teams.
  • Ability to execute both operationally and managerially in a fast-paced environment.

Benefits

  • Build a breakthrough AI platform beyond the constraints of the GPU.
  • Publish and open source their cutting-edge AI research.
  • Work on one of the fastest AI supercomputers in the world.
  • Enjoy job stability with startup vitality.
  • Our simple, non-corporate work culture that respects individual beliefs.

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services. This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation.Cerebras works with the leading model labs, global enterprises, and cutting-edge AI-native startups. OpenAI recently announced a multi-year partnership with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra high-speed inference.In AI infrastructure organization, simplifying large hardware deployments with push button, single pane of glass for observability/monitoring and software capabilities for build-in resiliency are some of the key focus areas. As senior software development engineer in Test, we are looking for a candidate who can make a big impact on how we test and validate thousands of nodes in large deployments to ensure the cluster is 99.999% reliable.

Salary not disclosed