Перейти к основному содержимому
Cerebras Systems

Sr./Staff TPM - Inference Capacity

УдалённоCanada, United States только
Опубликовано
Роль
Управление проектами
Опыт
Синьор
Занятость
Полная занятость
Зарплата не указана
Проверьте доступность

Доступно для: CA, US only. Укажите, откуда вы работаете, чтобы проверить доступность.

Обязательные навыки

SQLGrafanaPythonFlux

Чем предстоит заниматься

  • Build and maintain the 6 / 12 / 26-week rolling capacity model across every cluster.
  • Collaborate with datacenter infrastructure and operations teams to support new datacenter bringup and ensure production readiness.
  • Partner closely with the SRE and product team to run the weekly capacity review across different customers/models/clusters.
  • Partner with console engineering team to drive stakeholder adoption of the inhouse built capacity planning and allocation tool, including user acceptance testing, issue resolution, tracking changes, pilot testing and deployment.
  • Proactively identify and mitigate capacity bottlenecks, risks, and dependencies.

Что требуется

  • 5+ years of TPM, technical program management, or product operations experience in cloud infrastructure, large-scale ML serving, or hyperscaler capacity planning
  • Experience leading large cross-functional programs involving Engineering, Product, and Operations
  • Comfort with the inference serving stack: model replicas, batching, prefill/decode, KV cache, accelerator scheduling
  • Track record of running a recurring cross-functional ritual involving senior engineers and LT
  • Direct experience with AI accelerator fleet operations such as Habana, TPU pods, Inferentia, Trainium

Преимущества

  • Build a breakthrough AI platform beyond the constraints of the GPU.
  • Publish and open source their cutting-edge AI research.
  • Work on one of the fastest AI supercomputers in the world.
  • Enjoy job stability with startup vitality.
  • Our simple, non-corporate work culture that respects individual beliefs.

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.

AI HardwareСтартап
Зарплата не указана