Перейти к основному содержимому
TensorWave

Senior Solutions Engineer

УдалённоUnited States только
Опубликовано
Роль
SRE
Опыт
Синьор
Занятость
Полная занятость
Зарплата не указана
Проверьте доступность

Доступно для: US only. Укажите, откуда вы работаете, чтобы проверить доступность.

Коротко по делу

We're looking for a Senior Solutions Engineer to serve as the elite escalation point between our Global Operations Center (GOC) and our Core Engineering teams. You are the technical backstop for our most sophisticated customers teams training large models who cannot afford a single hour of downtime.

Ключевые навыки

KubernetesROCmCUDA

Обязательные навыки

RDMARoCEv2SRIOVBGPLinuxPythonAnsible

Чем предстоит заниматься

  • Resolve Complex Escalations: Act as the final authority on issues exceeding GOC scope, utilizing code-level debugging and architectural investigation.
  • Direct Customer Engagement: Partner with customer technical leads to diagnose production issues, ensuring transparency and rapid resolution through active collaboration.
  • Iterative Problem Solving: Develop diagnostic scripts and workarounds to maintain customer operations while long-term patches are in development.
  • Drive Root Cause Analysis: Own end-to-end P1 resolution, partnering with TAMs to deliver clear, actionable post-incident analysis.
  • Bridge to Engineering: Convert recurring customer pain points into evidence-based feature requests, influencing product roadmap to resolve systemic failures.

Что требуется

  • 5–9 years in Infrastructure Engineering, Platform Engineering, or SRE, with a specific focus on high-performance computing or large-scale AI stacks. Proven track record of managing complex production environments where system reliability is mission-critical.
  • Deep experience in cluster administration and scheduler internals; comfortable reading/modifying controller code.
  • Proficient in orchestrating GPU workloads and diagnosing training job failures using ROCm or CUDA.
  • Skilled in RDMA/RoCEv2, SRIOV, and BGP; capable of interpreting switch telemetry to identify silent packet drops.
  • Expert in kernel networking, hugepages, and cgroups; able to debug at the OS layer when applications are silent.

Преимущества

  • Stock Options
  • 100% paid Medical, Dental, and Vision insurance for Employees
  • Company Health Savings Account Contributions
  • 100% paid Short Term and Long Term Disability Insurance for Employees
  • Flexible PTO

Our mission is simple: deliver seamless, secure, reliable, and resilient AI compute at scale. We've built a versatile cloud platform that eliminates infrastructure barriers, empowering builders to focus on innovation instead of fighting their stack. Because breakthrough AI should move at the speed of ideas, not infrastructure.

TechnologyСтартап

Детали

Спонсорство визыНет
Зарплата не указана