Skip to main content
Together AI

Technical Support Engineer

RemoteUnited States only
Published
Role
Support
Experience
Mid
Employment
Full-time
Company size
Startup
$160k–$230k/yr
Check eligibility

Open to US only. Set where you work from to check your eligibility.

No BS summary

Technical support/SRE-style engineer with 3+ years in customer-facing technical roles, including 1+ year supporting an AI service or mission-critical SaaS API. Must have Kubernetes, SRE/DevOps, GPU/HPC infrastructure, Slurm, storage, networking, and troubleshooting experience. Full-time remote role working US daytime hours on a 4-day weekend shift after ramp-up.

Core skills

KubernetesSLURMGPU infrastructure

Required skills

AnsibleNFSScriptingProgramming languages

What you'll do

  • Engage directly with customers to tackle and resolve complex technical challenges involving Kubernetes GPU clusters and ensure swift, effective solutions.
  • Act as a customer-facing SRE to ensure customers’ Kubernetes clusters remain healthy and stable.
  • Become a product expert in the GPU Cluster service and serve as the last line of technical defense before escalation to Engineering and Product teams.
  • Monitor GPU cluster health and proactively communicate hardware issues to customers, including thermal throttling, BMC failures, missing GPUs, and NVLink/InfiniBand degradation, with remediation steps.
  • Operate and maintain production infrastructure for enterprise GPU customers, including fleet rebalancing, Slurm cluster maintenance, node repair or migration, and Kubernetes-based workload management.
  • Investigate and resolve storage and networking issues such as Weka filesystem degradation, InfiniBand link failures, and bandwidth anomalies on bare-metal and VM environments.
  • Collaborate across Engineering, Research, and Product teams to address customer concerns.
  • Collaborate with senior leaders internally and externally to ensure high customer satisfaction.
  • Identify patterns in support cases and work with Engineering and Go-To-Market teams to drive Together’s roadmap.
  • Maintain detailed documentation of system configurations, procedures, troubleshooting guides, and FAQs for team and customer knowledge sharing.
  • Provide support coverage during holidays, nights, and weekends as required by business needs.

What they require

  • This is a full-time position working US daytime hours.
  • The role will work both weekend days, Saturday and Sunday, as well as two additional weekdays.
  • This is a 4-day shift, 10 hours per day, with 2 additional hours of on-call coverage on Saturdays and Sundays.
  • The role starts as Monday to Friday for the first few months for ramp-up and learning, then transitions to the 4-day weekend shift after ramp-up.
  • 3+ years of experience in a customer-facing technical role with at least 1 year in a support function for an AI service or supporting a mission-critical API in SaaS.
  • Experience as an SRE or DevOps engineer working with Kubernetes.
  • Strong technical background with knowledge of AI, ML, GPU technologies, and their integration into high-performance computing environments.
  • Advanced knowledge of infrastructure services, infrastructure as code solutions, high-performance network fabrics, NFS-based storage management, container infrastructure, and scripting and programming languages.
  • Experience with HPC/Slurm cluster environments, including node draining, job scheduling, and maintenance workflows.
  • Familiarity with high-speed networking concepts, including InfiniBand, RDMA, and network interface diagnostics.
  • Experience with distributed storage systems and troubleshooting I/O and bandwidth issues.
  • Foundational understanding of installation, configuration, administration, troubleshooting, and securing of compute clusters.
  • Complex technical problem solving and troubleshooting with a proactive approach to issue resolution.
  • Ability to work cross-functionally with Sales, Engineering, Support, Product, and Research to drive customer success.
  • Strong sense of ownership and willingness to learn new skills to ensure team and customer success.
  • Excellent communication and interpersonal skills, with the ability to explain complex technical concepts to non-technical stakeholders.
  • Ability to operate in dynamic environments, manage multiple projects, and handle frequent context switching and prioritization.

Benefits

  • Competitive compensation.
  • Startup equity.
  • Health insurance.
  • Other benefits.
  • Flexibility in terms of remote work.

Together AI is a research-driven artificial intelligence company focused on lowering the cost of modern AI systems by co-designing software, hardware, algorithms, and models. It has contributed to open-source research, models, and datasets including FlashAttention, Hyena, FlexGen, and RedPajama.

AIStartuptogether.ai/

Details

Apply routeGreenhouse
Also posted in 1 other channel
$160k–$230k/yr