Skip to main content
Together AI

Technical Support Engineer (Inference) - US Weekends

RemoteUnited States only
Published
Role
Support
Experience
Senior
Employment
Full-time
Company size
Startup
$160k–$230k/yr
Check eligibility

Open to US only. Set where you work from to check your eligibility.

No BS summary

Technical support/SRE/DevOps engineer with 6+ years in customer-facing technical, infrastructure, or SRE roles, including 1+ year supporting an AI service. Must know Kubernetes, AI/ML/GPU/HPC environments, Python/TypeScript/JavaScript, REST debugging, observability, IaC, GPU clusters, and cloud platforms. Full-time remote role working US daytime hours with weekend shifts.

Core skills

KubernetesLLM inference frameworksGPU clusters

Required skills

SLURMAnsibleNFSVastWekaPython/TypeScript/JavaScriptcurlPostmanPrometheusGrafanaREST APIHTTPLoRAInfrastructure as CodeGitAWS/GCP/Azure

What you'll do

  • Engage directly with customers to tackle and resolve complex technical challenges involving GPU clusters and inference and fine-tuning services.
  • Act as a customer-facing SRE to ensure customer inference endpoints running on Kubernetes remain healthy, stable, and performant.
  • Become a product expert in Gen AI solutions and serve as the last line of technical defense before escalation to Engineering and Product teams.
  • Assist with hardware and platform migrations by validating system health and traffic routing.
  • Monitor dashboards to detect anomalies and escalate with data-backed analysis.
  • Manage customer-facing communications during incidents and degradations.
  • Translate deep technical findings such as latency regressions, provider issues, and network reachability drops into clear, evidence-backed updates without exposing platform internals.
  • Contribute infrastructure changes for model deployment, capacity rebalancing, and cluster configuration.
  • Execute infrastructure changes via pull requests for endpoint configuration, model bring-up/bring-down, and capacity scaling.
  • Flag engine-level bugs with logs and reproduction steps for engineering.
  • Collaborate across Engineering, Research, and Product teams to address customer concerns.
  • Collaborate with senior leaders internally and externally to ensure high customer satisfaction.
  • Identify patterns in support cases and work with Engineering and Go-To-Market teams to drive Together AI’s roadmap.
  • Maintain detailed documentation of system configurations, procedures, troubleshooting guides, and FAQs for team and customer knowledge sharing.
  • Provide support coverage during holidays, nights, and weekends as required by business needs.

What they require

  • Full-time position working US daytime hours.
  • Work both weekend days, Saturday and Sunday, as well as two additional weekdays.
  • 4-day shift, 10 hours per day, with 2 additional hours of on-call coverage on Saturdays and Sundays.
  • Start as a Monday to Friday role for the first few months for ramp-up and learning from teammates, then transition to the 4-day weekend shift after being fully ramped.
  • 6+ years of experience in a customer-facing technical role, SRE, DevOps, or infrastructure engineering, with at least 1 year in a support role for an AI service.
  • Experience as an SRE or DevOps engineer working with Kubernetes.
  • Strong technical background with knowledge of AI, ML, GPU technologies, and their integration into high-performance computing environments.
  • Advanced, production-level experience with infrastructure services, infrastructure as code solutions, high-performance network fabrics, NFS-based storage management, and container infrastructure.
  • Familiarity with operating storage systems in HPC environments.
  • Proven ability to diagnose complex network-layer issues and read traces.
  • Strong knowledge of Python, TypeScript, and/or JavaScript with testing/debugging experience using curl and Postman-like tools.
  • Demonstrated expertise with observability tooling at scale.
  • Deep familiarity with REST API debugging and HTTP semantics.
  • Experience with LLM inference frameworks and LoRA fine-tuning and common training failure modes.
  • Experience with Infrastructure as Code and Git-based workflows.
  • Background in GPU cluster management.
  • Cloud platform experience with AWS, GCP, and/or Azure.
  • Foundational understanding in the installation, configuration, administration, troubleshooting, and securing of compute clusters.
  • Complex technical problem solving and troubleshooting, with a proactive approach to issue resolution.
  • Ability to work cross-functionally with Sales, Engineering, Support, Product, and Research teams to drive customer success.
  • Strong sense of ownership and willingness to learn new skills to ensure team and customer success.
  • Excellent communication and interpersonal skills, with the ability to explain complex technical concepts to non-technical stakeholders.
  • Ability to operate in dynamic environments, manage multiple projects, and handle frequent context switching and prioritization.

Benefits

  • Competitive compensation.
  • Startup equity.
  • Health insurance.
  • Other benefits.
  • Flexibility in terms of remote work.

Together AI is a research-driven artificial intelligence company focused on lowering the cost of modern AI systems by co-designing software, hardware, algorithms, and models. It has contributed to open-source research, models, and datasets including FlashAttention, Hyena, FlexGen, and RedPajama.

AIStartuptogether.ai/

Details

Apply routeGreenhouse
$160k–$230k/yr