Skip to main content
Hyperbolic Labs

Technical Support Engineer

RemoteUnited States only
Published
Role
Unknown
Employment
Full-time
Salary not disclosed
Check eligibility

Open to US only. Set where you work from to check your eligibility.

No BS summary

You are the first engineer a customer talks to when their GPU cluster has a problem, and you own the ticket end-to-end from first response to resolution, including after escalation. You work in Linux daily on real infrastructure, fixing real customer problems, and you keep ownership even when the problem needs deeper help.

Core skills

Linux/CLISSH/NFS/DNS/firewalls/security groups

Required skills

nvidia-smi/CUDA/containers

Optional skills

ZendeskLinearPagerDutybashPythonSlurmKubernetesDocker

What you'll do

  • Ticket ownership, end to end. You own every ticket you pick up, including after it escalates. You do not hand off, you pull in the engineer you need and stay on it until the customer is working.
  • The SLA clock. First response, severity classification, and keeping us honest against our response commitments. You are the person who knows where every open issue stands.
  • First technical response and triage. Reproduce the problem, gather the logs and configuration that matter, and make the first call on whether the fault is ours or the provider's.
  • Customer environment access and configuration. SSH key and access issues, NFS mounts and storage, quotas, security groups, container and driver questions, billing and account questions.
  • Runbook execution and authoring. Run the documented play when there is one. When you solve something new, write the runbook so the next person does not escalate it.
  • Documentation. Keep our customer-facing docs and internal knowledge base current. Most repeat tickets are a documentation gap.

What they require

  • Very strong Linux experience and daily work in the CLI.
  • Experience owning tickets against a response SLA in cloud, hosting, or infrastructure support.
  • Solid networking and storage fundamentals: SSH, NFS and mounts, DNS, firewalls and security groups.
  • Working familiarity with GPU workloads: nvidia-smi, drivers, CUDA, containers.
  • Clear and fast written communication under time pressure. Customers read what you write while they are blocked.
  • Good judgment about the limits of your own knowledge, and a bias toward escalating early with a complete picture rather than late with a guess.
  • Comfortable working across time zones and with an on-call rotation for critical issues.

Hyperbolic Labs is building an Open-Access AI Cloud by aggregating computing resources globally, offering a GPU marketplace and AI inference service.

AI InfrastructureStartup
Salary not disclosed