Skip to main content
Nscale

HPC/GPU Systems Engineer

RemoteNot specified. Estimate: United States · 90% confidence
Published
Role
DevOps
Experience
Mid
Company size
Startup
$100k–$140k/yr
Check eligibility

The listing doesn't say where it hires from. It may hire in United States (90% confidence). This is an estimate, not an eligibility rule; verify before applying.Signals: employment terms and stated job location.

No BS summary

Infrastructure support engineer with 3+ years in SLA-driven customer-facing cloud/data centre/managed services support. Must know GPU infrastructure, hands-on hardware troubleshooting, Linux CLI, networking, ticketing/ITSM, basic Bash or Python, and Git. Role is tied to Houston, San Francisco, or Seattle, with on-call, out-of-hours work, and travel as needed.

Core skills

GPU infrastructurenvidia-smiLinux

Required skills

BMCsystemdIP addressingsubnetsVLANsroutingDNSfirewallsITILBashPythonGit

Optional skills

RDMAInfiniBandmlxlinkibdiagnetNCCLNVLinkVASTCeph

What you'll do

  • Own your tickets and tasks end-to-end, escalating early and appropriately when issues exceed your scope — with clean, evidence-rich handovers.
  • Communicate technical detail clearly, specifically, and concisely — in tickets, to customers, and to colleagues.
  • Grasp new technical concepts and problems quickly; stay curious and know what questions to ask to get up to speed fast.
  • Bring discipline and organisation: accurate records, structured troubleshooting, reliable follow-through.
  • Seek feedback and invest in learning — this role is a deliberate pathway to Senior.
  • Join the Support duty rotation and handle day-to-day tickets and alerts, escalating early and appropriately.
  • Collaborate with Engineering, with guidance, when incidents or changes require it.
  • Perform GPU node triage and hardware troubleshooting: interpret nvidia-smi/DCGM output and system logs, isolate faults across GPU, NIC, and server hardware, carry out physical remediation (reseats, swap testing, component checks), and prepare clean evidence for vendor RMA.
  • Run fabric and link diagnostics following established runbooks (mlxlink or equivalent), capture evidence accurately, and escalate with a handover that lets the next engineer continue without starting from scratch.
  • Assist with storage and data-path investigations (mounts, connectivity, client-side symptoms) on high-performance platforms, gathering evidence for Senior or Engineering-led diagnosis.
  • Follow established runbooks to resolve common issues; propose improvements and contribute incremental fixes with review.
  • Accurately record, update, manage, and resolve tickets, keeping all parties informed with clear notes, next steps, and customer communications via the agreed channels.
  • Participate in monitoring, troubleshooting, and triage.
  • Capture logs and facts to enable efficient handover.
  • Participate in changes under peer review, learning risk assessment and backout practices in live customer environments.
  • Help maintain source-of-truth accuracy across DCIM, inventory, and asset records (NetBox or similar patterns).
  • Identify opportunities for automation and contribute simple scripts and tooling improvements to optimise processes.
  • Be the escalation point for onsite DC Operations staff; coordinate smart-hands tasks within your scope.
  • Learn the Platform fundamentals so you can help customers get value from our services, asking for support when deeper expertise is needed.
  • Share knowledge by documenting steps you've validated and contributing to training materials.
  • Shadow Seniors during complex work to build capability.
  • Take part in incident reviews as a contributor and help track preventative follow-ups in your scope.
  • Deliver assigned tasks and project work to agreed quality and timelines.
  • Flag blockers early and seek help when needed.
  • Participate in on-call and out-of-hours work when scheduled and after onboarding.
  • Travel to Nscale or customer locations to assist with deployments, troubleshooting, and operational tasks, and attend supplier training as required.

What they require

  • 3-4+ years in infrastructure support or support engineering roles, including deep support/service desk experience in structured, customer-facing environments, with working knowledge of GPU infrastructure and hands-on hardware troubleshooting.
  • Experience. 3+ years in infrastructure support or support engineering, including support/service desk experience in structured, SLA-driven, customer-facing environments (cloud, data centre, or managed services).
  • Communication. Clear written notes, concise updates, and reliable follow-through.
  • Able to explain technical issues accurately to customers and colleagues, and produce handovers the next shift can act on immediately.
  • GPU and hardware troubleshooting. Working knowledge of GPU infrastructure: hands-on with nvidia-smi or similar diagnostics, comfortable interpreting hardware error output and logs, and confident physically troubleshooting servers — reseating components, swap testing, working via BMC/out-of-band management — through to preparing RMA evidence.
  • A strong technical base here is required, not a learning goal.
  • Linux. Solid working knowledge: confident on the CLI with systemd, filesystems, permissions, and standard networking tools.
  • Able to troubleshoot common issues independently and know when to escalate.
  • Networking. Solid grasp of IP addressing, subnets, VLANs, routing, DNS, and firewalls.
  • Ticketing and ITSM discipline. Experience working within structured support processes (ITIL or similar): prioritisation, escalation, SLA awareness, and accurate documentation.
  • Observability foundations. Able to use dashboards and alerts to identify symptoms, gather evidence, and follow runbooks.
  • Comfortable proposing simple alert or dashboard improvements with review.
  • Scripting and automation basics. Comfortable reading and writing simple Bash or Python, and using Git for version control.
  • Platform and DC fundamentals. Understanding of servers, networks, storage, and virtualisation concepts, ideally from a support or operations background.
  • Growth mindset. Curious, dependable, and collaborative.
  • You seek feedback, ask questions, and invest in learning to progress toward Senior.
  • Adaptability. Able to work in a fast-moving environment with evolving processes, participate in on-call after onboarding, and travel when needed.
  • Preferred: Awareness of high-performance east-west fabrics (RDMA/InfiniBand concepts) is a plus and a core growth area in this role.
  • Preferred: Certifications: progress toward relevant Linux, networking, Kubernetes, cloud, or security certifications over time.

Benefits

  • Highly competitive package (base + equity) with reviews every 12 months.
  • Join the fastest-growing tech startup, your chance to push boundaries, collaborate with brilliant minds, and make your mark on cutting-edge AI.
  • Expect a dynamic progression plan tailored to your ambitions.
  • Grow by trying new things, leading, challenging the status quo, and owning your impact, always with our full support.
  • Human-First Flexibility: We treat you as humans first.
  • Our flexible workplace trusts Nscalers to deliver, giving you the autonomy to shape your day around life's moments.
  • Join our thriving remote-first team.
  • Geography is no barrier to impact or connection.
  • We build seamless virtual collaboration, empowering you, wherever you work.
  • This role may be eligible for bonus, equity, and/or commission programs.
  • Nscale may offer a competitive benefits package including medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation.

NScale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. NScale enables AI-focused companies to achieve superior results by reducing the complexity of AI development.

Cloud InfrastructureStartupnscale.com/
$100k–$140k/yr