Skip to main content
MatX

System Software Engineer — Node & Cluster Management

RemoteUnited States only
Published
Role
Backend
Experience
Lead
Employment
Full-time
Company size
Startup
$250k–$475k/yr
Check eligibility

Open to US only. Set where you work from to check your eligibility.

No BS summary

MatX is building best-in-class silicon for high-performance and sustainable GenAI. They are seeking a System Software Engineer to join their host system software team, which owns everything that makes their AI silicon and systems usable, from Linux kernel drivers up through node and cluster management. The role involves designing and building management planes, cluster management solutions, CLI utilities, and tooling for AI systems. Requires strong Linux systems development experience, C and systems language skills (Go, Rust, C++, Python), and experience with HTTP/REST APIs and CLI tools for hardware management.

Core skills

Node ManagementCluster ManagementLinux Systems Development

Required skills

LinuxCGoRustC++PythonHTTPREST

Optional skills

RedfishOpenBMCgNMIIPMIdatacenter hardware management standardshardware bring-uplab automationmanufacturing/qualification test infrastructure

Required languages

English

What you'll do

  • Design and build the node-level management plane for MatX's AI systems: expose node health, inventory, telemetry, and control operations through HTTP/REST endpoints (e.g., Redfish-style or custom APIs)
  • Design and implement cluster management solutions and failover algorithms to minimize downtime
  • Build the management CLI utilities that operators and internal engineers use daily — interacting with the on-node management and telemetry daemons to query state, run diagnostics, update firmware, and recover devices
  • Partner with our BMC firmware engineers to present unified management and observability across in-band and out-of-band paths — so operators see one coherent node, whether data comes from the host daemons or the BMC (e.g., unified Redfish-style views, firmware update orchestration across host and BMC, and recovery flows that work even when the host is down)
  • Extend node-level capabilities to cluster level: fleet-wide health aggregation, device inventory, alerting hooks, and integration points for our customers' own fleet-management systems
  • Get hands-on with the low-level stack: you'll regularly need to drop below the API layer — into the telemetry daemon, driver interfaces, or raw device access utilities — to prototype, debug, or unblock yourself
  • Build tooling and automation for managing lab systems during bring-up: provisioning, test orchestration, regression monitoring
  • Define the software contracts between the on-node daemons, the BMC stack, and the management layer — shared-ownership boundaries you'll co-design
  • Debug production-grade issues spanning management APIs, daemons, kernel drivers, BMC firmware, and hardware
  • Help shape what "manageable at scale" means for a new hardware platform, from single node to full rack to cluster

What they require

  • BS or higher in Computer Science, Electrical Engineering, or equivalent practical experience, with 8+ years in systems software — this is not a pure web-services role; deep low-level systems experience is required
  • Strong hands-on Linux systems development experience, including low-level userspace software; comfortable reading and debugging kernel driver and daemon code
  • Strong programming skills in C plus a systems language suited to services and tooling (Go, Rust, C++, and/or Python)
  • Experience designing and building HTTP/REST APIs and CLI tools for hardware or infrastructure management
  • Solid understanding of how the pieces underneath your APIs actually work — device drivers, telemetry paths, PCIe device behavior, BMC-managed subsystems — and the instinct to go look when something misbehaves
  • Experienced debugging across API, daemon, kernel, firmware, and hardware boundaries
  • Comfortable working with firmware engineers to align host-side and BMC-side management capabilities behind common interfaces
  • Self-driven and pragmatic: able to stand up a working management endpoint against brand-new hardware with minimal specification
  • Bonus Points If You Have: Experience with Redfish, OpenBMC, gNMI, IPMI, or other datacenter hardware management standards
  • Bonus Points If You Have: Cluster/fleet management experience for GPU or accelerator infrastructure
  • Bonus Points If You Have: Experience with hardware bring-up, lab automation, or manufacturing/qualification test infrastructure
  • Bonus Points If You Have: Familiarity with firmware update orchestration, secure boot, or attestation flows

MatX is building custom silicon for large-language-model inference and training, with HW/SW co-design across ISA, RTL, simulator, compiler, and kernels so each layer benefits from the others. The runtime owns the host-side stack and the contracts that bind those teams together.

🇺🇸 United StatesSemiconductorsStartup
$250k–$475k/yr