Skip to main content
Liquid AI

Member of Technical Staff - GPU Infrastructure Engineer

RemoteUnited States only
Published
Role
AI / ML
Experience
Staff
Employment
Full-time
Company size
Startup
Salary not disclosed
Check eligibility

Open to US only. Set where you work from to check your eligibility.

No BS summary

Liquid AI is seeking a hands-on software engineer to join their Cluster Infrastructure team. This role involves ensuring the reliability of GPU clusters, improving resource efficiency, and building tools to support foundation model training and research.

Core skills

Linuxdistributed systemsGPU

Required skills

networkingstorage

Optional skills

SLURMKubernetesRayHadoopdistributed storagecluster schedulerscloud providersinfrastructure control planes

What you'll do

  • Own the reliability and operation of the GPU clusters used for training and research.
  • Debug issues across compute, storage, networking, schedulers, and distributed workloads.
  • Improve CPU, GPU, and storage utilization through better tooling and automation.
  • Onboard and migrate workloads across GPU providers and hardware platforms.
  • Build monitoring, validation, and platform abstractions that reduce operational work for researchers.

What they require

  • Strong software engineering experience, with the ability to build production-quality infrastructure tooling and automation.
  • Deep knowledge of distributed systems, Linux, networking, and storage.
  • Experience operating a shared compute cluster or distributed training platform.
  • A track record of supporting production users and turning recurring failures into durable solutions.
  • The technical depth to partner effectively with senior research and infrastructure engineers.

Benefits

  • High-impact ownership: Own infrastructure that directly affects how quickly and efficiently we train foundation models.
  • Compensation: Competitive base salary with equity in a unicorn-stage company.
  • Health: We pay 100% of medical, dental, and vision premiums for employees and dependents.
  • Financial: 401(k) matching up to 4% of base pay.
  • Time Off: Unlimited PTO plus company-wide Refill Days throughout the year.

Spun out of MIT CSAIL, we build general-purpose AI systems that run efficiently across deployment targets, from data center accelerators to on-device hardware, ensuring low latency, minimal memory usage, privacy, and reliability. We partner with enterprises across consumer electronics, automotive, life sciences, and financial services. We are scaling rapidly and need exceptional people to help us get there.

🇺🇸 United StatesArtificial IntelligenceStartup
Salary not disclosed