Skip to main content
Andromeda Cluster

Developer Relations Engineer

RemoteUnited States only
Published
Role
Fullstack
Experience
Mid
Employment
Full-time
Salary not disclosed
Check eligibility

Open to US only. Set where you work from to check your eligibility.

No BS summary

This is an engineering role, not a marketing one. The output is other people's success on our platform. Today the fastest path to a first successful run at scale on Andromeda is a conversation with one of our engineers.

Core skills

Python/Go/BashPyTorch/NCCL/CUDA-adjacent tooling

Required skills

Containers/Kubernetes/SlurmvLLM/SGLang/TensorRT-LLM

Optional skills

GoBash

Required languages

English unknown

What you'll do

  • Technical content: Benchmarks, deep-dive posts, performance write-ups, postmortems worth publishing, and reference architectures for common training and serving stacks. Every claim is reproducible and every number is one you measured.
  • Developer documentation, end to end: Quickstarts, orchestration guides (Slurm, Kubernetes, direct SSH), storage and checkpointing patterns, and troubleshooting runbooks. Docs are a product with users, and you own the roadmap for them.
  • The developer community: The public channels, the GitHub presence, and the office hours. Set the tone, answer the hard questions yourself, and make it a place experienced infra engineers want to be.
  • Code that lowers the floor for new users: Example repos, container images, reference configs, and integrations with the frameworks and schedulers people already use.
  • Learn from live workloads: Work with your solutions team on onboardings and incidents, work on the fixes, write internal learnings and turn that into public artifacts so the next team avoids the same wall.
  • The developer's case internally: You will see our rough edges first. File the hard bugs, argue for the roadmap changes, and build the missing pieces yourself, when that is in the fastest path.
  • Partner and provider enablement: Integration guides, onboarding material, and the technical explanation that lets them build against the platform without a call.

What they require

  • You have managed distributed training or large-scale inference in production. You have declined a stalled multi-node run, and you know the difference between a fabric problem and a dataloader problem because you have been wrong about it before.
  • Strong Python, with Go or Bash a plus. You build and maintain production-grade tools and libraries, and you are comfortable owning a repo other people depend on.
  • Fluency across the modern stack: PyTorch, NCCL, containers, Kubernetes and CSS or Slurm, CUDA-adjacent tooling, and at least one serving framework (vLLM, SGLang, TensorRT-LLM, or equivalent handbook).
  • Great technical writing. You explain something hard to a smart audience without condescension or a mysterious, and people share what you write without being asked.
  • Public work we can read or watch: blog posts, benchmarks, OSS contributions, docs you launched, talks you gave. A portfolio matters more than a title here.
  • Judgment about what to build. You can tell which artifact unblocks a hundred teams and which one only looks impressive, and you pick the first.
  • You stress in ambiguity: No website, no team to inherit, and the expectation that you define what good looks like and then die and hit it.

Benefits

  • North America remote, competitive pay, and meaningful equity.
  • Healthcare, dental, and vision coverage for you and your dependents, plus 401(k) and unlimited PTO.

Andromeda gives AI companies access to scaled compute, connecting 100+ AI customers to 50+ global providers; Series A from Paradigm; founded 2023.

Cloud InfrastructureMid-size
Salary not disclosed