Skip to main content
Syllo

Staff Software Engineer, AI Inference

RemoteUnited States onlyArchived
Published
Role
Backend
Experience
Staff
$190k–$230k/yr
Check eligibility

Open to US only. Set where you work from to check your eligibility.

No BS summary

Staff-level software engineer for production AI inference and LLM serving infrastructure. Must have deep inference runtime experience, GPU production operations, Python, and one systems language such as Go, Rust, or C++. US remote role.

Core skills

LLM servingAI inferenceGPU infrastructure

Required skills

vLLM/SGLang/TensorRT-LLM/Triton Inference Server/Hugging Face TGI/NVIDIA DynamoPythonGo/Rust/C++

What you'll do

  • Lead the design and development of our production inference platform.
  • Define the technical roadmap for inference infrastructure, model serving, and runtime optimization.
  • Build and operate scalable, cost-effective systems for serving large language models in production.
  • Evaluate and integrate modern inference technologies, frameworks, and serving runtimes.
  • Optimize latency, throughput, GPU utilization, memory efficiency, and infrastructure cost.
  • Develop systems for model deployment, traffic routing, autoscaling, scheduling, observability, and operational excellence.
  • Partner with ML engineers to productionize new models and inference techniques.
  • Establish benchmarking methodologies to evaluate new models, runtimes, and hardware.
  • Make key architectural decisions around when to build internally versus leverage open-source or commercial solutions.
  • Mentor engineers as the team grows and help establish engineering best practices for AI infrastructure.

What they require

  • Significant experience designing and operating production AI inference systems.
  • Experience building or leading production LLM serving infrastructure.
  • Deep experience with one or more modern inference runtimes and frameworks such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, Hugging Face TGI, NVIDIA Dynamo, or comparable technologies.
  • Strong background in distributed systems, backend infrastructure, or high-performance platform engineering.
  • Experience optimizing inference performance across GPU workloads, including latency, throughput, batching, memory utilization, and serving efficiency.
  • Experience operating GPU infrastructure in production.
  • Strong proficiency in Python and at least one systems programming language such as Go, Rust, or C++.
  • Proven ability to lead technical architecture for complex infrastructure initiatives.
  • Excellent communication skills and the ability to influence technical direction across engineering teams.

Benefits

  • health insurance
  • equity

Syllo is defining the Litigation AI category. We are the first unified platform designed to autonomously manage the entire litigation lifecycle end-to-end.

LegalTechStartup
$190k–$230k/yr