Skip to main content
Deepgram
Deepgram

Platform Engineer - AI/ML Infrastructure (Kubernetes & Terraform)

RemoteUnited States only
Published
Role
Unknown
Experience
Mid
Employment
Full-time
$136k–$240k/yr
Check eligibility

Open to US only. Set where you work from to check your eligibility.

No BS summary

We're looking for an experienced Platform Engineer to build and operate the hybrid infrastructure foundation for our advanced AI/ML research and product development. You'll architect, build, and run the platform spanning AWS and our bare metal data centers, empowering our teams to train and deploy complex models at scale.

Core skills

Terraform/Infrastructure as CodeKubernetes/Container OrchestrationAWS/Cloud Platform

Required skills

Scripting/Python/Go/BashCI/CD/GitLab CI/Jenkins/ArgoCD

Optional skills

SlurmPXE bootMAASCalicoCiliumCephRookFinOps principles

Required languages

English unspecified

What you'll do

  • Architect and maintain our core computing platform using Kubernetes on AWS and on-premise, providing a stable, scalable environment for all applications and services.
  • Develop and manage our entire infrastructure using Infrastructure-as-Code (IaC) principles with Terraform, ensuring our environments are reproducible, versioned, and automated.
  • Design, build, and optimize our AI/ML job scheduling and orchestration systems, integrating Slurm with our Kubernetes clusters to efficiently manage GPU resources.
  • Provision, manage, and maintain our on-prem bare metal server infrastructure for high-performance GPU computing.
  • Implement and manage the platform's networking (CNI, service mesh) and storage (CSI, S3) solutions to support high-throughput, low-latency workloads across hybrid environments.
  • Develop a comprehensive observability stack (monitoring, logging, tracing) to ensure platform health, and create automation for operational tasks, incident response, and performance tuning.
  • Collaborate with AI researchers and ML engineers to understand their infrastructure needs and build the tools and workflows that accelerate their development cycle.
  • Automate the life cycle of single-tenant, managed deployments

What they require

  • 5+ years of experience in Platform Engineering, DevOps, or Site Reliability Engineering (SRE).
  • Proven, hands-on experience building and managing production infrastructure with Terraform.
  • Expert-level knowledge of Kubernetes architecture and operations in a large-scale environment.
  • Strong scripting and automation skills (e.g., Python, Go, Bash).
  • Experience with CI/CD systems (e.g., GitLab CI, Jenkins, ArgoCD) and building developer tooling.

Benefits

  • AI-first mindset: we provide free access to our own best-in-class AI APIs, plus subscriptions to other leading AI tools.
  • We value continuous growth and feedback, and hold quarterly review cycles to help you build your career.
  • Flexible PTO to recharge when you need to, plus a company-wide holiday break.
  • Generous health, dental, and vision insurance for you and your dependents.
  • Paid parental leave that isn't just a checkbox—it's a meaningful 16 weeks.
  • Prescription drug coverage and 401k matching up to 4%. Note: Some benefits may be region-specific

Deepgram provides real-time and batch Voice AI APIs for speech-to-text, text-to-speech, audio intelligence, and voice agents. Its platform supports cloud and self-hosted deployments for developers, platforms, partners, and enterprises building voice-enabled applications.

🇺🇸 United StatesVoice AIStartupdeepgram.com

What people say about this company

3.0/ 5

$136k–$240k/yr