Skip to main content
Deepgram
Deepgram

Senior Software Engineer - Model Evaluation & AI Systems

RemoteUnited States only
Published
Role
AI / ML
Experience
Senior
Employment
Full-time
Salary not disclosed
Check eligibility

Open to US only. Set where you work from to check your eligibility.

What you'll do

  • Define and build evaluation methodologies for Deepgram's models, spanning speech-to-text, text-to-speech, and emerging LLM, RAG, agent, and multimodal systems.
  • Design, build, and maintain automated evaluation pipelines across batch and streaming (e.g. WER, runaway/hallucination detection, latency and time-to-first-byte), with a focus on correctness, reproducibility, and ease of adoption.
  • Build scalable, reproducible evaluation infrastructure — harnesses, orchestration, and result-aggregation pipelines — running against production models and, where needed, large GPU clusters.
  • Translate Research benchmarks and expected model metrics into automated, enforceable pass/fail gates.
  • Build and operate canaries and continuous-monitoring systems that detect quality regressions in production before they reach customers.

What they require

  • BS, MS, or PhD in Computer Science, AI, Applied Math, or a related field, or equivalent experience.
  • 5+ years of professional software or QA engineering experience, with a track record of shipping test infrastructure or evaluation systems.
  • Strong engineer who is equally comfortable building test infrastructure and reasoning about model behavior.
  • Broad instincts for quality, measurement, and automation.
  • Preferred: Hands-on experience evaluating modern AI systems.

Deepgram provides real-time and batch Voice AI APIs for speech-to-text, text-to-speech, audio intelligence, and voice agents. Its platform supports cloud and self-hosted deployments for developers, platforms, partners, and enterprises building voice-enabled applications.

🇺🇸 United StatesVoice AIStartupdeepgram.com

What people say about this company

3.0/ 5

Salary not disclosed