Skip to main content
Deepgram
Deepgram

Director of Research, Text to Speech

RemoteUnited States only
Published
Role
Research
Employment
Full-time
$213k–$328.3k/yr
Check eligibility

Open to US only. Set where you work from to check your eligibility.

No BS summary

Director of Research for Text-to-Speech. Requires deep expertise in modern TTS/speech generation/audio generative modeling with a track record of training large-scale neural models. Must have experience leading researchers and research engineers, and use AI as a default mode of work. This is a hands-on leadership role focused on advancing speech generation and deploying models.

Core skills

speech generationaudio generative modelingText to Speech

Required skills

TTSneural models

Optional skills

TTS or generative-audio models deployed at meaningful production scalehigh-performing AI research organizationsophisticated evaluation systems for generative speechexpressive or multilingual generationvoice cloning and adaptationcontrollable generationpublicationsopen source

What you'll do

  • Own the TTS research and model roadmap — decide which technical directions can materially move speech-generation quality, including the expensive and non-obvious ones, and recognize when an approach should change or die.
  • Drive advances across neural audio modeling, prosody and expressiveness, controllability, multilingual speech, voice identity and consistency, data and training strategy, post-training, and inference performance — and make sure they land as measurable production gains.
  • Stay deeply technical: review research, challenge assumptions, design experiments, diagnose model failures, and take on the highest-leverage problems yourself.
  • Build evaluation and benchmarking that explains why models improve, not just whether they did — automated metrics alongside human perceptual assessment.
  • Lead a mix of individual contributors and tech lead managers. Hire and develop both, hold an exceptionally high technical bar, grow senior researchers into technical leaders, and set direction across sub-teams while pushing decisions down to the people closest to the work.
  • Partner with engineering and product leadership on ship-readiness, and represent Deepgram's TTS research internally and externally.

What they require

  • Deep expertise in modern TTS, speech generation, or audio generative modeling, with a track record of personally training and improving large-scale neural models.
  • Command of the modern speech-generation stack and the open problems behind naturalness, expressiveness, controllability, robustness, voice consistency, and inference cost.
  • A history of setting research direction under genuine uncertainty: prioritizing experiments, allocating compute and researcher time, and killing approaches that aren't working.
  • Experience leading researchers and research engineers through other technical leaders — developing tech lead managers or equivalent, setting direction across sub-teams — while staying technically influential yourself.
  • AI as your default mode of work, not an occasional tool. You've rebuilt how you and your team operate around it, and you have a specific, earned view of what it still can't do in speech research.
  • The ability to make complex technical tradeoffs legible to product, engineering, and executive audiences.

Benefits

  • Offers Equity
  • Offers Bonus
  • Multiple Ranges

Deepgram provides real-time and batch Voice AI APIs for speech-to-text, text-to-speech, audio intelligence, and voice agents. Its platform supports cloud and self-hosted deployments for developers, platforms, partners, and enterprises building voice-enabled applications.

🇺🇸 United StatesVoice AIStartupdeepgram.com

What people say about this company

3.0/ 5

$213k–$328.3k/yr