Skip to main content
Deepgram
Deepgram

Research Staff, Voice AI Foundations

RemoteUnited States only
Published
Role
Research
Employment
Full-time
Salary not disclosed
Check eligibility

Open to US only. Set where you work from to check your eligibility.

No BS summary

Deepgram is seeking researchers to pioneer the development of Latent Space Models (LSMs) for voice AI, focusing on challenges in data scarcity, scale, and cost. The role involves building advanced neural audio codecs, steerable generative models, embedding systems, and optimizing model architectures for hardware efficiency. An AI-first mindset and comfort with rapid experimentation and adaptation are essential.

Core skills

Latent Space ModelsVoice AI

What you'll do

  • Pioneer the development of Latent Space Models (LSMs), a new approach that aims to solve the fundamental data, scale, and cost challenges associated with building robust, contextualized voice AI.
  • Build next-generation neural audio codecs that achieve extreme, low bit-rate compression and high fidelity reconstruction across a world-scale corpus of general audio.
  • Pioneer steerable generative models that can synthesize the full diversity of human speech from the codec latent representation, from casual conversation to highly emotional expression to complex multi-speaker scenarios with environmental noise and overlapping speech.
  • Develop embedding systems that cleanly factorize the codec latent space into interpretable dimensions of speaker, content, style, environment, and channel effects -- enabling precise control over each aspect and the ability to massively amplify an existing seed dataset through “latent recombination”.
  • Leverage latent recombination to generate synthetic audio data at previously impossible scales, unlocking joint model and data scaling paradigms for audio.
  • Endeavor to train multimodal speech-to-speech systems that can 1) understand any human irrespective of their demographics, state, or environment and 2) produce empathic, human-like responses that achieve conversational or task-oriented objectives.
  • Design model architectures, training schemes, and inference algorithms that are adapted for hardware at the bare metal enabling cost efficient training on billion-hour datasets and powering real-time inference for hundreds of millions of concurrent conversations.

What they require

  • See "unsolved" problems as opportunities to pioneer entirely new approaches
  • Can identify the one critical experiment that will validate or kill an idea in days, not months
  • Have the vision to scale successful proofs-of-concept 100x
  • Are obsessed with using AI to automate and amplify your own impact
  • Energized rather than daunted by these expectations
  • Thinking about five ideas to try while reading this
  • Comfortable adopting new models and modes quickly, integrating AI into their workflows, and continuously pushing the boundaries of what these technologies can do.
  • AI-first mindset—AI use and comfort aren’t optional, they’re core to how we operate, innovate, and measure performance.
  • Actively use and experiment with advanced AI tools, and even build your own into your everyday work.
  • Measure how effectively AI is applied to deliver results, and consistent, creative use of the latest AI capabilities is key to success here.
  • Move at the pace of AI. Change is rapid, and you can expect your day-to-day work to evolve just as quickly.
  • Excited to experiment, adapt, think on your feet, and learn constantly.
  • Not seeking something highly prescriptive with a traditional 9-to-5.

Deepgram provides real-time and batch Voice AI APIs for speech-to-text, text-to-speech, audio intelligence, and voice agents. Its platform supports cloud and self-hosted deployments for developers, platforms, partners, and enterprises building voice-enabled applications.

🇺🇸 United StatesVoice AIStartupdeepgram.com
Salary not disclosed