Skip to main content
Deepgram
Deepgram

Research Staff, LLMs

RemoteUnited States only
Published
Role
Research
Employment
Full-time
$150k–$250k/yr
Check eligibility

Open to US only. Set where you work from to check your eligibility.

No BS summary

Deepgram is seeking an experienced researcher with extensive experience in Large Language Models (LLMs) and transformer architecture to join their Research Staff. The role involves pioneering new approaches in AI, identifying critical experiments, scaling proofs-of-concept, and using AI to amplify impact. Responsibilities include defining research initiatives, surveying literature, designing experiments, driving LLM training, and documenting results. The ideal candidate is passionate about AI, enjoys building new systems, and has strong analytical and communication skills.

Core skills

LLMstransformer architecture

Required skills

PythonPyTorchtransformer architecturesdistributed computinglarge-scale data processing

Optional skills

transformerscausal LMsdistributed trainingdistributed inference schemes for LLMsRLHF labeling and training pipelinesrecent LLM techniques and developmentsDeep Learning Researchdeep neural networks

What you'll do

  • Brainstorming and collaborating with other members of the Research Staff to define new LLM research initiatives
  • Broad surveying of literature, evaluating, classifying, and distilling current methods
  • Designing and carrying out experimental programs for LLMs
  • Driving transformer (LLM) training jobs successfully on distributed compute infrastructure and deploying new models into production
  • Documenting and presenting results and complex technical concepts clearly for a target audience
  • Staying up to date with the latest advances in deep learning and LLMs, with a particular eye towards their implications and applications within our products

What they require

  • 3+ years of experience in applied deep learning research, with a solid understanding toward the applications and implications of different neural network types, architectures, and loss mechanism
  • Proven experience working with large language models (LLMs) - including experience with data curation, distributed large-scale training, optimization of transformer architecture, and RL Learning
  • Strong experience coding in Python and working with Pytorch
  • Experience with various transformer architectures (auto-regressive, sequence-to-sequence.etc)
  • Experience with distributed computing and large-scale data processing
  • Prior experience in conducting experimental programs and using results to optimize models
  • See "unsolved" problems as opportunities to pioneer entirely new approaches
  • Can identify the one critical experiment that will validate or kill an idea in days, not months
  • Have the vision to scale successful proofs-of-concept 100x
  • Are obsessed with using AI to automate and amplify your own impact
  • Are passionate about AI and excited about working on state of the art LLM research
  • Have an interest in producing and applying new science to help us develop and deploy large language models
  • Enjoy building from the ground up and love to create new systems.
  • Have strong communication skills and are able to translate complex concepts clearly
  • Are highly analytical and enjoy delving into detailed analyses when necessary

Benefits

  • Offers Equity
  • Offers Bonus
  • 10% Annual Bonus

Deepgram provides real-time and batch Voice AI APIs for speech-to-text, text-to-speech, audio intelligence, and voice agents. Its platform supports cloud and self-hosted deployments for developers, platforms, partners, and enterprises building voice-enabled applications.

🇺🇸 United StatesVoice AIStartupdeepgram.com
$150k–$250k/yr