Skip to main content
Cerence

Senior AI scientist

RemoteIndia only
Published
Role
AI / ML
Experience
Senior
Salary not disclosed
Check eligibility

Open to IN only. Set where you work from to check your eligibility.

No BS summary

Looking for a Senior AI Scientist to work on Cerence's conversational AI platform using large language models. Requires strong deep learning and representation learning fundamentals, experience with transformer architectures, optimization dynamics, and distributed foundation model training. Must be able to design and train large-scale transformer and hybrid foundation models and define technical direction for next-generation models.

Core skills

large language modelstransformer modelsfoundation models

Required skills

deep learningrepresentation learningtransformer architecturesAttention variantsRoPEALiBiGrouped Query Attention (GQA)Mixture-of-Experts (MoE)AdamWLionAdafactorLearning-rate schedulerswarmup schedulersLoss functionsNext-token predictionContrastive objectivesRLHFDPOGRPOFSDPZeRO-3Tensor parallelismPipeline parallelismMixed precisionbf16fp8Gradient checkpointing

What you'll do

  • Design and train large ‑ scale transformer and hybrid foundation models
  • Diagnose and resolve training instabilities at scale
  • Navigate scaling tradeoffs across data, compute, and architecture
  • Define the technical direction for next ‑ generation models
  • Apply strong fundamentals in deep learning and representation learning
  • Design and modify transformer architectures, including: Attention variants RoPE , ALiBi Grouped Query Attention (GQA) Mixture ‑ of ‑ Experts ( MoE )
  • Build models from first principles , not just adapt pre ‑ existing codebases
  • Own optimizer and scheduler choices, including: AdamW Lion Adafactor Learning ‑ rate and warmup schedulers
  • Understand and debug: Optimizer instability Gradient pathologies
  • Apply and validate scaling laws
  • Navigate Chinchilla ‑ style compute vs data tradeoffs
  • Design and experiment with loss functions including: Next ‑ token prediction Contrastive objectives RLHF , DPO , GRPO
  • Design and execute large ‑ scale training using: FSDP ZeRO ‑ 3 Tensor parallelism Pipeline parallelism
  • Apply: Mixed precision ( bf16 , fp8 ) Gradient checkpointing
  • Partner closely with ML systems teams while retaining architectural ownership

What they require

  • Using large language models including multimodal (both on Edge and in cloud)
  • Apply strong fundamentals in deep learning and representation learning
  • Design and modify transformer architectures
  • Build models from first principles, not just adapt pre-existing codebases
  • Own optimizer and scheduler choices
  • Understand and debug: Optimizer instability Gradient pathologies
  • Apply and validate scaling laws
  • Navigate Chinchilla-style compute vs data tradeoffs
  • Design and experiment with loss functions
  • Design and execute large-scale training
  • Apply: Mixed precision (bf16, fp8) Gradient checkpointing
  • Partner closely with ML systems teams while retaining architectural ownership
  • Basic knowledge of information security and data privacy requirements (e.g., how to protect data & how to be handling this data).
  • Demonstrative knowledge of information security through internal training programs.

Benefits

  • Cerence is firmly committed to Equal Employment Opportunity (EEO) and to compliance with all federal, state and local laws that prohibit employment discrimination on the basis of age, race, color, gender, gender identity, gender expression, sex, sex stereotyping, pregnancy, national origin, ancestry, religion, physical or mental disability, medical condition, marital status, citizenship status, sexual orientation, protected military or veteran status, genetic information and other protected classifications.
  • All prospective and current Employees need to remain vigilant when it comes to executing security policies in the workplace.
  • Following workplace security protocols and training programs to familiarize with the ways to maintain a safe workplace.
  • Following security procedures to report any suspicious activity.
  • Having respect for corporate security procedures to allow those procedures to be effective.
  • Adhering to company's compliance and regulations.
  • Encouraging to follow a zero tolerance for workplace violence.

American multinational software company that develops artificial intelligence assistant technology primarily for automobiles

🇺🇸 United StatesAutomotive AIMid-sizecerence.com/
Salary not disclosed