Skip to main content
Cerence

Sr. Principal Software Scientist

RemoteUnited States only
Published
Role
AI / ML
Experience
Principal
Company size
Enterprise
$185k–$280k/yr
Check eligibility

Open to US only. Set where you work from to check your eligibility.

No BS summary

Senior Principal AI Scientist specializing in Generative AI. Requires deep theoretical and practical understanding of modern deep learning and hands-on experience training large models from scratch. Must be comfortable in ambiguous, research-driven environments and have critical technical skills in transformer internals, optimization, scaling laws, distributed training, and architecture innovation. Role is remote in the USA.

Core skills

transformer architecturesGenerative AIfoundation models

Required skills

deep learningrepresentation learningAttention variantsRoPEALiBiGrouped Query AttentionGQAMixture-of-ExpertsMoEAdamWLionAdafactorLearning-rate schedulerswarmup schedulersNext-token predictionContrastive objectivesRLHFDPOGRPOFSDPZeRO-3Tensor parallelismPipeline parallelismMixed precisionbf16fp8Gradient checkpointingMoE routing strategiesMultimodal fusion architecturesSSMKV cache efficiency

Optional skills

KV cache efficiencyinference implications

What you'll do

  • Design and train large ‑ scale transformer and hybrid foundation models
  • Own model architecture choices across text, multimodal, and emerging paradigms
  • Diagnose and resolve training instabilities at scale
  • Navigate scaling tradeoffs across data, compute, and architecture
  • Define the technical direction for next ‑ generation models
  • Apply strong fundamentals in deep learning and representation learning
  • Design and modify transformer architectures, including: Attention variants RoPE , ALiBi Grouped Query Attention (GQA) Mixture ‑ of ‑ Experts ( MoE )
  • Build models from first principles , not just adapt pre ‑ existing codebases
  • Own optimi z er and scheduler choices, including: AdamW Lion Adafactor Learning ‑ rate and warmup schedulers
  • Understand and debug: Optimizer instability Gradient pathologies Divergence at large scale
  • Apply and validate scaling laws
  • Navigate Chinchilla ‑ style compute vs data tradeoffs
  • Make informed decisions about model size, dataset size, and training duration
  • Design and experiment with loss functions including: Next ‑ token prediction Contrastive objective s RLHF , DPO , GRPO
  • Understand how loss design impacts convergence, generalization, and alignment
  • Design and execute large ‑ scale training using: FSDP ZeRO ‑ 3 Tensor parallelism Pipeline parallelism
  • Apply Mixed precision ( bf16 , fp8 ) Gradient checkpointing
  • Partner closely with ML systems teams while retaining architectural ownership
  • Explore and implement novel model designs, including: MoE routing strategies Multimodal fusion architectures SSM / hybrid architectures
  • Design architectures with KV cache efficiency and inference implications in mind

What they require

  • Strongly Required Deep theoretical and practical understanding of modern deep learning
  • Hands ‑ on experience training large models from scratch
  • Ability to reason about optimization, not just tune hyperparameters
  • Comfort operating in ambiguous, research ‑ driven environments
  • Critical Technical Skills Transformer internals and attention mechanisms Optimi s ation algorithms and training dynamics Scaling laws and compute/data tradeoffs Distributed training strategies and mixed precision Architecture innovation for large, real ‑ world models
  • Common Problems You’ll Be Solving Why training diverges at scale How optimizer dynamics interact with architecture When scaling laws break down The real tradeoffs between data, compute, and model design

Benefits

  • Salary range $185,000.00 - $280,000.00
  • It is not typical for offers to be made at or near the top of the range. The actual salary will be determined based on experience and other job-related factors.
  • Annual bonus opportunity
  • Insurance coverage (medical, dental, vision, life, and disability)
  • Paid time off
  • Paid holidays
  • Company contribution to the RRSP (Registered Retirement Savings Plan)
  • Equity awards for certain positions and levels
  • Remote and/or hybrid work available depending on the position
  • All compensation and benefits are subject to the terms and conditions of the underlying plans or programs, as applicable, and may be amended, terminated, or replaced from time to time.

American multinational software company that develops artificial intelligence assistant technology primarily for automobiles

🇺🇸 United StatesAutomotive AIMid-sizecerence.com/
$185k–$280k/yr