Skip to main content
DeepL

Senior Research Scientist | Model Scaling

RemoteUnited Kingdom only
Published
Role
Research
Experience
Senior
Employment
Full-time
Company size
Mid-size
Salary not disclosed
Check eligibility

Open to GB only. Set where you work from to check your eligibility.

Core skills

large language modelsmodel scaling

Required skills

PythonPyTorchJAXTensorFlowLoRAPEFT

Optional skills

machine translationmultilingual NLPdocument-/layout-aware modellingMixture-of-LoRA-Experts

What you'll do

  • Drive the selection and evaluation of open foundation / open-weight models as the basis for our next-generation translation systems.
  • Lead model selection and general architecture decisions for scaling to hundreds of billions of parameters, including Mixture-of-Experts and other sparse or efficient designs.
  • Design multi-capability adaptation strategies using LoRA, PEFT, and related methods.
  • Own the modelling lifecycle for your work: prototyping, ablations, scaling experiments, evaluation, and delivery into production, with rigorous and reproducible evaluation.
  • Partner closely with post-training, RL/RLHF, and instruction-following specialists to integrate alignment and capability work into the base model.
  • Stay ahead of the open-model and scaling literature, and bring well-founded recommendations back to the team.

What they require

  • Strong hands-on experience adapting and scaling large language models via fine-tuning, instruction-tuning, or post-training of multi-billion-parameter models beyond black-box use.
  • Sound judgment about architecture trade-offs at scale (e.g. dense vs. MoE) and about which open-weight foundation models to build on.
  • Working knowledge of parameter-efficient and multi-capability adaptation (LoRA/PEFT and variants).
  • A hands-on builder who enjoys training models, running experiments, and debugging pipelines, and who can carry research results through to production with engineering.
  • Strong coding and experimentation skills (Python, PyTorch/JAX/Tensorflow).
  • Ability to communicate clearly, collaborate across teams, and align research work with product and engineering priorities.
  • Experience quantifying uncertainty in large models — calibration and confidence estimation via Bayesian methods, ensembling, steering, or prompt-based approaches.
  • Experience with machine translation, multilingual NLP, or document-/layout-aware modelling experience.
  • Familiarity with MoE-specific training and adaptation (e.g. expert routing, Mixture-of-LoRA-Experts) and large-scale data-mixture design.

Benefits

  • Diverse and internationally distributed team: joining our team means becoming part of a large, global community with people of more than 90 nationalities.
  • Open communication, regular feedback
  • Hybrid work, flexible hours
  • Virtual Shares - An ownership mindset in every role.
  • Regular in-person team events
  • Monthly full-day hacking sessions
  • 30 days of annual leave
  • Competitive benefits

AI-powered machine translation service developed by DeepL SE

AIEnterprisedeepl.com/
Salary not disclosed