Перейти к основному содержимому
DeepL

Senior Research Scientist | Model Scaling

УдалённоUnited Kingdom только
Опубликовано
Роль
Исследования
Опыт
Синьор
Занятость
Полная занятость
Размер компании
Средняя
Зарплата не указана
Проверьте доступность

Доступно для: GB only. Укажите, откуда вы работаете, чтобы проверить доступность.

Ключевые навыки

large language modelsmodel scaling

Обязательные навыки

PythonPyTorchJAXTensorFlowLoRAPEFT

Желательные навыки

machine translationmultilingual NLPdocument-/layout-aware modellingMixture-of-LoRA-Experts

Чем предстоит заниматься

  • Drive the selection and evaluation of open foundation / open-weight models as the basis for our next-generation translation systems.
  • Lead model selection and general architecture decisions for scaling to hundreds of billions of parameters, including Mixture-of-Experts and other sparse or efficient designs.
  • Design multi-capability adaptation strategies using LoRA, PEFT, and related methods.
  • Own the modelling lifecycle for your work: prototyping, ablations, scaling experiments, evaluation, and delivery into production, with rigorous and reproducible evaluation.
  • Partner closely with post-training, RL/RLHF, and instruction-following specialists to integrate alignment and capability work into the base model.
  • Stay ahead of the open-model and scaling literature, and bring well-founded recommendations back to the team.

Что требуется

  • Strong hands-on experience adapting and scaling large language models via fine-tuning, instruction-tuning, or post-training of multi-billion-parameter models beyond black-box use.
  • Sound judgment about architecture trade-offs at scale (e.g. dense vs. MoE) and about which open-weight foundation models to build on.
  • Working knowledge of parameter-efficient and multi-capability adaptation (LoRA/PEFT and variants).
  • A hands-on builder who enjoys training models, running experiments, and debugging pipelines, and who can carry research results through to production with engineering.
  • Strong coding and experimentation skills (Python, PyTorch/JAX/Tensorflow).
  • Ability to communicate clearly, collaborate across teams, and align research work with product and engineering priorities.
  • Experience quantifying uncertainty in large models — calibration and confidence estimation via Bayesian methods, ensembling, steering, or prompt-based approaches.
  • Experience with machine translation, multilingual NLP, or document-/layout-aware modelling experience.
  • Familiarity with MoE-specific training and adaptation (e.g. expert routing, Mixture-of-LoRA-Experts) and large-scale data-mixture design.

Преимущества

  • Diverse and internationally distributed team: joining our team means becoming part of a large, global community with people of more than 90 nationalities.
  • Open communication, regular feedback
  • Hybrid work, flexible hours
  • Virtual Shares - An ownership mindset in every role.
  • Regular in-person team events
  • Monthly full-day hacking sessions
  • 30 days of annual leave
  • Competitive benefits

AI-powered machine translation service developed by DeepL SE

AIКрупнаяdeepl.com/
Зарплата не указана