Skip to main content
Avomind

Principal Machine Learning Engineer

RemoteSingapore only
Published
Role
AI / ML
Experience
Principal
Employment
Full-time
Company size
Startup
Salary not disclosed
Check eligibility

Open to SG only. Set where you work from to check your eligibility.

No BS summary

Principal ML engineer for production-grade AI/LLM systems in Singapore. Needs deep learning, transformer architectures, large-scale model training/fine-tuning/deployment, PyTorch or JAX, distributed training/inference, and GPU optimization.

Core skills

GPU optimizationMachine LearningLLMs

Required skills

Deep LearningTransformer architecturesPyTorch/JAXDeepSpeed/FSDP/Megatron/ZeRO/RayQuantizationMixed precision

Optional skills

vLLMTensorRT-LLMFasterTransformerOpen-source machine learning librariesOpen-source systems librariesScientific computingCompiler technologiesGPU kernel development

What you'll do

  • Build and own end-to-end machine learning pipelines covering data processing, model training, evaluation, inference, and deployment.
  • Fine-tune and adapt models using modern techniques such as LoRA, QLoRA, Supervised Fine-Tuning (SFT), Direct Preference Optimization (DPO), and model distillation.
  • Design and operate scalable inference systems while balancing latency, cost, and reliability.
  • Develop and maintain data pipelines for both synthetic and real-world training datasets.
  • Build evaluation frameworks to assess model performance, robustness, safety, and bias in collaboration with research teams.
  • Optimize production deployments through GPU optimization, memory efficiency, latency reduction, and scaling strategies.
  • Collaborate with application engineering teams to integrate machine learning systems into backend, mobile, and desktop applications.
  • Continuously improve ML systems through rapid iteration and real-world performance monitoring while balancing production constraints such as latency, cost, reliability, and safety.

What they require

  • Strong background in deep learning and transformer-based architectures.
  • Hands-on experience training, fine-tuning, or deploying large-scale machine learning models in production.
  • Proficiency with modern machine learning frameworks such as PyTorch or JAX.
  • Experience with distributed training and inference frameworks, including technologies such as DeepSpeed, FSDP, Megatron, ZeRO, or Ray.
  • Strong software engineering skills with experience building robust, maintainable, production-grade systems.
  • Experience optimizing GPU workloads, including memory efficiency, quantization, and mixed precision.
  • Ability to independently own end-to-end machine learning systems in fast-moving environments.
  • Strong problem-solving skills with a focus on rapid iteration and continuous improvement.
  • Preferred: Open-source contributions to machine learning or systems libraries.
  • Preferred: Scientific computing, compiler technologies, or GPU kernel development.
  • Preferred: Reinforcement Learning from Human Feedback (RLHF) pipelines, including PPO, DPO, or ORPO.
  • Preferred: Training or deploying multimodal or diffusion models.

Dynamic and fast-growing multinational company in the food and beverage industry; a leading FMCG producer in Southeast Europe with over 20 brands available in over 40 markets worldwide.

🇭🇰 Hong Kong SAR ChinaFood And BeverageEnterprise
Salary not disclosed