Skip to main content
Recursion

Engineering Manager - Machine Learning

RemoteCanada, United States only
Published
Role
Engineering Management
Employment
Full-time
CAD 210.1k–CAD 282.9k/yr
Check eligibility

Open to CA, US only. Set where you work from to check your eligibility.

No BS summary

Engineering manager for ML infrastructure/MLOps, distributed systems, and model deployment at scale. Must be based in Toronto, Canada and comfortable leading teams across AI/ML, LLM, agentic systems, GPU infrastructure, and cloud/on-prem supercomputing. Life sciences or drug discovery fluency is a plus, not required.

Core skills

MLOpsDistributed SystemsMachine Learning Infrastructure

What you'll do

  • Lead a team working to build, scale, and optimize the machine learning infrastructure that powers Recursion's drug discovery platform.
  • Ensure ML models can operate at massive scale across Recursion's supercomputing infrastructure, both on prem and in the cloud.
  • Work cross-functionally across ML engineering, data science, and research teams to translate requirements into robust, scalable ML infrastructure solutions.
  • Enable AI/ML, LLM, and Agentic Systems teams for scale.
  • Build and operate platforms that allow data scientists and ML engineers to train, deploy, and monitor models across Recursion's massive datasets.
  • Work closely with researchers and ML engineers to understand their infrastructure needs and build scalable solutions for model development, training, and deployment.
  • Act as a mentor, coach, and sponsor.
  • Share your technical, leadership and managerial skills in MLOps, distributed computing, and infrastructure engineering, delivering impact, learning, and growth across teams at Recursion.
  • Partner with ML research, platform engineering, and business teams.
  • Enable a model-driven culture.
  • Work with stakeholders across the business to ensure ML infrastructure supports rapid experimentation, reliable model deployment, and continuous improvement.
  • Work on problems ranging from optimizing GPU cluster utilization to implementing Agentic orchestration and establishing company-wide MLOps standards.

What they require

  • Experience in a hands-on technical role as a tech lead or a manager with a focus on infrastructure, MLOps and distributed systems.
  • Excitement for deeply engaging in technical details with your team around machine learning, orchestration and agentic systems.
  • A people-first mindset.
  • Deliver in a way that prioritizes supporting coworkers in their growth and experience and understand how Conway's Law shapes ML system outcomes.
  • Demonstrated past record of learning from and teaching peers in areas of ML infrastructure, model deployment, distributed compute, GPU optimization, and MLOps system architecture.
  • Excitement to learn parts of the ML tech stack that you might not already know.
  • Current ML infrastructure includes: Python, PyTorch, Docker, Kubernetes, Ray, Weights & Biases, Prefect, BigQuery, Postgres, GCP, CUDA, and various model serving frameworks.
  • Preferred: Fluency in life sciences or drug discovery.

Benefits

  • Eligible for an annual bonus.
  • Eligible for equity compensation.
  • Comprehensive benefits package.
  • Accommodations are available on request for candidates taking part in all aspects of the selection process.

Recursion (NASDAQ: RXRX) is a clinical-stage TechBio company decoding biology to radically improve lives. Recursion is advancing a portfolio of differentiated investigational medicines across its wholly owned and partnered pipeline in oncology, rare disease, neuroscience, immunology, and other therapeutic areas with significant unmet need.

🇺🇸 United StatesBiotechEnterprise

Details

Apply routeGreenhouse
CAD 210.1k–CAD 282.9k/yr