Principal ML System Engineer
- Role
- AI / ML
- Experience
- Principal
- Employment
- Full-time
Open to US only. Set where you work from to check your eligibility.
Core skills
Required skills
Optional skills
At PointClickCare our mission is simple: to help providers deliver exceptional care. And that starts with our people. As a leading health tech company that’s founder-led and privately held, we empower our employees to push boundaries, innovate, and shape the future of healthcare.
With the largest long-term and post-acute care dataset and a Marketplace of 400+ integrated partners, our platform serves over 30,000 provider organizations, making a real difference in millions of lives. We also reinvest a significant percentage of our revenue back into research and development, ensuring our employees have the resources to innovate and make a lasting impact. Recognized by Forbes as a top private cloud company and honored as one of Canada’s Most Admired Corporate Cultures, we offer flexibility, growth opportunities, and meaningful work.
At PointClickCare, we empower our people to be the architects of a smarter healthcare future; one that is human-first and accelerated by AI to create meaningful and lasting change. Employees harness AI as a catalyst for creativity, productivity, and thoughtful decision-making. By integrating AI tools into our daily workflows, collaboration is enhanced, outcomes are improved, and every team member has the proficiency to maximize their impact. It all starts with our hiring practices where we uncover AI expertise that complements our mission, and we continue to invest in training and development to nurture innovation throughout the employee journey.
Join us in redefining healthcare — so it doesn’t just survive, it thrives. To learn more about PointClickCare, check out Life at PointClickCare and connect with us on Glassdoor and LinkedIn.
**Travel to Office expectations** For Remote Roles: If this role is remote, there will be in-office events that will require travel to and from the Mississauga and/or Salt Lake City office. These will include, but not limited to, onboarding, team events, semi-annual and annual team meetings.
For Hybrid Roles: If this role is Hybrid, there will be an expectation to reside within commutable distance to the office/location specified in the job listing. This will include, but not limited to, weekly/bi-weekly/monthly events in the office with your specific team. This is a requirement for this role.
Team Summary This team will serve as the product owner for the machine learning platform capabilities within PointClickCare, working closely with other engineering teams across the organization to identify, build and support traditional machine learning (ML) and hybrid ML/LLMsolutions. This centralized team with deep specialization will closely integrate with key horizontal partners to ensure delivery of safe, scalable, and high-impact AI products. Job Summary The Principal AI Machine Learning Platform Engineer will set the technical vision and strategy for the machine learning platform that powers ML and generative AI development across PointClickCare, partnering with Product and Engineering leadership to align that direction with product and business goals. As the technical authority for the ML platform, the Principal Engineer will define the reference architectures, standards, and roadmap for the pipelines, tooling, and infrastructure used for model training, deployment, serving, and monitoring, and will provide technical leadership and mentorship to raise the engineering bar for ML systems company-wide. Key Responsibilities Partner with product and engineering leadership to translate business and product objectives into a multi-quarter technical strategy and roadmap for the ML platform. Define the reference architectures and standards for scalable data and ML pipelines spanning model training, evaluation, deployment, and serving that engineering teams across the organization build upon. Set the direction and best practices for MLOps across the company — including CI/CD for models, model registry, feature stores, and experiment tracking — and drive build-vs-buy decisions for core platform components. Establish the practices and architecture for reliability, observability, and performance of ML systems in production, including monitoring, alerting, and automated remediation. Establish the security architecture for the ML platform, including authentication, role-based access control, audit logging, and compliance monitoring, and ensure adoption across teams. Define secure, cost-efficient integration and infrastructure patterns for connecting the platform with existing systems, APIs, and data sources at scale. Provide technical leadership and mentorship across engineering teams, guiding senior engineers and influencing the org-wide technical roadmap for ML infrastructure. Qualifications & Skills Expert level in Python and Java with strong software engineering fundamentals. Deep experience designing and building ML platforms and ML Ops workflows at scale, familiarity with tools such as MLFlow, Kubeflow, Ray, and model-serving frameworks or equivalents. Extensive experience with cloud platforms (AWS, Azure, and/or GCP), containerization, and orchestration (Docker, Kubernetes). Demonstrated track record of setting technical direction and driving org-wide technical initiatives across multiple teams. Preferred Bachelor’s degree or higher in Computer Science, Machine Learning, or a related field. Sufficient familiarity with Azure Machine Learning components, Databricks processing and serverless environments, and ML Frameworks to support strategic decision making Demonstrable history of leading and sustaining build out of critical cross-team systems Experience implementing security at scale including role-based access control, multi-factor authentication, network security best practices, and compliance monitoring. Experience optimizing large model training and inference (including LLM serving) for performance and cost. #LI-AJ1 #LI-remote
What you'll do
- Partner with product and engineering leadership to translate business and product objectives into a multi-quarter technical strategy and roadmap for the ML platform.
- Define the reference architectures and standards for scalable data and ML pipelines spanning model training, evaluation, deployment, and serving that engineering teams across the organization build upon.
- Set the direction and best practices for MLOps across the company — including CI/CD for models, model registry, feature stores, and experiment tracking — and drive build-vs-buy decisions for core platform components.
- Establish the practices and architecture for reliability, observability, and performance of ML systems in production, including monitoring, alerting, and automated remediation.
- Establish the security architecture for the ML platform, including authentication, role-based access control, audit logging, and compliance monitoring, and ensure adoption across teams.
What they require
- Expert level in Python and Java with strong software engineering fundamentals.
- Deep experience designing and building ML platforms and ML Ops workflows at scale, familiarity with tools such as MLFlow, Kubeflow, Ray, and model-serving frameworks or equivalents.
- Extensive experience with cloud platforms (AWS, Azure, and/or GCP), containerization, and orchestration (Docker, Kubernetes).
- Demonstrated track record of setting technical direction and driving org-wide technical initiatives across multiple teams.
- Preferred: Bachelor’s degree or higher in Computer Science, Machine Learning, or a related field.
Benefits
- Benefits starting from Day 1!
- Retirement Plan Matching
- Flexible Paid Time Off
- Wellness Support Programs and Resources
- Parental & Caregiver Leaves
Health Tech