Skip to main content
Protege

Applied Healthcare Researcher

RemoteUnited States only
Published
Role
Research
Experience
Mid
Employment
Full-time
Company size
Startup
Salary not disclosed
Check eligibility

Open to US only. Set where you work from to check your eligibility.

No BS summary

Healthcare applied researcher with advanced quantitative/ML background and hands-on ML/LLM systems experience. Must know real-world healthcare data, Python, SQL, dataset evaluation, and customer-facing research.

Core skills

MLLLMSQL

Required skills

Python

What you'll do

  • Serve as the primary technical and research point of contact for healthcare customer conversations.
  • Translate a lab's model-development goals into concrete, feasible data strategies.
  • Help customers scope opportunities and identify the highest-value data available to them.
  • Explain data limitations, tradeoffs, and potential biases to technically sophisticated stakeholders while grounding conversations in what real-world data actually looks like.
  • After delivery, answer the research questions customers raise about the data we provided. Delivery is not the end of the relationship.
  • Develop and evaluate methods — fine-tuning, LLM-based extraction, classification, rules-based approaches, or whatever the problem calls for — to demonstrate that a dataset can support a customer's training or evaluation objective.
  • Design and run feasibility research pre-contract: can this data support this model objective, at what quality, with what caveats.
  • Build the evidence base that makes a data strategy credible — benchmarks, validation analyses, error characterization, and honest assessments of where the data falls short.
  • Partner with the Assessments team on healthcare benchmarks across modalities.
  • Evaluate whether requested variables, labels, or cohort definitions are achievable with available healthcare data.
  • Identify proxy variables or alternative dataset structures when the ideal variable doesn't exist.
  • Analyze partner and source datasets — schema, field availability, quality, completeness, and required transformations.
  • Contribute to our point of view on which healthcare data matters most for which modality and which stage of model development.
  • Help evaluate new data partners and identify datasets worth acquiring before a customer asks for them.
  • Produce reusable research, evidence, and technical collateral rather than starting from scratch for each opportunity.
  • Identify where a successful one-off approach should become a repeatable workflow, and work with Product and Engineering to operationalize it.
  • Help expand proven healthcare datasets across multiple customers instead of selling them once.
  • Work with Solutions and FDEs from the beginning of an opportunity.
  • Coordinate with Healthcare Data Partnerships on sourcing and with Product and Engineering on tooling.

What they require

  • Advanced degree (PhD or Master's plus 2+ years industry experience) in machine learning, computer science, biomedical informatics, epidemiology, statistics, or a related quantitative field — or equivalent applied experience.
  • Hands-on experience building and evaluating ML or LLM-based systems for extraction, classification, or prediction on real-world data.
  • Experience working with healthcare data: claims, EMR/EHR, clinical notes, imaging, registries, or similar. You understand why real-world clinical data is messy and what that means for model training.
  • Strong Python and SQL, with the ability to work independently against large datasets.
  • Experience designing evaluations — measuring data quality and dataset representativeness.
  • Demonstrated ability to work directly with technical stakeholders and translate ambiguous goals into concrete, defensible research plans.
  • Comfort operating on a customer's timeline without lowering the standard of the research.
  • Is energized by working directly with customers, and specifically by working with other researchers as peers.
  • Moves fast on messy, real-world problems and knows which corners can and cannot be cut.
  • Is rigorous about what the data can and cannot support, and willing to tell a customer when the answer is no.
  • Enjoys the full arc — scoping a vague problem, doing the research, and showing the result to the person who asked for it.
  • Thinks about leverage: builds the reusable version rather than the one-off when it's worth doing.

Protege is building a platform for secure, efficient, privacy-centric exchange of AI training data. DataLab is Protege's research arm focused on data for AI.

HealthcareStartup
Salary not disclosed