Skip to main content
Protege

Technical Product Manager, Data Ingestion & Quality

RemoteNot specified
Published
Role
Product
Experience
Mid
Employment
Full-time
Company size
Startup
Salary not disclosed
Check eligibility

The listing doesn't say where it hires from. Check the description or the employer's site before applying.

No BS summary

Product Manager with 4–7 years owning data pipelines, data quality systems, or ingestion platforms. Must be hands-on with SQL, pipeline logs, schema mismatches, external messy data, and technical requirements for engineering. Best fit has worked on raw-data-to-trusted-data infrastructure, not BI dashboards or customer-facing analytics.

Core skills

SQL

Optional skills

dbtGreat Expectationsdata contracts

What you'll do

  • Own the supply side of Protege's data platform — the pipeline that takes raw data from a partner and turns it into something catalog-ready, trustworthy, and usable
  • Own the infrastructure that makes data trustworthy enough to build products from: validation gates, metadata generation pipelines, QA standards, de-identification transformations, and catalog-readiness criteria
  • Work across healthcare, media, and any other vertical Protege enters
  • Write SQL, review pipeline outputs, define what "good" looks like at each stage of ingestion, and translate those standards into platform requirements that engineering can build against
  • Define the stages, validation gates, and quality checks that data passes through from partner arrival to catalog-ready
  • Own the platform requirements that make ingestion repeatable across modalities and verticals
  • Own the product decisions around what metadata gets extracted or generated at ingestion, including transcripts, tags, confidence scores, schema inference, thresholds, and how it gets stored and surfaced
  • Define what "catalog-ready" means
  • Build the tooling that enforces QA standards
  • Get into the data directly to validate that standards are being met
  • Run queries and review pipeline outputs, not just read dashboards
  • Work with vertical stakeholders to translate their readiness requirements into consistent platform-level standards that don’t require custom engineering per deal
  • Build a clear understanding of Protege’s current data ingestion workflow, including how raw partner data moves from arrival to catalog-ready
  • Get hands-on with pipeline outputs, schemas, metadata, validation checks, and QA processes to understand where quality risk shows up in practice
  • Build context with engineering, vertical stakeholders, GTM / delivery, and DataLab on where ingestion quality is most manual, inconsistent, or risky today
  • Own the first clear version of what “catalog-ready” means across the ingestion pipeline, including validation gates, metadata requirements, QA standards, and readiness criteria
  • Translate the highest-priority ingestion quality gaps into product requirements engineering can build against
  • Own the roadmap for improving ingestion quality, metadata generation, QA tooling, de-identification workflows, and catalog readiness
  • Create a repeatable operating rhythm for reviewing pipeline outputs, quality signals, and ingestion risks with the right cross-functional partners

What they require

  • 4–7 years of PM experience where the core product was a data pipeline, data quality system, or data ingestion platform — you’ve owned the "raw data in, trusted data out" problem before
  • Hands-on technical depth — you can write SQL, read pipeline logs, spot a schema mismatch, and understand the tradeoffs in a data validation architecture; you look at data directly to verify things are working, not just at metrics
  • Experience with external data — you’ve worked on a product that ingested messy, inconsistently formatted data from third-party partners and had to make it trustworthy; you know what that problem actually feels like
  • Build-versus-partner judgment — you’ve made vendor decisions in a fast-moving technical domain; you know how to evaluate a tool against requirements that will change, and how to structure relationships that preserve flexibility
  • Cross-functional credibility — you’ll be writing requirements that multiple engineering teams and vertical PMs depend on; you can hold a technical conversation and a product conversation in the same meeting
  • Preferred: Experience with data quality frameworks, metadata standards, or catalog tooling
  • Preferred: Familiarity with de-identification approaches for sensitive data — PHI, PII, or confidential enterprise data
  • Preferred: Background in healthcare data operations, financial data infrastructure, or any domain where data quality has real downstream consequences
  • Preferred: Exposure to ML training pipelines or AI data workflows, where data fitness affects model outcomes
  • Preferred: Experience with data governance strategies
  • Not primarily owned data products from the customer side — analytics dashboards, BI tooling, or data visualization
  • You have actually dug into raw data files to find out why a pipeline produced the wrong output

Protege is building a platform for secure, efficient, privacy-centric exchange of AI training data. DataLab is Protege's research arm focused on data for AI.

HealthcareStartup
Salary not disclosed