Skip to main content
H1

Staff Data Engineer- Data Lake

RemoteUnited States only
Published
Role
Data Engineering
Experience
Staff
Employment
Full-time
$190k–$220k/yr
Check eligibility

Open to US only. Set where you work from to check your eligibility.

No BS summary

Staff Data Engineer for Data Lake team. Requires 8+ years of experience in data engineering/software engineering, with a focus on building and scaling distributed data platforms. Must have strong Python (PySpark), SQL, and Spark expertise, preferably in AWS. Experience with containerization and orchestration tools is needed. The role involves technical leadership, mentoring, and contributing to platform strategy, with an interest in growing into an Engineering Manager track.

Core skills

Data LakeETLELT

Required skills

PythonPySparkSQLApache SparkAWSEMRGlueS3AthenaRedshiftArgoAirflowParquetAvroORCDockerKubernetes

Optional skills

large-scale data cleaningparsingnormalizationvalidation workflowshealthcare datasetslife sciences datasetspublication datasetslarge-scale entity-resolution datasets

Required languages

English

What you'll do

  • Architect, build, and scale distributed ETL/ELT pipelines and large-scale ingestion frameworks across structured and unstructured healthcare datasets.
  • Lead the evolution of H1’s Data Lake architecture with a focus on scalability, observability, reliability, and cost optimization.
  • Own and improve data quality, validation, normalization, and standardization workflows across thousands of global data sources.
  • Design and optimize batch and near real-time data processing frameworks using cloud-native distributed systems.
  • Optimize distributed compute and storage systems, including Spark workloads, query performance, partitioning strategies, and infrastructure efficiency.
  • Drive improvements in monitoring, governance, operational excellence, and production reliability across the platform.
  • Troubleshoot complex production data and infrastructure issues across distributed systems.
  • Partner closely with Product, Infrastructure, Security, Compliance, and downstream engineering teams to support scalable and secure data delivery.
  • Mentor engineers through technical leadership, architecture reviews, and engineering best practices.
  • Help define technical roadmap priorities and contribute to long-term platform strategy and execution planning.
  • Support production operations, incident response, and platform health as part of overall ownership of the Data Lake ecosystem.

What they require

  • 8+ years of experience in data engineering, software engineering, or related fields with significant experience building and scaling distributed data platforms.
  • Demonstrated technical leadership experience with interest in or experience mentoring and leading engineers.
  • Strong proficiency in Python (PySpark), Java, Scala, or similar programming languages.
  • Advanced SQL expertise, including performance tuning and optimization across large datasets.
  • Deep experience with Apache Spark and cloud-native big data platforms, preferably within AWS environments (EMR, Glue, S3, Athena, Redshift, or similar).
  • Experience designing and scaling modern cloud-native data lake architectures and large-scale ingestion frameworks.
  • Experience with orchestration and workflow management tools such as Argo, Airflow, or similar technologies.
  • Strong understanding of distributed storage systems, partitioning strategies, and file formats such as Parquet, Avro, and ORC.
  • Experience with Docker, Kubernetes, and modern containerization technologies.
  • Experience implementing monitoring, observability, and data quality frameworks within production environments.
  • Experience with large-scale data cleaning, parsing, normalization, and validation workflows preferred.
  • Experience working with healthcare, life sciences, publication, or large-scale entity-resolution datasets preferred.
  • Exposure to ML/AI-driven data enrichment, parsing, or validation workflows is a plus.
  • Experience using AI-assisted coding tools (e.g., GitHub Copilot, Claude Code) to accelerate development while maintaining quality is encouraged

Benefits

  • Full suite of health insurance options, in addition to generous paid time off
  • Pre-planned company-wide wellness holidays
  • Retirement options
  • Health & charitable donation stipends
  • Impactful Business Resource Groups
  • Flexible work hours & the opportunity to work from anywhere
  • The opportunity to work with leading biotech and life sciences companies in an innovative industry with a mission to improve healthcare around the globe

H1

H1 provides a platform to inform doctor interactions globally, promoting health equity and trust in healthcare systems.

🇺🇸 United StatesHealthcareMid-size
$190k–$220k/yr