Skip to main content
Innovaccer

Principal Software Engineer - Data Platform (Iceberg/Trino)

RemoteUnited States only
Published
Role
Data Engineering
Experience
Principal
Employment
Full-time
Salary not disclosed
Check eligibility

Open to US only. Set where you work from to check your eligibility.

No BS summary

Principal-level data platform engineer with 12+ years building large-scale data platforms or distributed systems. Needs deep Trino/Presto or Spark SQL internals, Apache Iceberg or similar lakehouse table formats, catalog services, S3-compatible storage, and Java and/or Python. US-based role; on-premise, regulated, air-gapped, or healthcare data experience is a plus.

Core skills

Apache Iceberg/Delta Lake/HudiTrino/PrestoSpark SQL

Required skills

Polaris/Nessie/Hive MetastoreS3Snowflake/BigQuery/RedshiftJava/Python

Optional skills

Delta LakeHudi

What you'll do

  • Own the lakehouse reference architecture: Iceberg table design, Trino cluster topology, catalog service, Spark transform compute, and object-storage layout.
  • Design on-premise replacements for cloud-managed warehouse capabilities that have no direct equivalent: change-data-capture streams, scheduled tasks, and write-back paths into operational stores.
  • Run proof-of-concept validation of the catalog and query engine at expected data volumes, and define evidence-based triggers for placement decisions (VM-based versus Kubernetes-native operators).
  • Set platform-wide standards for table layout, partitioning, file sizing, and Iceberg maintenance: compaction, snapshot expiry, and orphan-file cleanup.
  • Lead the SQL dialect strategy for porting existing warehouse workloads to Trino and Spark SQL.
  • Mentor senior engineers across data workstreams, review designs, and raise the bar on engineering quality.
  • Partner with platform engineering on storage sizing, resource isolation, and capacity planning for the lakehouse footprint.

What they require

  • B.E., B.Tech., M.Sc. degree in Computer Science or a related technical field.
  • 12+ years of industry experience building and operating large-scale data platforms or distributed systems.
  • Deep, hands-on expertise with distributed SQL engines: Trino/Presto or Spark SQL internals, query planning, and performance engineering.
  • Production experience with Apache Iceberg (or Delta Lake/Hudi with willingness to go deep on Iceberg): table spec, merge-on-read versus copy-on-write, and table maintenance at scale.
  • Working knowledge of Iceberg catalog services (REST catalogs such as Polaris or Nessie, or Hive Metastore) and S3 compatible object storage.
  • Strong understanding of cloud warehouse internals (Snowflake, BigQuery, or Redshift) sufficient to design functional equivalents on open-source infrastructure.
  • Professional software development experience with Java and/or Python.
  • Preferred: Experience delivering data platforms in on-premise, regulated, or air-gapped environments is a strong plus.
  • Preferred: healthcare data experience is a plus.

Benefits

  • Generous Paid Time Off: Recharge and relax with 20 days of fixed time off per year, in addition to company holidays—because we believe work-life balance fuels performance.
  • Best-in-Class Parental Leave: Spend quality time with your growing family. We offer one of the industry’s most generous parental leave policies to support you during life’s most important moments.
  • Recognition & Rewards: We celebrate wins—big and small. Get rewarded with monetary incentives and company-wide recognition for your impact and dedication. Your hard work won’t go unnoticed.
  • Comprehensive Insurance Coverage: Stay covered with medical, dental, and vision insurance, plus 100% company-paid short- and long-term disability and basic life insurance.
  • Optional perks include discounted legal aid and pet insurance.

Innovaccer activates the flow of healthcare data, empowering providers, payers, and government organizations to deliver intelligent and connected experiences that advance health outcomes. The Healthcare Intelligence Cloud equips every stakeholder in the patient journey to turn fragmented data into proactive, coordinated actions that elevate the quality of care and drive operational performance.

Healthcareinnovaccer.com/
Salary not disclosed