Перейти к основному содержимому
УдалённоUnited States только
Опубликовано
Роль
Инженерия данных
Опыт
Синьор
Занятость
Полная занятость
$134.6k–$194.5k/yr
Проверьте доступность

Доступно для: US only. Укажите, откуда вы работаете, чтобы проверить доступность.

Коротко по делу

Senior Data Engineer needed to build and operate enterprise-scale data solutions on AWS Lakehouse platform. Requires 5+ years of experience in data engineering, Python, PySpark, and cloud platforms (AWS preferred). Must have strong SQL and understanding of data warehousing and Lakehouse architectures.

Ключевые навыки

PythonPySparkAWS

Обязательные навыки

AWS GluedbtApache SparkSQLGitAgile

Желательные навыки

NoSQLApache IcebergTerraformApache AirflowMWAAAWS Step FunctionsRESTful APIsKafka

Чем предстоит заниматься

  • Design and implement ELT/ETL solutions for batch and streaming ingestion, integration, refinement, and publish patterns on the Lakehouse.
  • Develop reusable data processing frameworks and configuration-driven pipelines using Python and PySpark (EMR, Glue, or comparable Spark runtimes).
  • Build and maintain scalable orchestration workflows (e.g., Airflow) for production data delivery, including retries, historical loads, and operational runbooks.
  • Implement data quality checks, validation frameworks, and monitoring so data products meet defined contracts and SLAs.
  • Apply DataOps practices: Git-based development, CI/CD/CT for data pipelines, automated testing, and controlled promotion across environments.
  • Contribute to data lifecycle practices (retention, archival, disaster recovery / resiliency considerations) in partnership with platform and governance teams.
  • Support platform modernization and cloud migration of legacy data flows into Lakehouse patterns (Iceberg on S3, governed catalog access).
  • Collaborate with stakeholders to map technical designs to business processes, non-functional requirements, and consumption needs (Athena, Redshift, APIs, exports, streams).
  • Establish and document standards, naming/conventions, and engineering practices; participate in Agile ceremonies and cross-team delivery.
  • Provide technical leadership: mentor engineers, conduct design and code reviews, and continuously improve reliability, performance, and cost efficiency.

Что требуется

  • Bachelor's degree in Computer Science, Information Systems, Engineering, or a related technical field, or equivalent work experience.
  • 8+ years of overall IT experience.
  • 5+ years of hands-on experience designing and developing enterprise-scale data engineering solutions.
  • Strong experience developing scalable data pipelines and reusable frameworks using Python and PySpark.
  • Experience implementing enterprise data ingestion, integration, and transformation solutions using AWS Glue, dbt, Apache Spark, or comparable technologies.
  • Strong understanding of data warehousing concepts, dimensional modeling, and modern data lake / Lakehouse architectures.
  • Experience with cloud platforms such as AWS, Microsoft Azure, or Google Cloud Platform (AWS preferred).
  • Strong SQL expertise with relational databases; familiarity with NoSQL databases is a plus.
  • Experience with distributed data processing technologies such as Apache Spark, Amazon EMR, or Hadoop-based platforms.
  • Experience using Git-based source control and Agile software development methodologies.
  • Strong analytical, problem-solving, and communication skills, with the ability to collaborate across technical and business teams.
  • Preferred Qualifications: Experience designing cloud-native data platforms using AWS services such as S3, Glue, EMR, Athena, Redshift, Lambda, Lake Formation, IAM, and CloudWatch.
  • Hands-on experience with Apache Iceberg (or similar open table formats) and governed Lakehouse patterns.
  • Experience implementing CI/CD pipelines, DataOps practices, continuous testing (CT), and infrastructure automation (e.g., Terraform).
  • Experience with workflow orchestration tools such as Apache Airflow, MWAA, AWS Step Functions, or similar platforms.
  • Experience developing RESTful APIs or other access layers and integrating with enterprise applications for data-product consumption.
  • Experience building and supporting real-time or streaming platforms using Kafka, Kinesis, or Spark Structured Streaming.
  • Strong understanding of data quality, observability, monitoring, and automated validation frameworks.
  • Knowledge of data governance, metadata management, lineage, and enterprise data catalog solutions.
  • Strong understanding of data security, encryption, access controls, and healthcare regulatory compliance (HIPAA/PHI).
  • Experience optimizing distributed workloads for scalability, reliability, and cloud cost efficiency.
  • Experience mentoring engineers, conducting design and code reviews, and establishing engineering best practices.

Преимущества

  • medical, dental and vision coverage
  • incentive and recognition programs
  • life insurance
  • 401k contributions (all benefits are subject to eligibility requirements)
HealthcareКрупная
$134.6k–$194.5k/yr