Skip to main content

Senior Data Engineer

RemoteUnited States only
Published
Role
Data Engineering
Experience
Senior
Employment
Full-time
Salary not disclosed
Check eligibility

Open to US only. Set where you work from to check your eligibility.

No BS summary

Senior data engineer for a remote Herndon, VA role requiring Public Trust Clearance. Needs 5+ years SQL/T-SQL and cloud ELT/ETL, 3+ years Azure Synapse/Azure Machine Learning, and Python/Pandas.

Core skills

Azure SynapseAzure Machine LearningPython

Required skills

SQLT-SQLELTETLPandas

Optional skills

DP-203PySparkPolarsTerraformBicepCLIREST APIs

What you'll do

  • Provide authoritative expertise on data engineering methods and best practices, including code first development approaches and modern pipeline design patterns.
  • Design, implement, and maintain the data architecture that supports products and end users, with all assets managed under source control.
  • Design, implement, and maintain ELT and ETL pipelines for efficient processing of source data in Azure Synapse and Azure Machine Learning, using both SDK V1 and SDK V2.
  • Migrate source data identified by SBA OIG into Azure Data Lake Storage.
  • Normalize entity attributes such as addresses, phone numbers, and other common fields.
  • Review, maintain, and improve existing architecture and pipelines, including periodic audits addressing bottlenecks, deprecated dependencies, and architecture drift.
  • Establish quality controls across all pipelines and introduce error handling, logging mechanisms, and validation checks.
  • Incorporate source control across all pipelines and analytics codebases so code can evolve iteratively without destabilizing the architecture.
  • Optimize ingestion, processing, and storage across a wide variety of datasets and data types, including modern columnar formats such as Parquet.
  • Develop self service capabilities that let SBA OIG analysts query and export data for investigations and audits.
  • Author robust standard operating procedures governing the authoring, development, validation, publishing, execution, and monitoring of all data pipelines and assets in the Azure environment.
  • Produce detailed documentation of the data architecture, including data dictionaries, entity relationship diagrams, and pipeline process maps.
  • Maintain and expand the environment with additional datasets and services on request, following a defined intake and testing process before production deployment.
  • Stay current with emerging AI tooling relevant to data engineering and contribute to exploratory work evaluating automation and language model assisted capabilities.

What they require

  • Must have an Public Trust Clearance
  • Bachelor's degree in data engineering, computer science, data science, machine learning, mathematics, or a related field. Alternatively, five years of applied work experience in any of the same fields.
  • 5 years - Maintaining SQL databases and conducting advanced operations in SQL and T-SQL.
  • 5 years - Designing, implementing, and maintaining ELT and ETL processes in cloud based data analytics environments.
  • 3 years - Working in Azure Synapse and Azure Machine Learning with the modern data stack.
  • 3 years -Manipulating data in Python. Pandas is required.
  • Preferred: Certifications preferred, DP-203 or equivalent.
  • Preferred: Experience developing reusable, modular code preferred.
  • Preferred: DP-203, Microsoft Certified Azure Data Engineer Associate, or an equivalent current certification.
  • Preferred: Implementing pipelines and infrastructure using code first approaches: Python SDK, CLI, REST APIs, or infrastructure as code tooling such as Terraform or Bicep.
  • Preferred: Implementing source control and continuous integration and delivery workflows for data assets.
  • Preferred: Demonstrated familiarity with AI coding assistants and large language model integration patterns.
  • Preferred: PySpark or Polars at production scale.
  • Preferred: Entity resolution and attribute normalization across records with inconsistent addresses, names, and identifiers.
  • Preferred: Building self service analytic access for non engineering users.

Benefits

  • Medical
  • Dental
  • Vision
  • Basic Life
  • Health Saving Account
  • 401K matching
  • Three weeks of PTO/Sick
  • 11 Paid Holidays
  • Pre-Approved Online Training
Salary not disclosed