Skip to main content
Irth Solutions

Data Engineer RD-017

RemoteCanada only
Published
Role
Data Engineering
Experience
Mid
Employment
Full-time
CAD 75k+/yr
Check eligibility

Open to CA only. Set where you work from to check your eligibility.

No BS summary

Mid-level Data Engineer (3–5 yrs) to build and operate Databricks-based data ingestion and processing pipelines. Must be proficient in Python, SQL, Spark/PySpark and Delta Lake; experience integrating high-volume external APIs. Remote role limited to Canada with strong preference for candidates based in Quebec; French is an asset.

Core skills

DatabricksPySparkDelta Lake

Required skills

Apache SparkPythonSQLDelta Live TablesAPI integrationData pipeline developmentDelta Lake optimization (OPTIMIZE, Z-ORDER, VACUUM)Databricks WorkflowsGitCI/CDAzure/AWS/GCP

Optional skills

LLM prompting / prompt engineeringAI/NLP basicsMedallion architecture / lakehouse best practicesOrchestration frameworks (ADF, Databricks Workflows, Airflow)CI/CD tools and GitHub ActionsDatabricks certification (Data Engineer Associate)Basic security practices (RBAC, encryption, credential management)

Optional languages

English Fluent (required)French Fluent (asset)

What you'll do

  • Design, build, and maintain ingestion pipelines from high-volume external APIs capable of running continuously and reliably at scale.
  • Implement ingestion and transformation workflows using Databricks (Spark/PySpark, SQL, Delta Live Tables) applying medallion architecture (Bronze → Silver → Gold).
  • Build infrastructure for deduplication and relevance filtering of incoming content and implement filtering logic and quality criteria with the Data Scientist.
  • Implement schema evolution handling and data validation rules as sources and formats change over time.
  • Configure and manage Delta Lake storage structures, tables, partitions, and optimization routines (OPTIMIZE, Z-ORDER, VACUUM).
  • Design and evolve data schemas balancing query performance, cost, and maintainability.
  • Maintain clear metadata and documentation of table structures for consumption by Data Science and application teams.
  • Ensure pipeline reliability and observability: error handling, retries, monitoring, and alerting for continuously running systems.
  • Adapt pipelines to evolving external API contracts, rate limits, authentication changes, and new data sources.
  • Troubleshoot pipeline failures, perform recovery, and tune performance.
  • Build, schedule, and monitor workflows using Databricks Workflows, Delta Live Tables, or similar orchestration tools.
  • Contribute to CI/CD pipelines for code deployment, versioning, and environment management.
  • Collaborate with the Data Scientist to expose clean, well-structured data feeding LLM/NLP pipelines and downstream models.
  • Participate in technical decisions around data architecture and propose structuring solutions as needs evolve.
  • Document pipelines, data dictionaries, job schedules, and transformation logic.
  • Support onboarding of new data sources and pipelines as the product expands.

What they require

  • 3 to 5 years of experience in data engineering, with solid experience building and operating production-grade data pipelines.
  • Familiarity with data modeling, data quality, and schema evolution.
  • Solid understanding of data pipeline reliability practices: monitoring, alerting, and handling failures in a continuously running system.
  • Hands-on experience with Databricks or an equivalent Spark-based environment: schema design, Delta Lake, performance tuning, and pipeline orchestration.
  • Experience with at least one major cloud provider (Azure preferred; AWS/GCP also beneficial).
  • Experience integrating with external APIs at scale: authentication, pagination, rate limiting, retries, and error handling.
  • Strong proficiency in Python and SQL.
  • Comfortable working with unstructured/semi-structured text data at scale.
  • Strong preference for candidates residing in Quebec; fluency in French (spoken and written) is a strong asset in addition to English.

Benefits

  • Competitive salary (base)
  • Medical, dental, and vision insurance
  • 401(k) plan with company match
  • Generous paid time off (PTO)
  • Company-paid holidays
  • Flexible work options / work-from-home opportunities
  • On-call compensation for eligible shifts

Irth Solutions is a software product company building cutting-edge technology platforms that continuously set industry benchmarks. With a strong product culture, collaborative environment, and high growth potential, Irth offers an exciting opportunity to work on modern, enterprise-scale data platforms.

🇺🇸 United StatesSoftwareStartup

Details

Posting languageEnglish
CAD 75k+/yr