Skip to main content
GoGuardian

Data Engineer II

RemoteUnited States only
Published
Role
Data Engineering
Experience
Mid
$130k–$150k/yr
Check eligibility

Open to US only. Set where you work from to check your eligibility.

No BS summary

Data Engineer II for a US-based remote role. Requires 2–4 years building large-scale data systems, strong Python/SQL, and hands-on PySpark, DBT, and Airflow/Dagster/Prefect. MLOps, Databricks and Terraform experience are pluses.

Core skills

PySpark/pandasDBTDatabricks

Required skills

PythonSQLAWSTerraformAirflow/Dagster/Prefect

Optional skills

TerraformAirflow

What you'll do

  • Design, build, and optimize ETL pipelines that power analytics, data science, and ML workflows using tools such as Databricks, PySpark, and Airflow.
  • Develop and maintain labeling and retraining pipelines for machine learning models, ensuring quality, reproducibility, and observability.
  • Implement and support MLOps practices, including model versioning, CI/CD for ML, and model monitoring in production environments.
  • Collaborate with data scientists to productionize and scale model training, inference, and evaluation pipelines.
  • Contribute to the design and evolution of the data lakehouse, including schema design, partitioning strategies, and performance optimization.
  • Document and communicate data architecture, lineage, and dependencies to ensure transparency and maintainability across teams.
  • Champion data quality and governance, ensuring that datasets are accurate, well-structured, and compliant with organizational standards.
  • Leverage infrastructure-as-code and containerization to build reproducible, maintainable environments.
  • Participate in code reviews and continuous improvement of engineering best practices within the team.

What they require

  • Bachelor’s degree in Computer Science, Engineering, or related field.
  • 2–4 years of experience building and operating large-scale data systems, ideally supporting analytics and ML workloads.
  • Proficiency in Python and SQL, with experience in PySpark, pandas, or similar data processing frameworks.
  • Experience with DBT
  • Experience with modern data warehousing and lakehouse platforms, preferably Databricks.
  • Experience with workflow orchestration tools such as Airflow, Dagster, or Prefect.
  • Strong understanding of data modeling, ETL design, and distributed data systems.
  • Experience with AWS data and compute services (S3, Lambda, ECS, CloudWatch, etc.) or equivalent cloud platforms.
  • Familiarity with MLOps concepts (e.g., feature stores, model registries, CI/CD for ML)
  • Experience using Infrastructure as Code, preferably Terraform.
  • Excellent problem-solving, collaboration, and communication skills; comfortable working in a dynamic, fast-paced environment.

Benefits

  • Competitive pay, complete health insurance, 401(k) matching, and an employee equity plan.
  • Flexible time off, paid holidays, paid parental leave, and a paid year-end holiday break.
  • A robust catalog of benefits that support your professional growth and personal wellbeing, including work from home funds, fertility & adoption reimbursement, and more…
  • A varied and challenging role in an innovative, global company.
  • Supportive, driven colleagues who have your back and share your passion.

GoGuardian builds learning solutions for K-12, trusted by school leaders to promote effective teaching and equitable engagement while helping empower educators to keep students safe.

EdTechMid-sizegoguardian.com
$130k–$150k/yr