Skip to main content
Torc Robotics

Senior, Software Engineer - ML Data Delivery

RemoteUnited States only
Published
Role
Data Engineering
Experience
Senior
Employment
Full-time
$160.8k–$193k/yr
Check eligibility

Open to US only. Set where you work from to check your eligibility.

No BS summary

Senior software engineer for ML data delivery and pseudo-labeling data pipelines, with Python, MLOps/data processing tooling, and 6+ years with a BS or 3+ years with an MS. Must be US-based/eligible for a Remote - US role; autonomous driving/offline perception experience is important.

Core skills

ParquetPythonMLOps

Required skills

MLflowWeights and BiasesPyArrowDaftPandasVDICIGitHub ActionsDockerPyTorch/Lightning/Ray

What you'll do

  • Design, implement, test and deploy tooling and pipelines for internal quality control and issue identification of pseudo-labeled data, ensuring annotations meet the bar required by downstream model training.
  • Provide statistical support and report on pseudo-label quality, coverage, and pipeline health to internal stakeholders and leadership.
  • Support secondary data selection (virtual packaging) to curate and feed on-demand data to online model training.
  • Support the build-out of the ML data delivery system that enables the online perception team to train models on-demand.
  • Drive general pipeline improvement and optimization across the pseudo-labeling and data delivery stack, identifying and resolving bottlenecks in throughput, quality, or reliability.
  • Demonstrate project management skills, serving as project lead guiding less experienced team members in multiple facets of project execution.
  • Stay up to date with the latest developments in offline perception, data pipeline engineering, and ML data infrastructure for autonomous driving.
  • Independently develop tools, services, and algorithms using disciplined software development processes, making recommendations for developing new code or re-using existing code, implementing version control, and maintaining documentation of created applications.
  • Define and implement ingestion, data preparation, curation, and governance of large, multi-faceted data sets supporting analytics and ML training workflows.
  • Proactively assess current capabilities to identify areas for improvement, proposing solutions that align with core strategy and operation.
  • Guide and produce information products, supporting visualization and data accessibility in a customer-centric manner.
  • Evaluate and make recommendations regarding technical advances that improve productivity and quality, reduce flow times, and enhance operational surety.
  • Develop guidelines and standards for data quality control, data delivery systems, and their deployment, and associated processes.
  • Provide technical guidance or business process expertise, technical leadership, coaching and mentoring to team members.

What they require

  • Considered highly skilled and proficient in discipline; conducts complex, important work under minimal supervision and with wide latitude for independent judgment.
  • Scope of Influence: Expected to drive alignment across team interfaces to the rest of the organization.
  • Designs, maintains and owns team technical solutions and drives consensus.
  • Mentors and guides engineers within the group.
  • Bachelor’s Degree in Computer Science, Robotics, Electrical Engineering or related technical field plus demonstrated competencies and technical proficiencies typically acquired through 6+ years of experience OR; Master’s Degree in Computer Science, Robotics, Electrical Engineering or related technical field plus demonstrated competencies and technical proficiencies typically acquired through 3+ years of experience OR;
  • Familiarity with the offline perception stack in general, and knowledge of how pseudo-label data is produced and related best practices.
  • Strong software engineering background building and operating data pipelines and services at scale.
  • Statistical analysis and reporting skills, with the ability to translate data quality findings into actionable insights.
  • Scaled ML Operations (MLOps) and Tooling – ML Frameworks, experiment tracking, model registry, MLflow, Weights and Biases, ML Metrics and Evaluation / Quality.
  • Model Data Curation – Parquet data processing (PyArrow, Daft, Pandas, etc).
  • Development Tools & Eco-System (at scale) – Proficiency in Python software development. Also, VDI and cloud-based development environments, CI Systems (GitHub Actions), and Docker.
  • Experience with distributed data processing and/or ML frameworks – PyTorch, Lightning, Ray, or similar.
  • Preferred: Experience with large-scale data delivery systems and associated quality control.
  • Preferred: Pseudo-labeling experience in general.

Benefits

  • A competitive compensation package that includes a bonus component and stock options
  • 100% paid medical, dental, and vision premiums for full-time employees
  • 401K plan with a 6% employer match
  • Flexibility in schedule and generous paid vacation (available immediately after start date)
  • Company-wide holiday office closures
  • AD+D and Life Insurance
  • Torc's total compensation package will also include our corporate bonus and stock option plan.
  • Dependent on the position offered, sign-on payments, relocation, and other forms of compensation may be provided as part of a total compensation package, in addition to a full range of medical, financial, and/or other benefits.

Torc develops autonomous driving software for automated trucks. A leader in autonomous driving since 2007, it is now part of the Daimler family and partners directly with a truck manufacturer.

Autonomous VehiclesMid-sizetorc.ai
$160.8k–$193k/yr