Python Engineer — ML Tooling & Annotation Platform
- Role
- AI / ML
- Experience
- Senior
- Employment
- Full-time
The listing doesn't say where it hires from. Check the description or the employer's site before applying.
No BS summary
Build and extend our annotation platform: task routing, consensus review, model-assisted pre-labeling with YOLO inference in the loop. Own dataset infrastructure — versioned datasets, class taxonomy management, stratified splits, and the pipelines that feed training runs.
Core skills
Required skills
The role Model quality scales with data quality, and data quality scales with tooling. Our internal annotation and dataset platform is how millions of aerial frames become training data: labeling interfaces, QA queues, automated pre-labeling with existing models, dataset versioning, and the pipelines that feed training runs. As the ML team grows, this platform is the multiplier — and it needs an owner-builder. What you'll do Build and extend our annotation platform: task routing, consensus review, model-assisted pre-labeling with YOLO inference in the loop. Own dataset infrastructure — versioned datasets, class taxonomy management, stratified splits, and the lineage from raw capture to trained model. Write the glue that makes training reproducible: experiment tracking, config management, GPU job submission to our own clusters. Ship clean, tested Python services and CLIs that ML engineers and annotation teams rely on daily. What you'll bring 4+ years of production Python — services, APIs, and data pipelines, not just scripts. Experience supporting ML workflows (annotation systems, dataset management, or training infrastructure). Pragmatism about tooling: you build what the team needs, measure whether it's used, and delete what isn't.
What you'll do
- Build and extend our annotation platform: task routing, consensus review, model-assisted pre-labeling with YOLO inference in the loop.
- Own dataset infrastructure — versioned datasets, class taxonomy management, stratified splits, and the pipelines that feed training runs.
- Write the glue that makes training reproducible: experiment tracking, config management, GPU job submission to our own clusters.
- Ship clean, tested Python services and CLIs that ML engineers and annotation teams rely on daily.
What they require
- 4+ years of production Python — services, APIs, and data pipelines, not just scripts.
- Experience supporting ML workflows (annotation systems, dataset management, or training infrastructure).
- Pragmatism about tooling: you build what the team needs, measure whether it's used, and delete what isn't.
RAAD designs and builds its own servers, including GPU processing nodes, storage arrays, and edge appliances, and deploys them across its own racks, public cloud, and hardened on-prem installations for clients in energy, defense-adjacent, and critical infrastructure environments.