Software Engineer, ML Platform
- Role
- AI / ML
Open to AT only. Set where you work from to check your eligibility.
No BS summary
Software engineer for ML platform work: Python production systems, PyTorch training, cloud/containers/orchestration, distributed computing, data pipelines, and large versioned datasets. Needs to work directly with researchers on reproducible experiments, evaluation, inference, releases, and secure handling of sensitive clinical speech data.
Core skills
Required skills
Optional skills
About the role As a Software Engineer on the ML Platform, you will build the systems behind every nyra labs experiment and model release. You will work across data processing, distributed training, experiment management, evaluation, inference, and release infrastructure. Your goal is to give a small research team the leverage to run ambitious experiments quickly, reproducibly, and reliably. This is not a conventional backend role. You will work directly with researchers, understand how models are developed, and turn recurring research bottlenecks into dependable platform capabilities. Why we need you Strong research depends on more than strong ideas. Training data must be versioned and traceable. Experiments need to be reproducible. Evaluations must run consistently. Models need to move from a researcher’s environment into efficient inference and public releases without fragile manual steps. The sensitivity and scale of clinical speech data add another challenge: the platform must enable fast research while maintaining strict standards for security, privacy, and data governance. We need an engineer who sees infrastructure as a force multiplier for research. About the company At nyra health, we build software that supports clinics, therapists, and patients throughout neurorehabilitation. myReha delivers personalized therapy, while nyra insights helps clinical teams manage and understand patient progress. nyra labs is the research arm of nyra health. We turn difficult problems encountered in practice into open models, datasets, benchmarks, and research that the wider community can build on. If that resonates with you, we would love to hear from you. What you’ll shape Data platform: Build reliable pipelines for ingesting, validating, transforming, versioning, and accessing large speech datasets. Training infrastructure: Improve distributed training, orchestration, checkpointing, resource scheduling, and failure recovery. Experiment systems: Create tooling for configuration, tracking, comparison, reproducibility, and artifact management. Evaluation platform: Make it easy to run benchmarks, inspect regressions, compare releases, and understand model behavior. Inference: Optimize models for efficient cloud and on-device use where relevant. Release infrastructure: Automate model packaging, documentation, validation, and open-source publishing. Developer experience: Build internal tools that remove friction from the daily work of researchers and engineers. Reliability and security: Establish observability, access controls, and operational practices appropriate for sensitive clinical data. What sets you up for success Strong software engineering: Excellent Python skills and experience building maintainable production systems. ML systems experience: Familiarity with PyTorch training, model evaluation, GPU workloads, and the ML development lifecycle. Distributed systems: Experience with cloud infrastructure, containers, orchestration, job scheduling, or distributed computing. Data engineering: Experience building reliable pipelines and working with large, versioned datasets. Operational mindset: You care about observability, debuggability, failure recovery, and clear system boundaries. Platform thinking: You build reusable capabilities instead of solving the same problem repeatedly. Research empathy: You understand that research workflows change quickly and infrastructure must support exploration. AI-native workflow: You use coding agents, automation, and custom tooling to increase your own leverage and that of the team. Beyond your CV Pragmatic: You know what needs a platform and what needs a small script. Leverage-oriented: You look for improvements that make the entire team faster. Reliable: You treat reproducibility and data integrity as core product requirements. Self-directed: You can identify bottlenecks and own the solution end to end. Collaborative: You enjoy working closely with researchers and translating experimental needs into durable systems. Why nyra labs Build the infrastructure behind open speech models and research Work directly with researchers instead of operating as a separate platform team Solve difficult systems problems involving large-scale audio and sensitive data High ownership with visible impact on every experiment and release Direct collaboration with founders and the wider nyra health engineering team Attractive compensation, Phantom Stock Options, and company benefits A beautiful office in Vienna’s First District with a hybrid working model To apply Please include: Your resume A short description of a technically challenging ML, data, or infrastructure system you built and the tradeoffs you made A link to your GitHub or other relevant work, if available The process Intro call, approximately 30 minutes: Your background, expectations, and an introduction to nyra labs. Technical deep-dive: A discussion of a system you built and a platform-design scenario relevant to ML research. Meet the founders and team: Discuss working style, technical direction, and how you would increase the lab’s research velocity.
What you'll do
- Build the systems behind every nyra labs experiment and model release.
- Work across data processing, distributed training, experiment management, evaluation, inference, and release infrastructure.
- Work directly with researchers, understand how models are developed, and turn recurring research bottlenecks into dependable platform capabilities.
- Build reliable pipelines for ingesting, validating, transforming, versioning, and accessing large speech datasets.
- Improve distributed training, orchestration, checkpointing, resource scheduling, and failure recovery.
- Create tooling for configuration, tracking, comparison, reproducibility, and artifact management.
- Make it easy to run benchmarks, inspect regressions, compare releases, and understand model behavior.
- Optimize models for efficient cloud and on-device use where relevant.
- Automate model packaging, documentation, validation, and open-source publishing.
- Build internal tools that remove friction from the daily work of researchers and engineers.
- Establish observability, access controls, and operational practices appropriate for sensitive clinical data.
What they require
- Excellent Python skills and experience building maintainable production systems.
- Familiarity with PyTorch training, model evaluation, GPU workloads, and the ML development lifecycle.
- Experience with cloud infrastructure, containers, orchestration, job scheduling, or distributed computing.
- Experience building reliable pipelines and working with large, versioned datasets.
- You care about observability, debuggability, failure recovery, and clear system boundaries.
- You build reusable capabilities instead of solving the same problem repeatedly.
- You understand that research workflows change quickly and infrastructure must support exploration.
- You use coding agents, automation, and custom tooling to increase your own leverage and that of the team.
- You know what needs a platform and what needs a small script.
- You look for improvements that make the entire team faster.
- You treat reproducibility and data integrity as core product requirements.
- You can identify bottlenecks and own the solution end to end.
- You enjoy working closely with researchers and translating experimental needs into durable systems.
- Please include: Your resume.
- Please include: A short description of a technically challenging ML, data, or infrastructure system you built and the tradeoffs you made.
- Please include: A link to your GitHub or other relevant work, if available.
Benefits
- Build the infrastructure behind open speech models and research
- Work directly with researchers instead of operating as a separate platform team
- Solve difficult systems problems involving large-scale audio and sensitive data
- High ownership with visible impact on every experiment and release
- Direct collaboration with founders and the wider nyra health engineering team
- Attractive compensation, Phantom Stock Options, and company benefits
- A beautiful office in Vienna’s First District with a hybrid working model
At nyra health, we build software that supports clinics, therapists, and patients throughout neurorehabilitation. myReha delivers personalized therapy, while nyra insights helps clinical teams manage and understand patient progress. nyra labs is the research arm of nyra health. We turn difficult problems encountered in practice into open models, datasets, benchmarks, and research that the wider community can build on.