Skip to main content
Cantina

Machine Learning Engineer, Speech - Joint Audio-Video Modeling

RemoteEurope+United States
Published
Role
AI / ML
Employment
Full-time
$200k–$220k/yr
Check eligibility

Open to Anywhere in Europe + US. Set where you work from to check your eligibility.

No BS summary

ML/research engineer for large-scale speech and audio generation, focused on joint audio-video modeling. Needs deep hands-on experience with diffusion or flow-matching transformers, audio VAEs/neural codecs/vocoders, PyTorch, distributed multi-GPU training, and production speech/audio or multimodal generative models. Remote role hiring in the U.S. or Europe.

Core skills

Diffusion transformers/Flow-matching transformersSpeech generationAudio-video modeling

Required skills

Audio VAEsNeural audio codecsVocodersFSDP/DeepSpeedPyTorchCUDA/Triton/C++

Optional skills

Video diffusion transformersVideo flow-matching transformersVideo VAEsSelf ForcingSelf Forcing++

What you'll do

  • Design, train, and improve audio VAEs, neural codecs, and vocoders for generative models, including latent design, reconstruction and perceptual objectives, and compression-vs-fidelity tradeoffs.
  • Architect, implement, pre-train, fine-tune, and post-train/alignment diffusion and flow-matching transformers for large-scale audio and video generation.
  • Design audio conditioning and cross-modal alignment inside joint audio-video models, including audio latents alongside video latents, reference-audio and multi-speaker conditioning, and multi-shot generation audio/video modeling.
  • Design, run, and analyze scientific experiments to advance understanding of the models.
  • Define data requirements and collaborate on acquisition, curation, AV-sync and quality filtering, annotation quality, and synthetic data strategies for paired audio-video and speech corpora.
  • Design automated objective and subjective evaluations for audio fidelity, intelligibility metrics, AV-sync, listening and viewing tests, robustness and bias checks, and red-team studies.
  • Drive distillation, step-count reduction, quantization, and kernel/memory optimization to meet interactive latency and cost targets.
  • Harden the training-to-evaluation-to-inference pipeline; profile latency, memory, and cost; and meet production SLAs with robust monitoring and rollback.
  • Partner with infrastructure to run distributed training and inference on cloud fleets and productionize models with reliability and observability.
  • Independently lead small research projects while collaborating on larger team initiatives, including cross-team work with video generation.
  • Develop and improve developer tooling to enhance team productivity.
  • Contribute to safety and consent guardrails, watermarking, and misuse/abuse mitigation for responsible voice and likeness technology.

What they require

  • Exceptional research/development experience with large-scale audio models greater than 8B parameters and greater than 500k hours of data.
  • Deep hands-on experience with diffusion and/or flow-matching transformers, including practical knowledge of samplers, schedules, conditioning mechanisms, and distillation.
  • Deep hands-on experience training audio VAEs, neural audio codecs, and vocoders, including latent/tokenizer design, reconstruction and perceptual objectives, and adversarial training.
  • Strong experience with multi-node, multi-GPU distributed training using FSDP/DeepSpeed or equivalent.
  • Strong software engineering skills with a proven track record of building complex systems.
  • Strong with PyTorch and performance work including profiling and CUDA/Triton/C++ as needed, and writing reliable production-quality code.
  • Shipped large-scale speech/audio or multimodal generative models to production.
  • Background working with large-scale ML data and ability to iterate on data and triangulate quality using subjective and objective signals.
  • Experience with voice cloning, speech control/steerability, or expressive speech generation.
  • Notable publications and/or open-source contributions in speech/audio/ML.
  • Sees research and engineering as two sides of the same coin and enjoys owning work end-to-end.
  • Excited to work across modalities and collaborate closely with a video generation team rather than staying inside audio.
  • Results-oriented, flexible, and willing to pick up whatever moves the needle.
  • Likes collaborating closely with infrastructure, data, and product to ship measurable improvements.
  • Enjoys designing experiments, listening tests, and metrics that correlate with user-perceived quality.
  • Eager to learn every day and to find and solve unique large-scale problems.
  • Preferred: Experience with multimodal audio-video modeling, including joint AV generation of multi-shot, multi-speaker scenes with dialogue, music, and sound design generated jointly with video, and cross-modal alignment that keeps them in sync.
  • Preferred: Experience with video generation, including video diffusion/flow-matching transformers, video VAEs, conditioned and multi-shot generation, and building data pipelines for video models.
  • Preferred: Streaming or real-time generation and causal distillation.

Benefits

  • Competitive salary and generous company equity.
  • Medical, dental, and vision insurance with 99.99% of premiums covered by Cantina.
  • 42 days of paid time off.
  • 15 PTO days.
  • 10 sick days.
  • 15 company holidays.
  • 2 floating holidays.
  • Generous parental leave and fertility support.
  • 401(k) retirement savings plan.
  • Lifestyle spending account of $500/month.
  • Complimentary lunch and snacks for in-office employees.
  • One Medical membership.

Cantina is a new social platform founded by Sean Parker with an advanced AI character creator. Its bots can interact across voice, video, and text.

Social AiStartup

Details

Apply routeDom
$200k–$220k/yr