Skip to main content
Cantina

Machine Learning Engineer - Voice Conversion

RemoteEurope+United States
Published
Role
AI / ML
Employment
Full-time
Company size
Startup
$200k–$220k/yr
Check eligibility

Open to Anywhere in Europe + US. Set where you work from to check your eligibility.

No BS summary

Research/ML engineer for large-scale speech and voice conversion systems. Needs hands-on experience with large audio models, diffusion/flow-matching transformers, audio VAEs/codecs/vocoders, distributed training, and PyTorch. Remote role limited to the U.S. or Europe.

Core skills

PyTorchSpeech modelsVoice conversion

Required skills

Diffusion transformers/Flow-matching transformersAudio VAEsNeural audio codecsVocodersFSDP/DeepSpeedCUDATritonC++

What you'll do

  • Architect, implement, pre-train, fine-tune, and post-train/alignment for large-scale speech models
  • Design, run, and analyze scientific experiments
  • Develop and improve developer tooling
  • Contribute across the stack from low-level optimizations to high-level model design
  • Define data requirements and collaborate on acquisition, curation, augmentation, labeling quality, and synthetic data strategies
  • Design automated objective and subjective evaluations, including listening tests and ASR-based metrics
  • Harden the training, evaluation, and inference pipeline
  • Profile latency, memory, and cost and meet production SLAs
  • Contribute to safety and consent guardrails and abuse mitigation for speech technology

What they require

  • Exceptional research/development experience with large-scale audio models over 8B parameters and over 500k hours of data
  • Deep hands-on experience with diffusion and/or flow-matching transformers, including samplers, schedules, conditioning mechanisms, and distillation
  • Deep hands-on experience training audio VAEs, neural audio codecs, and vocoders, including latent/tokenizer design, reconstruction and perceptual objectives, and adversarial training
  • Strong experience with multi-node, multi-GPU distributed training using FSDP/DeepSpeed or equivalent
  • Strong software engineering skills with a proven track record of building complex systems
  • Strong with PyTorch and performance work, including profiling, CUDA/Triton/C++ as needed, and writing reliable production-quality code
  • Shipped large-scale speech/audio or multimodal generative models to production
  • Background working with large-scale ML data and triangulating quality using subjective and objective signals
  • Experience with voice cloning, speech control/steerability, or expressive speech generation
  • Notable publications and/or open-source contributions in speech/audio/ML

Benefits

  • Competitive salary and company equity
  • Medical, dental, and vision insurance with 99.99% of premiums covered by Cantina
  • 42 days of paid time off
  • 15 PTO days
  • 10 sick days
  • 15 company holidays
  • 2 floating holidays
  • Generous parental leave and fertility support
  • 401(k) retirement savings plan
  • Lifestyle spending account of $500/month
  • Complimentary lunch and snacks for in-office employees
  • One Medical membership

Cantina is a new social platform founded by Sean Parker with an advanced AI character creator. Its bots can interact across voice, video, and text.

Social AiStartup

Details

Apply routeDom
Also posted in 1 other channel
$200k–$220k/yr