Machine Learning Engineer - Voice Conversion

Cantina · Remote (U.S. or Europe) · $200k–$220k

Posted
19 days ago
Last confirmed live
1 day ago
Published range
$200k–$220k

What this role involves

Cantina is seeking a Research/ML Engineer to join their Speech Team, focusing on building state-of-the-art speech systems end-to-end, including voice conversion and controllable TTS. The role involves model building, experimental design, tool development, and full-stack contribution, with a strong emphasis on large-scale audio models and production deployment. The position offers a salary range of $200,000-$220,000.

Skills this posting asks for

  • pytorch
  • diffusion transformers
  • flow-matching transformers
  • audio vae
  • neural audio codecs
  • vocoders
  • distributed training
  • fsdp
  • deepspeed
  • cuda
  • triton
  • c++
  • voice cloning
  • speech control
  • expressive speech generation
  • grpo
  • dpo
  • asr
  • wer
  • sv

Requirements

  • Level: senior

From the employer’s posting

About Cantina: Cantina is a new social platform founded by Sean Parker with the most advanced AI character creator. Our bots are lifelike, social creatures that can interact wherever people are online—across voice, video, and text. Create yourself, imagine someone new, or choose from thousands of…

Read the full description on Cantina’s careers page

Apply without filling the form

Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.

Other roles at Cantina

All 8 roles at Cantina