Machine Learning Engineer, Speech - Joint Audio-Video Modeling

Cantina · Remote (U.S. or Europe)

Posted
24 days ago
Last confirmed live
1 day ago

What this role involves

Cantina Labs is seeking a Research/ML Engineer to join their Speech Team, focusing on building state-of-the-art speech and audio generation systems with joint audio-video modeling. The role involves owning the audio side of multimodal generation, including audio representations, generative backbones, and conditioning for characters to speak, sing, and emote in sync with video. Responsibilities include designing and training audio VAEs, neural codecs, diffusion/flow-matching transformers, and collaborating on data, evaluation, and inference efficiency.

Skills this posting asks for

  • audio vae
  • neural codecs
  • vocoders
  • diffusion transformers
  • flow-matching transformers
  • tts
  • voice conversion
  • voice cloning
  • multi-speaker conditioning
  • grpo
  • dpo
  • distillation
  • quantization
  • kernel optimization
  • memory optimization
  • distributed training
  • inference optimization
  • evaluation design
  • data curation
  • synthetic data
  • audio fidelity metrics
  • av-sync
  • red-team studies

From the employer’s posting

About Cantina: Cantina Labs is a social AI company, developing a suite of advanced real-time models that push the boundaries of expression, personality, and realism. We bring characters to life, transforming how people tell stories, connect, and create. We build and power ecosystems. Cantina, our…

Read the full description on Cantina’s careers page

Apply without filling the form

Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.

Other roles at Cantina

All 8 roles at Cantina