Machine Learning Engineer, Speech - Joint Audio-Video Modeling

Cantina

Remote (U.S. or Europe)RemoteFullTimePosted 2w ago
Apply on company site

This role is aggregated from the employer's public careers feed - the button opens their application page.

About Cantina: Cantina Labs is a social AI company, developing a suite of advanced real-time models that push the boundaries of expression, personality, and realism. We bring characters to life, transforming how people tell stories, connect, and create. We build and power ecosystems. Cantina, our flagship social AI platform, is just the beginning. If you're excited about the potential AI has to shape human creativity and social interactions, join us in building the future!   About the Role: We're looking for a Research / ML Engineer to join our Speech Team to build state-of-the-art speech and audio generation systems end-to-end from data specs through production inference with a focus on joint audio-video modeling. You'll own the audio side of multimodal generation: the representations (audio VAEs, neural codecs), the generative backbone (diffusion / flow-matching transformers), and the

Stop searching - get matched.

Upload your CV and our AI will surface roles like this one, with the reasons they fit you.

Upload CV - it's free