Likeremote

Subscribe to the latest remote jobs:

  • Likeremote jobs on https://LinkedIn.com/
  • Likeremote jobs on https://telegram.org/
  • Likeremote jobs on Reddit.com

AI Researcher (Multimodal Audio/Video Generation)

Tavus
🇺🇸 United States
On-site
Staff / Principal
4 months ago
  • AI
  • PyTorch
  • Machine Learning
Not scoredNo CV on file. Upload one and this job gets a score out of 100.Upload CV

About Us

Tavus is a research lab pioneering human computing. We’re building AI Humans: a new interface that closes the gap between people and machines, free from the friction of today’s systems. Our real-time human simulation models let machines see, hear, respond, and even look real—enabling meaningful, face-to-face conversations. AI Humans combine the emotional intelligence of humans with the reach and reliability of machines, making them capable, trusted agents available 24/7, in every language, on our terms.

Imagine a therapist anyone can afford. A personal trainer that adapts to your schedule. A fleet of medical assistants that can give every patient the attention they need. With Tavus, individuals, enterprises, and developers can all build AI Humans to connect, understand, and act with empathy at scale.

We’re a Series A company backed by world-class investors includingSequoia Capital, Y Combinator, and Scale Venture Partners.

Be part of shaping a future where humans and machines truly understand each other.

The Role


We’re hiring aSenior AI Researcher to lead research inaudio-visual avatar generation. This role is for someone who thrives in ambiguity, has a track record of pushing generative models to new frontiers, and wants to define what human–AI interaction looks like in practice.

Your Mission 🚀

  • Lead research efforts onaudio-visual generation for avatars (Neural Avatars, Talking-Heads), with a focus on conversational settings.

  • Design models that are coupled withconversation flow — capturing and generating verbal + non-verbal signals in sync.

  • Drive innovation indiffusion models, long-video generation, and audio-visual modeling.

  • Translate research into production by partnering with Applied ML and engineering.

  • Mentor researchers, set research directions, and publish impactful work.

You’ll Bring:

  • A PhD or equivalent research experience, plus2–3+ years of hands-on experience applying generative models at scale.

  • Expertise indiffusion models and awareness of the latest efficiency techniques.

  • Experience inmultimodal generation — spanning video, audio, and language.

  • Proven innovation inlong-video generation and/oraudio generation.

  • Excellent programming skills — fluent inPyTorch and GPU-optimized workflows.

  • Track record of publications in top-tier venues (CVPR, NeurIPS, BMVC, ICASSP, etc.).

  • Experienceleading research activities or mentoring teams.

Nice-to-Haves:

  • Skills in3D graphics, Gaussian splatting, or large-scale training setups.

  • Broad exposure to generative AI models beyond your specialty.

  • Familiarity with software development best practices.

Location:


Preferred:San Francisco (hybrid) orLondon.

Remote withinU.S. orEurope considered for exceptional candidates.

AI Researcher (Multimodal Audio/Video Generation) · Tavus

Auto apply with Likeremote