Senior AI Researcher Multimodal Audio Video Generation
Tavus is an AI research lab and operating company focused on Human Computing, offering real-time human-like AI PALs, multimodal models, developer APIs, and a no-code PAL Maker.
About Tavus
Tavus builds foundational AI models for perception, conversational understanding, voice, and human rendering. Its current offerings include the Conversational Video Interface developer API, PAL Maker for no-code AI-human creation and deployment, and managed enterprise PAL solutions for use cases such as recruiting, healthcare, sales, education, and onboarding.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
Lead research on audio-visual generation for conversational avatars, designing synchronized verbal and non-verbal signal models, advancing diffusion and long-video methods, partnering with Applied ML and engineering, mentoring researchers, setting research direction, and publishing impactful work.
Requirements
- PhD or equivalent research experience
- 2–3+ years of hands-on experience applying generative models at scale
- Expertise in diffusion models and awareness of efficiency techniques
- Experience in multimodal generation across video, audio, and language
- Proven innovation in long-video generation and/or audio generation
- Excellent programming skills in PyTorch and GPU-optimized workflows
- Track record of publications in top-tier venues
- Experience leading research activities or mentoring teams
Responsibilities
- Lead research on audio-visual generation for conversational avatars
- Design models coupled with conversation flow to generate synchronized verbal and non-verbal signals
- Drive innovation in diffusion models, long-video generation, and audio-visual modeling
- Translate research into production with Applied ML and engineering
- Mentor researchers, set research directions, and publish impactful work
