Conversational Modelling Research Engineer
TavusVisit Tavus website
AI video generation platform (digital avatars and dubbing).
About Tavus
Tavus (tavus.io) is an AI video generation platform building digital avatars, dubbing and video-personalization APIs.
Skills
About the Role
Research foundation multimodal conversational models for conversational avatars, develop real-time low-latency methods for verbal and non-verbal conversation and avatar control, experiment with fine-tuning and conditioning techniques, and move prototypes into production.
Requirements
- PhD or near completion in a relevant field, or equivalent research experience.
- Hands-on experience with large multimodal models and generative language models.
- Experience with visual question answering, audio and video understanding, captioning, behavioural analysis, translation, or speech-to-speech systems.
- Experience fine-tuning or adapting vision-language models for control, conditioning, or downstream tasks.
- Background in deep learning and foundation models.
- Strong PyTorch skills and experience building deep learning pipelines.
Responsibilities
- Conduct research on large multimodal models for conversational avatars.
- Develop low-latency methods to model verbal and non-verbal conversation and control avatar behaviour in real time.
- Experiment with fine-tuning, adaptation, and conditioning techniques for audiovisual multimodal models.
- Partner with the Applied ML team to take research prototypes into production.
- Stay current with cutting-edge advancements and help define future research directions.
Benefits
- Remote work within the United States or Europe for exceptional candidates.
