Staff Research Engineer Multimodal Generative Modelling

Synthesia is an active London-based enterprise AI video platform that lets businesses create, localize, manage, and publish videos using AI avatars and voiceovers.

London, United Kingdom
About Synthesia

Synthesia Limited provides browser-based AI video creation for business communications, training, sales enablement, marketing, and support. Its platform includes AI-assisted creation, avatars, voiceovers, translation/localization, collaboration, and publishing workflows.

View jobs by Synthesia

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will define research directions and build interactive multimodal systems that combine text, voice, and video. You will develop, evaluate, train, and deploy conversational models, improve their responsiveness and emotional expressiveness, curate data, and optimize production runtime.

Requirements

  • Generative modeling expertise, ideally with sequential or multimodal data
  • Experience with large language models or transformer-based architectures
  • PyTorch proficiency
  • Distributed training and model optimization
  • Time-series modeling and tokenization knowledge
  • Experience training deep learning models end to end
  • Software engineering skills

Responsibilities

  • Shape the roadmap for new multimodal model capabilities
  • Propose multimodal system architectures for text and voice
  • Develop and evaluate low-latency streaming conversational systems
  • Design solutions for emotional expressiveness and natural interaction
  • Implement models from pretraining through post-training
  • Integrate and test neural codecs, diffusion, and flow-matching architectures
  • Define evaluation metrics for conversational systems
  • Curate datasets
  • Lead post-training initiatives including DPO, fine-tuning, and distillation
  • Deploy optimized models to production and address customer feedback