Member of Technical Staff VLM

Frontier AI research lab building the FLUX family of multimodal visual-intelligence models and delivering them through an API, playground, and open weights.

Freiburg im Breisgau, Germany
About Black Forest Labs

Black Forest Labs develops generative AI models for image, video, audio, and action prediction, alongside production API and enterprise deployment offerings.

View jobs by Black Forest Labs

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will lead development and training of multimodal vision-language models. You will design fine-tuning strategies for creative use cases, research VLM and LLM integrations with diffusion and flow pipelines, and evaluate emerging multimodal architectures to turn research into practical improvements.

Requirements

  • Experience pretraining or significantly advancing a deployed or publicly released vision-language model
  • Record of advancing multimodal architectures through publications or production work
  • Knowledge of tokenization, alignment, grounding, cross-modal attention, and failure modes
  • Experience with multi-node distributed training
  • Experience with diffusion or flow-based generative models is preferred

Responsibilities

  • Lead development and training of multimodal vision-language models
  • Design fine-tuning strategies for specialized creative use cases
  • Research VLM and LLM integrations with diffusion and flow pipelines
  • Evaluate emerging multimodal architectures and implement practical improvements

Benefits

  • Equity
  • Reasonable travel costs covered
Member of Technical Staff VLM at Black Forest Labs | JobStash