Member of Technical Staff VLM
Black Forest LabsVisit Black Forest Labs website
Frontier AI research lab building the FLUX family of multimodal visual-intelligence models and delivering them through an API, playground, and open weights.
Freiburg im Breisgau, Germany
Funding history
About Black Forest Labs
Black Forest Labs develops generative AI models for image, video, audio, and action prediction, alongside production API and enterprise deployment offerings.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will lead development and training of multimodal vision-language models. You will design fine-tuning strategies for creative use cases, research VLM and LLM integrations with diffusion and flow pipelines, and evaluate emerging multimodal architectures to turn research into practical improvements.
Requirements
- Experience pretraining or significantly advancing a deployed or publicly released vision-language model
- Record of advancing multimodal architectures through publications or production work
- Knowledge of tokenization, alignment, grounding, cross-modal attention, and failure modes
- Experience with multi-node distributed training
- Experience with diffusion or flow-based generative models is preferred
Responsibilities
- Lead development and training of multimodal vision-language models
- Design fine-tuning strategies for specialized creative use cases
- Research VLM and LLM integrations with diffusion and flow pipelines
- Evaluate emerging multimodal architectures and implement practical improvements
Benefits
- Equity
- Reasonable travel costs covered
