ML Researcher - Image / Video Diffusion
Krea is an active generative-AI creative platform for generating, editing, and enhancing images, video, and 3D content.
San Francisco, United States
Funding history
About Krea
Krea provides browser-based AI creative tools and a developer API. Its offerings include image and video generation, enhancement and editing tools, 3D generation, and Krea 2, its in-house image foundation model.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will train image and video diffusion models on large GPU clusters. You will optimize distributed training, improve model quality and reliability, debug training failures, and evaluate architecture, data, optimizer, and algorithm choices.
Requirements
- Experience working with image or video models at scale
- PyTorch
- Distributed training including FSDP, CP, SP, USP, TP, and EP
- Profiling and debugging large distributed training
- Low-precision training and inference in FP8, NVFP4, and MXFP8
- Knowledge of diffusion-model training pipelines
- Knowledge of LLM, VLM, representation learning, and robotics research
- Experience designing custom data pipelines
Responsibilities
- Train diffusion models for image and video generation on large GPU clusters
- Optimize and profile distributed training runs
- Implement distributed training strategies including FSDP, CP, SP, TP, and EP
- Improve model quality and reliability through data, architecture, pipelines, experiments, and evaluations
- Debug distributed-training errors and implement fault tolerance
- Evaluate architecture, attention, optimizer, data, and algorithm choices
Benefits
- Equity packages
- Health insurance
- Dental and vision insurance
- Health FSA accounts
- Long-term disability coverage
- Flexible PTO
- 401k with a 4% company-sponsored match
- Office meals
- Uber transit to and from the office
- Visa sponsorship
