Helix AI Engineer Video Pretraining
Figure is an AI robotics company developing general-purpose humanoid robots and its Helix vision-language-action AI system.
Funding history
About Figure
Figure develops and deploys general-purpose humanoid robots for commercial and household tasks. Its current Figure 03 robot is powered by Helix, an onboard generalist vision-language-action model; the company also operates the Index data-collection service.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will design and train video foundation models using internet-scale and robot-collected data. You will develop video pretraining strategies, scalable data pipelines, distributed training systems, and evaluation frameworks for perception, prediction, and control applications.
Requirements
- Experience training large-scale models on video data or high-dimensional sequential modalities
- Understanding of deep learning architectures for video, vision, or multimodal systems
- Experience with large-scale pretraining, dataset curation, training dynamics, and scaling laws
- Proficiency in Python and PyTorch
- Experience with distributed training systems and large GPU clusters
- Experimental rigor in model design and training iteration
- Software engineering skills for scalable, reliable systems
Responsibilities
- Design and train large-scale video foundation models
- Develop pretraining strategies for temporal dynamics, motion, and object interaction
- Build transferable representations for perception, tracking, prediction, and control
- Explore transformer-based and diffusion-based video architectures
- Implement high-throughput video data pipelines and distributed training strategies
- Optimize model compute, memory, and training efficiency
- Design evaluation frameworks and benchmarks for temporal understanding, prediction quality, and generalization
