Helix AI Engineer Training Performance
Figure is an AI robotics company developing general-purpose humanoid robots and its Helix vision-language-action AI system.
Funding history
About Figure
Figure develops and deploys general-purpose humanoid robots for commercial and household tasks. Its current Figure 03 robot is powered by Helix, an onboard generalist vision-language-action model; the company also operates the Index data-collection service.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will optimize large-scale model training across GPU clusters, develop custom kernels, improve data pipelines and fault tolerance, and build performance-monitoring tooling. You will evaluate accelerator architectures and tune distributed parallelism for efficient training.
Requirements
- Bachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or a related field
- 3+ years in AI performance engineering
- Understanding of GPU architecture and performance characteristics
- Proficiency with Nsight Systems, Nsight Compute, PyTorch Profiler, HTA, or similar profiling tools
- Knowledge of NCCL, RDMA, NVLink, InfiniBand, RoCE, and topology-aware placement
- Python and CUDA/C++ skills
- Experience debugging performance regressions and instability at scale
- Experience using MFU or HFU hardware-efficiency metrics
Responsibilities
- Optimize training performance for 100B+ parameter models across 100k+ GPUs
- Inform accelerator, cluster topology, scheduling, and hardware procurement decisions
- Write and optimize Triton and CUDA kernels
- Build tooling and dashboards for performance monitoring, regression detection, and root-cause analysis
- Optimize data loading and preprocessing pipelines
- Improve checkpointing, fault tolerance, and elastic restart
- Co-design performant model architectures and training recipes
- Contribute to kernel compilers
- Build agentic systems for generating and benchmarking custom kernels
- Evaluate accelerator architectures and lead proof-of-concept ports and benchmarks
- Explore model and data parallelism configurations
