Staff AI Inference and Acceleration Engineer
Figure is an AI robotics company developing general-purpose humanoid robots and its Helix vision-language-action AI system.
Funding history
About Figure
Figure develops and deploys general-purpose humanoid robots for commercial and household tasks. Its current Figure 03 robot is powered by Helix, an onboard generalist vision-language-action model; the company also operates the Index data-collection service.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will own on-board inference architecture and map AI workloads across accelerators. You will optimize toolchains, models, kernels, memory, scheduling, and power use; profile bottlenecks; define compute budgets; evaluate acceleration hardware; and coordinate inference constraints with hardware, software, and AI teams.
Requirements
- Master’s or PhD in Computer Engineering, Electrical Engineering, Computer Science, or a related field, or equivalent industry experience
- At least 8 years of industry experience in hardware acceleration, machine-learning systems, or compute architecture
- Knowledge of AI and machine-learning inference, model formats, inference runtimes, and deployment pipelines
- Experience optimizing models for edge or embedded hardware with quantization, pruning, and operator-level tuning
- Knowledge of computer architecture, memory hierarchies, data movement, and heterogeneous compute
- Experience profiling and benchmarking inference workloads across CPU, GPU, NPU, and DSP
- Familiarity with TVM, MLIR, TensorRT, Torch, SNPE, QNN, JAX, CUDA, or ROCm
- C++ and Python software engineering skills
- Cross-functional communication skills
Responsibilities
- Own on-board inference architecture and map models to NPU, GPU, DSP, and CPU accelerators
- Partition inference workloads across heterogeneous compute resources
- Define and maintain compute budgets for robot inference tasks
- Evaluate acceleration hardware and contribute to future compute-platform requirements
- Optimize inference toolchains from model export through runtime execution
- Apply quantization, pruning, operator fusion, and compression techniques
- Profile inference pipelines and eliminate latency, memory-bandwidth, and power bottlenecks
- Optimize kernel scheduling, memory layout, and data movement
- Define hardware-friendly model constraints with AI and machine-learning teams
- Support runtime integration, scheduling, and power management
- Engage with silicon vendors and research teams on accelerator roadmaps
