ML Runtime Optimization Engineer
Applied Intuition is a physical-AI company that provides software platforms for developing, validating, deploying, and operating intelligent vehicles and machines.
About Applied Intuition
Founded in 2017, Applied Intuition builds physical-AI tooling and infrastructure, a Vehicle OS, and a Self-Driving System for automotive, defense, trucking, construction, mining, agriculture, and robotics applications.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will optimize machine-learning performance across embedded runtime environments. You will develop compute strategies for efficient, low-latency inference, prune and quantize models, profile target platforms, identify bottlenecks, and collaborate on efficient model architectures and deployments.
Requirements
- Bachelor’s degree in Electrical Engineering or Computer Science, or a BSc in Computer Science, Mathematics, Physics, or a related field
- 3+ years of experience with ML accelerators, GPU, CPU, SoC architecture, and micro-architecture
- Strong software development skills focused on embedded programming
- Experience profiling and optimizing model performance on embedded compute platforms
- Experience with deep-learning frameworks such as PyTorch, JAX, or ONNX
Responsibilities
- Drive ML performance optimization for ADAS and autonomous-driving stacks on embedded compute platforms
- Develop compute usage strategies to optimize model-inference efficiency and latency
- Prune and quantize models for memory-constrained platforms
- Collaborate on efficient model architecture solutions
- Profile model performance on embedded platforms and identify bottlenecks
Benefits
- Health, dental, vision, life, and disability insurance coverage
- 401k retirement benefits with employer match
- Learning and wellness stipends
- Paid time off
- Equity in the form of options and/or restricted stock units
