Research Engineer AI RL Infrastructure

Applied Intuition is a physical-AI company that provides software platforms for developing, validating, deploying, and operating intelligent vehicles and machines.

Sunnyvale, United States
About Applied Intuition

Founded in 2017, Applied Intuition builds physical-AI tooling and infrastructure, a Vehicle OS, and a Self-Driving System for automotive, defense, trucking, construction, mining, agriculture, and robotics applications.

View jobs by Applied Intuition

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will design and operate large-scale ML infrastructure for training, evaluation, data, and deployment. You will orchestrate GPU clusters, build benchmarking and data-curation systems, enable distributed training, diagnose systems issues, and translate AI research into production-ready systems.

Requirements

  • Have experience building and operating production-grade systems across ML training, evaluation, data, and deployment
  • Have experience with performance engineering and compute acceleration for large-scale ML training
  • Have systems-level debugging skills for large-scale distributed training
  • Be familiar with open-source ML and systems ecosystems
  • Have technical experience with PyTorch, CUDA, Ray, Flyte, and Kubernetes
  • Hold a Master's or PhD in machine learning and computer vision with autonomy and robotics applications or a closely related field
  • Have 1+ year of experience in autonomy and robotics applications or a closely related field

Responsibilities

  • Design and build training and evaluation infrastructure for AI research
  • Orchestrate GPU clusters to process multimodal sensor data
  • Build benchmarking, continuous evaluation, and regression tracking systems
  • Develop data sampling, dataset generation, and data-curation pipelines
  • Enable distributed training across cloud environments
  • Diagnose and resolve distributed-training issues across model code, data pipelines, runtimes, and cluster infrastructure
  • Collaborate with research, autonomy, and platform teams to productionize research

Benefits

  • Equity
  • Health insurance
  • Dental insurance
  • Vision insurance
  • Life insurance
  • Disability insurance
  • 401k retirement benefits with employer match
  • Learning stipends
  • Wellness stipends
  • Paid time off