Member of Technical Staff ML Performance

Odyssey is an AI lab building general-purpose world models that understand and simulate physical and virtual environments.

Palo Alto, United States
About Odyssey

Odyssey Systems, Inc. develops foundation world models for applications including robotics, autonomous driving, AI training, drones, and gaming. Its current public work includes Odyssey-3, Starchild-1, Agora-1, and PROWL-1.

View jobs by Odyssey

Skills

About the Role

You will optimize real-time ML models and improve distributed training and inference performance. You will design training strategies for large GPU clusters, develop tools to diagnose bottlenecks and stability issues, and shape performant model architectures and infrastructure.

Requirements

  • 8+ years of software engineering experience with significant ML performance work
  • Deep understanding of modern machine learning architectures and performance optimization
  • Experience with distributed training and inference
  • Track record of owning projects end to end
  • Proficiency with PyTorch or TensorFlow or JAX
  • Proficiency with Triton
  • Knowledge of NVIDIA GPU ecosystems and optimization stacks
  • Metric-based approach

Responsibilities

  • Optimize models for real-time use at scale
  • Design and implement distributed training strategies
  • Partner on performant model architectures
  • Develop tools to identify performance bottlenecks and stability issues
  • Pioneer performance frameworks and system designs
  • Make autonomous technical decisions
  • Use latest-generation GPUs
Member of Technical Staff ML Performance at Odyssey | JobStash