Member of Technical Staff ML Performance
OdysseyVisit Odyssey website
Odyssey is an AI lab building general-purpose world models that understand and simulate physical and virtual environments.
Palo Alto, United States
Funding history
About Odyssey
Odyssey Systems, Inc. develops foundation world models for applications including robotics, autonomous driving, AI training, drones, and gaming. Its current public work includes Odyssey-3, Starchild-1, Agora-1, and PROWL-1.
Skills
About the Role
You will optimize real-time ML models and improve distributed training and inference performance. You will design training strategies for large GPU clusters, develop tools to diagnose bottlenecks and stability issues, and shape performant model architectures and infrastructure.
Requirements
- 8+ years of software engineering experience with significant ML performance work
- Deep understanding of modern machine learning architectures and performance optimization
- Experience with distributed training and inference
- Track record of owning projects end to end
- Proficiency with PyTorch or TensorFlow or JAX
- Proficiency with Triton
- Knowledge of NVIDIA GPU ecosystems and optimization stacks
- Metric-based approach
Responsibilities
- Optimize models for real-time use at scale
- Design and implement distributed training strategies
- Partner on performant model architectures
- Develop tools to identify performance bottlenecks and stability issues
- Pioneer performance frameworks and system designs
- Make autonomous technical decisions
- Use latest-generation GPUs
