Member of Technical Staff - Pre-Training Infra
Reflection is an AI research lab building open frontier models and a full AI stack for developers, enterprises, and public-sector users.
Funding history
About Reflection
Reflection develops open-weight AI models, open-source software for customizing and running agents, AI-factory infrastructure, and related solutions. Its current research emphasizes large language models, reinforcement learning, and agentic reasoning.
Skills
About the Role
You will build distributed systems for frontier-model pre-training and operate large-scale training runs. You will optimize throughput, stability, memory, communication, and GPU utilization; maintain pipelines for datasets and checkpointing; work with researchers to productionize experiments; and debug bottlenecks across training stacks and runtimes.
Requirements
- Distributed training
- Machine learning
- Megatron
- DeepSpeed
- Model parallelism
- GPU optimization
- NCCL
- GPU communication
- Debugging
- Data pipeline
- Foundation model
Responsibilities
- Build and scale distributed training systems
- Design and operate large-scale foundation model training runs
- Develop infrastructure for training across thousands of GPUs
- Optimize training throughput stability and efficiency
- Productionize experimental training workflows with researchers
- Improve communication memory usage and GPU utilization
- Build and maintain training pipelines for datasets checkpointing and experiments
- Debug distributed training and GPU performance bottlenecks
Benefits
- Stock options
- Medical insurance
- Dental insurance
- Vision insurance
- Life insurance
- Annual wellness allowance
- Daily office lunch and dinner
- 22 weeks of paid parental leave
- Unlimited paid time off in the U.S.
- 30 days of vacation in the U.K.
- Visa sponsorship
- Regular off-sites
- Happy hours
- Team celebrations
