Research Engineer Large Scale Training
Together AIVisit Together AI website
Together AI operates an AI-native cloud platform for open and custom AI models.
San Francisco, United States
About Together AI
Together AI provides production AI infrastructure spanning inference, accelerated compute, model training and fine-tuning, and secure code sandboxes for AI development.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will build, optimize, and maintain large-scale training infrastructure. You will integrate and validate new model architectures, profile distributed workloads, eliminate performance bottlenecks, run experiments, and productionize training methods with reliable, scalable experimental systems.
Requirements
- Python
- PyTorch
- Large neural network training
- Multi-GPU training
- Multi-node training
- ML systems
- GPU architecture
- Mixed-precision training
- Distributed training
- Data parallelism
- Tensor parallelism
- Pipeline parallelism
- Expert parallelism
- Communication
- AI research
Responsibilities
- Design, implement, and optimize core large-scale training infrastructure components
- Integrate model architectures and optimize production fine-tuning workloads
- Profile distributed training workloads and eliminate compute, memory, and communication bottlenecks
- Design and execute performance experiments and benchmarks
- Productionize novel training methods
- Enable support for newly released open-source foundation models
- Build and maintain reliable, scalable experimental infrastructure
Benefits
- Startup equity
- Health insurance
