Research Engineer Large Scale Training

Together AI operates an AI-native cloud platform for open and custom AI models.

San Francisco, United States
About Together AI

Together AI provides production AI infrastructure spanning inference, accelerated compute, model training and fine-tuning, and secure code sandboxes for AI development.

View jobs by Together AI

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will build, optimize, and maintain large-scale training infrastructure. You will integrate and validate new model architectures, profile distributed workloads, eliminate performance bottlenecks, run experiments, and productionize training methods with reliable, scalable experimental systems.

Requirements

  • Python
  • PyTorch
  • Large neural network training
  • Multi-GPU training
  • Multi-node training
  • ML systems
  • GPU architecture
  • Mixed-precision training
  • Distributed training
  • Data parallelism
  • Tensor parallelism
  • Pipeline parallelism
  • Expert parallelism
  • Communication
  • AI research

Responsibilities

  • Design, implement, and optimize core large-scale training infrastructure components
  • Integrate model architectures and optimize production fine-tuning workloads
  • Profile distributed training workloads and eliminate compute, memory, and communication bottlenecks
  • Design and execute performance experiments and benchmarks
  • Productionize novel training methods
  • Enable support for newly released open-source foundation models
  • Build and maintain reliable, scalable experimental infrastructure

Benefits

  • Startup equity
  • Health insurance
Research Engineer Large Scale Training at Together AI | JobStash