Systems Research Engineer Intern GPU Programming Winter 2027

Together AI operates an AI-native cloud platform for open and custom AI models.

San Francisco, United States
About Together AI

Together AI provides production AI infrastructure spanning inference, accelerated compute, model training and fine-tuning, and secure code sandboxes for AI development.

View jobs by Together AI

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will develop and optimize GPU-accelerated kernels and algorithms for machine-learning and AI applications. You will co-design GPU kernels and model architecture, support efficient GPU architecture and programming models, integrate accelerated solutions into software, and stay current with GPU programming techniques. This internship runs from January through April on site in San Francisco.

Requirements

  • GPU programming and parallel computing expertise
  • CUDA or Triton knowledge
  • Machine learning and AI application knowledge
  • GPU performance profiling and optimization tool knowledge
  • Problem-solving and analytical skills

Responsibilities

  • Optimize and fine-tune GPU code for performance and scalability
  • Collaborate with cross-functional teams to integrate GPU-accelerated solutions into software systems
  • Stay current with GPU programming techniques and technologies

Benefits

  • Housing stipend
Systems Research Engineer Intern GPU Programming Winter 2027 at Together AI | JobStash