Lead/Manager AI Infrastructure Systems Engineering Team

Together AI operates an AI-native cloud platform for open and custom AI models.

San Francisco, United States
About Together AI

Together AI provides production AI infrastructure spanning inference, accelerated compute, model training and fine-tuning, and secure code sandboxes for AI development.

View jobs by Together AI

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will lead, coach, and manage an SRE team while operating user-facing production systems. You will build infrastructure and monitoring, respond to incidents, improve reliability and performance, debug production issues, and plan infrastructure growth.

Requirements

  • 7+ years of professional SRE or related experience
  • Ideally 2 years as a Lead SRE
  • Expert knowledge of Ansible, Terraform, and Kubernetes
  • Programming or scripting proficiency
  • Monitoring and observability experience
  • Cloud services knowledge

Responsibilities

  • Participate in an on-call rotation and respond to availability incidents
  • Manage, develop, and coach the SRE team
  • Build and operate infrastructure with Ansible, Terraform, and Kubernetes
  • Build monitoring systems
  • Design and implement deployment and upgrade processes
  • Debug production issues across the stack
  • Improve architecture for reliability, performance, and availability
  • Plan infrastructure growth
Lead/Manager AI Infrastructure Systems Engineering Team at Together AI | JobStash