Staff Software Engineer AI Compute Together Cloud
Together AIVisit Together AI website
Together AI operates an AI-native cloud platform for open and custom AI models.
San Francisco, United States
About Together AI
Together AI provides production AI infrastructure spanning inference, accelerated compute, model training and fine-tuning, and secure code sandboxes for AI development.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will set the architecture and technical direction for GPU and network virtualization, in-data-center infrastructure, and global management systems. You will build reliable distributed control planes, automate fault remediation, lead cross-team design decisions, mentor engineers, and create testing and developer documentation.
Requirements
- 7+ years of professional software development experience
- Expert-level proficiency in at least one backend language
- Experience owning the architecture of large distributed systems from conception to production at scale
- Deep experience building and operating globally distributed high-performance microservice architectures across cloud providers
- Expert systems knowledge across compute, networking, and storage
- Demonstrated technical leadership through mentoring, design reviews, and cross-team alignment
- Excellent communication and diplomacy skills
- Experience building reliable customer-facing production systems and infrastructure automation, observability, and CI/CD
Responsibilities
- Own the GPU and network virtualization stack
- Own and build the in-data-center IaaS layer
- Design GPU scheduling and the global management plane
- Architect monitoring and automated remediation for fault tolerance
- Set technical direction across teams
- Mentor senior and junior engineers
- Create testing frameworks, tools, and developer documentation
Benefits
- Startup equity
- Health insurance
- Remote work flexibility
