AI Infrastructure Systems Engineer
Together AI operates an AI-native cloud platform for open and custom AI models.
About Together AI
Together AI provides production AI infrastructure spanning inference, accelerated compute, model training and fine-tuning, and secure code sandboxes for AI development.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will build software that automates the full GPU-cluster lifecycle, including provisioning, validation, deployment, upgrades, repairs, and retirement. You will develop infrastructure agents and monitoring platforms to detect failures, support incident triage, and improve GPU availability, utilization, performance, and reliability. You will also create validation systems and internal developer tools for software-managed infrastructure.
Requirements
- 3+ years building distributed systems, infrastructure platforms, or large-scale backend software
- Software engineering skills in Python, Go, or Rust
- Experience building platforms, automation systems, or developer infrastructure
- Experience with Linux, Kubernetes, Terraform, Ansible, or similar infrastructure technologies
- Ability to understand problems across hardware and software
- Automation-first approach to repeated tasks
Responsibilities
- Design and build fleet automation systems for GPU cluster provisioning, validation, deployment, upgrades, repair, and retirement
- Build AI infrastructure agents for deployment automation, root-cause analysis, incident triage, and autonomous remediation
- Develop fleet intelligence platforms to monitor hardware health, firmware, networking, storage, thermals, and workload performance
- Build software that improves GPU availability, utilization, performance, and reliability
- Create automated validation systems for GPUs, networking fabrics, storage, and distributed AI workloads
- Build internal platforms and developer tools for software-managed infrastructure
- Improve deployment velocity, reliability, and operational efficiency through automation
- Partner with hardware, networking, platform, and AI teams on infrastructure improvements
Benefits
- Startup equity
- Health insurance
