Staff Software Engineer Inference Compute Infrastructure Engineering
Together AIVisit Together AI website
Together AI operates an AI-native cloud platform for open and custom AI models.
San Francisco, United States
About Together AI
Together AI provides production AI infrastructure spanning inference, accelerated compute, model training and fine-tuning, and secure code sandboxes for AI development.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will build a Kubernetes-native control plane that provisions and manages GPU inference fleets. You will develop declarative APIs, controllers, reconciliation loops, health and remediation systems, and scheduling improvements that make capacity self-service, reliable, and efficient.
Requirements
- Software engineering experience with Go, Python, Rust, or a similar language
- Experience with durable workflow orchestration tools such as Temporal or Cadence
- Experience building software control planes or orchestration systems that reconcile state over time
- Experience building event-driven systems using message queues, event streams, or pub/sub
- Experience building internal platforms or APIs for engineering teams
- Product mindset focused on developer experience
Responsibilities
- Build provisioning state machines for physical host lifecycles
- Design declarative self-service APIs and a control plane for inference clusters
- Automate detection, draining, repair, replacement, and reintroduction of failed nodes
- Ensure pipeline reliability through idempotency, retries, rollback, and drift detection
- Partner with the inference and ML platform team to encode cluster requirements into platform abstractions
- Develop infrastructure software with typing, tests, code review, versioning, and CI/CD
