Junior / Senior / Staff Software Engineer, Inference / Compute Infrastructure Engineering
Together AIVisit Together AI website
Together AI operates an AI-native cloud platform for open and custom AI models.
San Francisco, United States
About Together AI
Together AI provides production AI infrastructure spanning inference, accelerated compute, model training and fine-tuning, and secure code sandboxes for AI development.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will build Kubernetes-native control planes, declarative APIs, and workflow systems for GPU inference infrastructure. You will automate provisioning, self-healing, lifecycle management, capacity optimization, and reliability while operating the systems you ship in production.
Requirements
- Software engineering experience with Go, Python, Rust, or similar
- Experience with durable workflow orchestration tools
- Experience building control planes or orchestration systems
- Experience with event-driven systems
- Experience building internal platforms or APIs
Responsibilities
- Build provisioning state machines for physical-host lifecycles
- Design declarative self-service APIs and control planes
- Automate node health detection, remediation, and capacity recovery
- Ensure idempotency, retries, rollback, and drift detection
- Encode inference cluster topology and scheduling requirements
- Build typed, tested, versioned infrastructure software with CI/CD
Benefits
- Startup equity
- Health insurance
