Senior Backend Engineer Inference Platform
Together AIVisit Together AI website
Together AI operates an AI-native cloud platform for open and custom AI models.
San Francisco, United States
About Together AI
Together AI provides production AI infrastructure spanning inference, accelerated compute, model training and fine-tuning, and secure code sandboxes for AI development.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will build low-latency routing, load-balancing, auto-scaling, and traffic-shaping systems for inference workloads. You will optimize caching and system performance, productionize new model architectures, and work with researchers on scalable model serving.
Requirements
- Distributed systems
- API microservices
- Fault tolerance
- Scalability
- System stability
- Operating systems
- Multithreading
- Memory management
- Networking
- Storage performance
- Rust
- Go
- Python
- TypeScript
Responsibilities
- Build and optimize global and local request routing
- Ensure low-latency load balancing across data centers and model engine pods
- Develop auto-scaling systems to meet SLOs
- Design multi-tenant traffic-shaping, resource-allocation, and rate-limiting systems
- Engineer latency and throughput trade-offs
- Optimize prefix caching
- Bring new model architectures into production at scale
- Profile system performance, identify bottlenecks, and implement optimizations
Benefits
- Equity
- Health insurance
