Senior Backend Engineer Inference Platform

Together AI operates an AI-native cloud platform for open and custom AI models.

San Francisco, United States
About Together AI

Together AI provides production AI infrastructure spanning inference, accelerated compute, model training and fine-tuning, and secure code sandboxes for AI development.

View jobs by Together AI

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will build low-latency routing, load-balancing, auto-scaling, and traffic-shaping systems for inference workloads. You will optimize caching and system performance, productionize new model architectures, and work with researchers on scalable model serving.

Requirements

  • Distributed systems
  • API microservices
  • Fault tolerance
  • Scalability
  • System stability
  • Operating systems
  • Multithreading
  • Memory management
  • Networking
  • Storage performance
  • Rust
  • Go
  • Python
  • TypeScript

Responsibilities

  • Build and optimize global and local request routing
  • Ensure low-latency load balancing across data centers and model engine pods
  • Develop auto-scaling systems to meet SLOs
  • Design multi-tenant traffic-shaping, resource-allocation, and rate-limiting systems
  • Engineer latency and throughput trade-offs
  • Optimize prefix caching
  • Bring new model architectures into production at scale
  • Profile system performance, identify bottlenecks, and implement optimizations

Benefits

  • Equity
  • Health insurance
Senior Backend Engineer Inference Platform at Together AI | JobStash