Staff Software Engineer Inference Compute Infrastructure Engineering

Together AI operates an AI-native cloud platform for open and custom AI models.

San Francisco, United States
About Together AI

Together AI provides production AI infrastructure spanning inference, accelerated compute, model training and fine-tuning, and secure code sandboxes for AI development.

View jobs by Together AI

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will build a Kubernetes-native control plane that provisions and manages GPU inference fleets. You will develop declarative APIs, controllers, reconciliation loops, health and remediation systems, and scheduling improvements that make capacity self-service, reliable, and efficient.

Requirements

  • Software engineering experience with Go, Python, Rust, or a similar language
  • Experience with durable workflow orchestration tools such as Temporal or Cadence
  • Experience building software control planes or orchestration systems that reconcile state over time
  • Experience building event-driven systems using message queues, event streams, or pub/sub
  • Experience building internal platforms or APIs for engineering teams
  • Product mindset focused on developer experience

Responsibilities

  • Build provisioning state machines for physical host lifecycles
  • Design declarative self-service APIs and a control plane for inference clusters
  • Automate detection, draining, repair, replacement, and reintroduction of failed nodes
  • Ensure pipeline reliability through idempotency, retries, rollback, and drift detection
  • Partner with the inference and ML platform team to encode cluster requirements into platform abstractions
  • Develop infrastructure software with typing, tests, code review, versioning, and CI/CD