Machine Learning Engineer Inference

Together AI operates an AI-native cloud platform for open and custom AI models.

San Francisco, United States
About Together AI

Together AI provides production AI infrastructure spanning inference, accelerated compute, model training and fine-tuning, and secure code sandboxes for AI development.

View jobs by Together AI

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will design and build production systems for AI inference, optimize runtime inference services, and create supporting services and tools. You will work with researchers and product partners on new capabilities, review designs and code, document systems, and implement fault-tolerant data ingestion and processing.

Requirements

  • 3+ years writing high-performance, well-tested, production-quality code
  • Python proficiency
  • PyTorch proficiency
  • Experience building high-performance libraries and tooling
  • Understanding of multithreading, memory management, networking, storage, performance, and scale
  • Knowledge of AI inference systems such as TGI, vLLM, TensorRT-LLM, or Optimum preferred
  • Knowledge of speculative decoding preferred
  • CUDA or Triton programming knowledge preferred
  • Rust, Cython, and compiler knowledge are nice to have

Responsibilities

  • Design and build production systems for the inference engine
  • Develop and optimize runtime inference services for large-scale AI applications
  • Collaborate on new features and research capabilities
  • Conduct design and code reviews
  • Create services, tools, and developer documentation
  • Implement robust, fault-tolerant systems for data ingestion and processing

Benefits

  • Startup equity
  • Health insurance
Machine Learning Engineer Inference at Together AI | JobStash