Staff Machine Learning Engineer Voice AI

Together AI operates an AI-native cloud platform for open and custom AI models.

San Francisco, United States
About Together AI

Together AI provides production AI infrastructure spanning inference, accelerated compute, model training and fine-tuning, and secure code sandboxes for AI development.

View jobs by Together AI

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will own the voice model-serving roadmap for speech-to-text, text-to-speech, and speech-to-speech workloads. You will optimize inference performance, GPU utilization, batching, streaming pipelines, memory management, reliability, and latency. You will build evaluation frameworks, enable emerging audio model architectures, lead model-partner integrations, diagnose performance bottlenecks, influence serving architecture, and define voice fine-tuning capabilities.

Requirements

  • 8+ years of machine learning engineering experience
  • Production-scale model serving, inference optimization, or machine learning infrastructure experience
  • Experience with vLLM, SGLang, TensorRT-LLM, or equivalent serving engines
  • Python and PyTorch proficiency
  • GPU optimization knowledge, including CUDA kernels, memory hierarchies, and profiling toolchains
  • System design experience
  • Technical leadership
  • Developer tooling product intuition
  • Speech and audio machine learning knowledge
  • ASR and TTS architecture knowledge
  • Audio signal processing
  • Audio codec and tokenization familiarity, including SNAC, Encodec, and DAC
  • Speech-model training or fine-tuning experience

Responsibilities

  • Own the voice inference roadmap for speech-to-text, text-to-speech, and speech-to-speech models
  • Architect and implement voice inference systems for latency, throughput, and GPU utilization
  • Design serving architectures for serverless and dedicated endpoints
  • Build streaming inference pipelines, batching strategies, and memory management
  • Build voice-model evaluation frameworks and benchmarks
  • Enable support for emerging audio model architectures
  • Lead model-partner integrations from implementation through optimization
  • Profile, diagnose, and resolve performance bottlenecks
  • Influence platform serving architecture for real-time voice APIs
  • Define voice-model fine-tuning capabilities

Benefits

  • Startup equity
  • Health insurance
Staff Machine Learning Engineer Voice AI at Together AI | JobStash