Senior Machine Learning Engineer Voice AI
Together AIVisit Together AI website
Together AI operates an AI-native cloud platform for open and custom AI models.
San Francisco, United States
About Together AI
Together AI provides production AI infrastructure spanning inference, accelerated compute, model training and fine-tuning, and secure code sandboxes for AI development.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will own and optimize the voice-model serving stack across speech-to-text, text-to-speech, and speech-to-speech workloads. You will improve GPU performance, streaming inference, batching, evaluation, model integrations, and fine-tuning capabilities for real-time voice APIs.
Requirements
- ML engineering
- Model serving
- Inference optimization
- ML infrastructure
- LLM serving
- vLLM
- SGLang
- TensorRT-LLM
- Python
- PyTorch
- GPU profiling
- CUDA
- Memory management
- Kernel debugging
- Production ML systems
Responsibilities
- Optimize inference performance for voice models
- Productionize voice models on serverless and dedicated endpoints
- Design batching strategies, streaming inference, and memory management for audio workloads
- Build and maintain voice-model evaluation frameworks
- Enable new model architectures in the serving stack
- Integrate and optimize partner models
- Profile and debug the full inference stack and ship performance improvements
- Ensure the serving layer meets real-time voice API latency and reliability requirements
- Contribute to voice-model fine-tuning capabilities
Benefits
- Startup equity
- Health insurance
