Staff / Principal Machine Learning Engineer, Serving
Inworld AI develops AI products for growing applications, helping developers go from prototype to production faster. Their offerings include advanced text-to-speech (TTS) technology and an upcoming Runtime product, aimed at enhancing consumer applications with expressive, real-time voice AI.
Maintainer signals as of 8/23/2026
Funding history
About Inworld AI
Inworld develops AI products for consumer applications. They offer a text-to-speech model that aims for high quality with better pricing, lower latency, more control, local serving options, and open training code. They also have a product called Inworld Runtime, which is currently in private preview. Their services are used by partners like XBOX, Ubisoft, NVIDIA, and Meta. They focus on helping developers go from prototype to production faster and increase experimentation velocity to deploy new AI improvements daily.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
Own the full lifecycle of research models by containerizing, optimizing, and operating their serving infrastructure in production. Build sub-second multimodal inference systems using model acceleration techniques and distributed infrastructure capable of handling thousands of concurrent connections.
Requirements
- Deep understanding of modern serving frameworks like vLLM or TRT-LLM
- Hands-on experience with quantization, distillation, caching strategies, continuous batching, paged attention, and speculative decoding
- Proficiency in C++, CUDA, Rust, or highly optimized Python
- Experience with Kubernetes, Ray, custom load balancing, and multi-GPU/multi-node inference
- Public work such as open-source contributions or technical write-ups
- Full-cycle ownership experience taking models to production
- PhD in CS, Physics, Math, or equivalent practical experience
- Professional fluency in English
- Legal right to work in Switzerland
Responsibilities
- Take models from the research team to production
- Containerize and optimize model serving
- Optimize inference using quantization, distillation, caching, continuous batching, paged attention, and speculative decoding
- Build distributed serving infrastructure across multiple GPUs and nodes
- Handle thousands of concurrent connections reliably
- Profile code and optimize performance on NVIDIA GPUs
- Design benchmarks and prototypes to resolve open questions
- Contribute to open-source projects and technical write-ups
Benefits
- Remote work
- Full U.S. visa and relocation support may be available for those interested in relocating to the San Francisco Bay Area
