Senior / Lead Machine Learning Engineer, Serving

Inworld AI develops AI products for growing applications, helping developers go from prototype to production faster. Their offerings include advanced text-to-speech (TTS) technology and an upcoming Runtime product, aimed at enhancing consumer applications with expressive, real-time voice AI.

Seed6 current maintainers5 active leads8 lead step-downsTeam intelligence

Maintainer signals as of 8/23/2026

Distributed
About Inworld AI

Inworld develops AI products for consumer applications. They offer a text-to-speech model that aims for high quality with better pricing, lower latency, more control, local serving options, and open training code. They also have a product called Inworld Runtime, which is currently in private preview. Their services are used by partners like XBOX, Ubisoft, NVIDIA, and Meta. They focus on helping developers go from prototype to production faster and increase experimentation velocity to deploy new AI improvements daily.

View jobs by Inworld AI

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

Work on optimizing realtime inference for state-of-the-art voice models, taking models from research into production by containerizing, optimizing, and operating them reliably. Build distributed systems for high-throughput inference and treat performance, latency, and reliability as first-class product features.

Requirements

  • Deep understanding of modern serving frameworks and inference optimization techniques.
  • Hands-on experience with quantization, distillation, caching, continuous batching, paged attention, and speculative decoding.
  • Proficiency in C++, CUDA, Rust, or highly optimized Python.
  • Experience with Kubernetes, Ray, custom load balancing, and multi-GPU or multi-node inference.
  • Non-trivial systems programming projects, open-source contributions, or deep-dive technical write-ups.
  • Full-cycle experience taking models from research to production.
  • PhD in computer science, physics, or mathematics, or equivalent practical experience building backend or ML systems.
  • Professional fluency in written and spoken English.

Responsibilities

  • Optimize realtime inference using serving frameworks such as vLLM and TRT-LLM.
  • Apply quantization, distillation, caching, continuous batching, paged attention, and speculative decoding.
  • Profile and optimize NVIDIA GPU performance using C++, CUDA, Rust, or optimized Python.
  • Build and scale distributed systems with Kubernetes, Ray, and custom load balancing.
  • Support reliable multi-GPU and multi-node inference for thousands of concurrent connections.
  • Take models from research through containerization, serving optimization, and reliable production operation.
  • Design benchmarks and prototypes to validate architectural decisions.
  • Collaborate daily with US-based leadership and engineering teams.

Benefits

  • Full U.S. visa and relocation support may be available for candidates interested in relocating to the San Francisco Bay Area, subject to business needs and applicable legal and work authorization requirements.
Senior / Lead Machine Learning Engineer, Serving at Inworld AI | JobStash