Senior / Lead Machine Learning Engineer, Serving

Inworld AI develops AI products for growing applications, helping developers go from prototype to production faster. Their offerings include advanced text-to-speech (TTS) technology and an upcoming Runtime product, aimed at enhancing consumer applications with expressive, real-time voice AI.

Seed6 current maintainers5 active leads8 lead step-downsTeam intelligence

Maintainer signals as of 8/23/2026

Distributed
About Inworld AI

Inworld develops AI products for consumer applications. They offer a text-to-speech model that aims for high quality with better pricing, lower latency, more control, local serving options, and open training code. They also have a product called Inworld Runtime, which is currently in private preview. Their services are used by partners like XBOX, Ubisoft, NVIDIA, and Meta. They focus on helping developers go from prototype to production faster and increase experimentation velocity to deploy new AI improvements daily.

View jobs by Inworld AI

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

Work on optimizing realtime inference and serving of state-of-the-art voice models at massive scale. Take models from research, containerize and optimize their serving, ensure reliable production operation, and improve latency, throughput, and performance across distributed NVIDIA GPU systems.

Requirements

  • Deep understanding of modern serving frameworks such as vLLM or TRT-LLM
  • Hands-on experience with quantization, distillation, caching strategies, continuous batching, paged attention, and speculative decoding
  • Proficiency in C++, CUDA, Rust, or highly optimized Python
  • Experience with Kubernetes, Ray, custom load balancing, and multi-GPU or multi-node inference
  • Experience reliably handling thousands of concurrent connections
  • Non-trivial systems programming projects, major inference-engine open-source contributions, or deep-dive technical write-ups
  • Ability to take a model from research to production, containerize it, and optimize its serving
  • PhD in CS, Physics, or Math, or equivalent practical experience building backend or ML systems
  • Professional fluency in written and spoken English

Responsibilities

  • Optimize realtime inference and model serving at scale
  • Take models from the research team, containerize them, and optimize their serving
  • Ensure models run reliably in production
  • Apply model acceleration techniques including quantization, distillation, caching, continuous batching, paged attention, and speculative decoding
  • Profile code and optimize performance on NVIDIA GPUs
  • Handle distributed systems and scaling using Kubernetes, Ray, and custom load balancing
  • Manage multi-GPU and multi-node inference while handling thousands of concurrent connections
  • Collaborate with US-based leadership and engineering teams

Benefits

  • Full U.S. visa and relocation support may be available for candidates interested in relocating to the San Francisco Bay Area, subject to business needs and applicable legal and work authorization requirements