Staff / Principal Machine Learning Engineer, Serving

Inworld AI develops AI products for growing applications, helping developers go from prototype to production faster. Their offerings include advanced text-to-speech (TTS) technology and an upcoming Runtime product, aimed at enhancing consumer applications with expressive, real-time voice AI.

Seed6 current maintainers5 active leads8 lead step-downsTeam intelligence

Maintainer signals as of 8/23/2026

Distributed
About Inworld AI

Inworld develops AI products for consumer applications. They offer a text-to-speech model that aims for high quality with better pricing, lower latency, more control, local serving options, and open training code. They also have a product called Inworld Runtime, which is currently in private preview. Their services are used by partners like XBOX, Ubisoft, NVIDIA, and Meta. They focus on helping developers go from prototype to production faster and increase experimentation velocity to deploy new AI improvements daily.

View jobs by Inworld AI

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

Take models from the research team, containerize and optimize their serving, and ensure reliable production operation. Work on inference optimization, model acceleration, and high-performance distributed systems powering realtime multimodal AI at scale.

Requirements

  • Deep understanding of modern serving frameworks such as vLLM or TRT-LLM
  • Hands-on experience with quantization, distillation, caching strategies, continuous batching, paged attention, and speculative decoding
  • Proficiency in C++, CUDA, Rust, or highly optimized Python
  • Experience with Kubernetes, Ray, custom load balancing, multi-GPU or multi-node inference, and thousands of concurrent connections
  • Non-trivial systems programming projects, open-source contributions to major inference engines, or deep-dive technical write-ups
  • Full-cycle ownership from research model to production serving
  • PhD in CS, Physics, or Math, or equivalent practical experience building backend or ML systems
  • Legal right to work in the United Kingdom

Responsibilities

  • Optimize realtime inference and serving frameworks
  • Containerize research models and ensure reliable production deployment
  • Apply quantization, distillation, caching, continuous batching, paged attention, and speculative decoding
  • Profile code and optimize performance on NVIDIA GPUs
  • Handle multi-GPU and multi-node inference and thousands of concurrent connections
  • Design benchmarks and prototypes to validate technical decisions
  • Share work and contribute to open-source projects

Benefits

  • Equity
  • Benefits