Staff / Principal Machine Learning Engineer, Serving

Inworld AI develops AI products for growing applications, helping developers go from prototype to production faster. Their offerings include advanced text-to-speech (TTS) technology and an upcoming Runtime product, aimed at enhancing consumer applications with expressive, real-time voice AI.

Seed6 current maintainers5 active leads8 lead step-downsTeam intelligence

Maintainer signals as of 8/23/2026

Distributed
About Inworld AI

Inworld develops AI products for consumer applications. They offer a text-to-speech model that aims for high quality with better pricing, lower latency, more control, local serving options, and open training code. They also have a product called Inworld Runtime, which is currently in private preview. Their services are used by partners like XBOX, Ubisoft, NVIDIA, and Meta. They focus on helping developers go from prototype to production faster and increase experimentation velocity to deploy new AI improvements daily.

View jobs by Inworld AI

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

Optimize sub-second multimodal inference for realtime voice models, taking models from research into production through containerization, serving optimization, and reliable operation. Apply model acceleration techniques and build high-performance distributed systems capable of handling thousands of concurrent connections.

Requirements

  • Deep understanding of modern serving frameworks such as vLLM or TRT-LLM.
  • Hands-on experience with model acceleration techniques.
  • Proficiency in C++, CUDA, Rust, or highly optimized Python.
  • Experience profiling code and optimizing NVIDIA GPU performance.
  • Experience with Kubernetes, Ray, custom load balancing, and multi-GPU or multi-node inference.
  • Experience handling thousands of concurrent connections reliably.
  • Non-trivial systems programming projects, open-source contributions, or deep-dive technical write-ups.
  • Ability to take models from research to production.
  • PhD in CS, Physics, Math, or equivalent practical experience building backend or ML systems.

Responsibilities

  • Optimize realtime inference serving with frameworks such as vLLM and TRT-LLM.
  • Apply quantization, distillation, caching, continuous batching, paged attention, and speculative decoding.
  • Build high-performance systems in C++, CUDA, Rust, or optimized Python and profile GPU performance.
  • Design and scale distributed systems using Kubernetes, Ray, custom load balancing, and multi-GPU or multi-node inference.
  • Handle thousands of concurrent connections reliably.
  • Take models from research to production, including containerization and serving optimization.
  • Contribute to open-source projects and publish technical write-ups.

Benefits

  • Relocation assistance
  • Bonus
  • Equity
  • Benefits