Research Engineer, Inference Foundation

Paris-based AI company building frontier models, AI applications, developer tools, and compute infrastructure for enterprise and public-sector deployments.

Series DRecently funded0 current maintainers0 active leadsTeam intelligence

Maintainer signals as of 9/25/2026

Paris, France
About Mistral AI

Mistral AI develops open-weight and commercial language models and provides a full-stack AI platform spanning agents, application development, custom-model training, APIs, and AI cloud infrastructure.

View jobs by Mistral AI

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will develop, tune, and operate the core inference engine and orchestrator for large-scale LLM serving. You will improve throughput, latency, release reliability, fleet efficiency, capacity elasticity, topology, and infrastructure that supports reinforcement-learning and post-training workloads. You will debug across the systems stack and contribute upstream improvements when appropriate.

Requirements

  • Experience building and operating ML or LLM services at scale with latency and availability targets
  • Experience with inference engines such as vLLM, SGLang, or TensorRT-LLM
  • Knowledge of inference internals including prefill, decode, KV cache, batching, scheduling, speculative decoding, and parallelism strategies
  • Familiarity with distributed and disaggregated serving architectures
  • Ability to debug CUDA, NCCL, kernels, containers, networking, and storage
  • Python for systems tooling and backend services
  • PyTorch
  • Kubernetes
  • GPU and networking fundamentals, including CUDA runtime, NCCL, and InfiniBand/RDMA

Responsibilities

  • Develop and fix the inference engine and orchestrator
  • Select, configure, and tune serving features for performance at scale
  • Own validated serving-stack releases through automated performance gates and progressive rollout
  • Drive appropriate improvements and fixes upstream to open-source inference engines
  • Optimize fleet serving efficiency, pod startup time, caching, and offloading
  • Maintain optimal serving topology, placement, connectivity, and routing
  • Build serving infrastructure for reinforcement-learning and post-training workloads
  • Optimize inference performance across workloads

Benefits

  • Healthcare coverage
  • Parental leave
  • Retirement plans
  • Relocation support
  • Wellness programs
  • Meal allowances
  • Transportation allowances
  • Location-specific perks