Research Engineer ML Platform

Paris-based AI company building frontier models, AI applications, developer tools, and compute infrastructure for enterprise and public-sector deployments.

Paris, France
About Mistral AI

Mistral AI develops open-weight and commercial language models and provides a full-stack AI platform spanning agents, application development, custom-model training, APIs, and AI cloud infrastructure.

View jobs by Mistral AI

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will build and operate the ML platform for large-scale training, evaluation, fine-tuning, and batch inference. You will develop platform services and APIs, orchestrate GPU workloads, manage compute capacity, improve multi-cluster execution and reliability, and participate in on-call troubleshooting.

Requirements

  • 4+ years of ML infrastructure, distributed systems, Kubernetes platform engineering, or related experience
  • Python or Go
  • Kubernetes controllers, operators, CRDs, scheduling, networking, storage, and resource management
  • Kueue, Karpenter, Volcano, and Kyverno
  • Distributed ML workloads
  • GPU infrastructure
  • PyTorch
  • CUDA
  • NCCL
  • High-performance networking
  • Performance and reliability diagnosis

Responsibilities

  • Develop ML platform services, APIs, controllers, and tooling
  • Build GPU workload scheduling and admission-control systems
  • Manage heterogeneous GPU compute capacity
  • Enable multi-cluster workload execution
  • Create self-service workflows for distributed workloads
  • Optimize infrastructure performance and GPU utilization
  • Develop observability, recovery, and operational tooling
  • Participate in on-call rotations and troubleshoot infrastructure issues

Benefits

  • Healthcare coverage
  • Parental leave
  • Retirement plans
  • Relocation support
  • Wellness programs
  • Meal allowances
  • Transportation allowances