Site Reliability Engineer

Paris-based AI company building frontier models, AI applications, developer tools, and compute infrastructure for enterprise and public-sector deployments.

Paris, France
About Mistral AI

Mistral AI develops open-weight and commercial language models and provides a full-stack AI platform spanning agents, application development, custom-model training, APIs, and AI cloud infrastructure.

View jobs by Mistral AI

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will design and operate scalable, fault-tolerant infrastructure for web services and machine-learning workloads. You will improve monitoring, alerting, incident response, CI/CD, orchestration, and automation; support production operations and on-call response; collaborate on reproducible training environments; and document operational processes.

Requirements

  • Master’s degree in Computer Science, Engineering, or a related field.
  • 7+ years of experience in a DevOps or SRE role.
  • Experience with cloud computing and highly available distributed systems.
  • Experience with production troubleshooting, root cause analysis, and on-call rotations.
  • Experience with observability, alerting, and SLAs.
  • Experience with CI/CD, Docker, and Kubernetes.
  • Knowledge of Prometheus, Grafana, ELK Stack, Datadog, monitoring, logging, alerting, and observability.
  • Familiarity with Terraform or CloudFormation.
  • Proficiency in Python, Go, Bash, and software development best practices.
  • Knowledge of networking, security, and system administration.

Responsibilities

  • Design, build, and maintain scalable, highly available, fault-tolerant infrastructure.
  • Maintain availability for platform, inference, and model-training environments across HPC clusters.
  • Operate production systems and troubleshoot incidents.
  • Implement monitoring, alerting, and incident-response systems.
  • Build and maintain CI/CD, containerization, orchestration, monitoring, logging, and alerting workflows.
  • Participate in on-call rotations and perform root cause analysis.
  • Improve infrastructure automation, deployment, and orchestration using Kubernetes, Flux, and Terraform.
  • Enable safe and reproducible model-training experiments.
  • Build cloud-agnostic platform abstractions.
  • Develop workflows and tooling to improve reliability, availability, and performance.
  • Ensure infrastructure meets security and compliance requirements.
  • Document processes and procedures.
  • Contribute to open-source projects, research publications, blog articles, and conferences.

Benefits

  • Healthcare coverage
  • Parental leave
  • Retirement plans
  • Relocation support
  • Wellness programs
  • Meal allowances
  • Transportation allowances