Site Reliability Engineer
Paris-based AI company building frontier models, AI applications, developer tools, and compute infrastructure for enterprise and public-sector deployments.
About Mistral AI
Mistral AI develops open-weight and commercial language models and provides a full-stack AI platform spanning agents, application development, custom-model training, APIs, and AI cloud infrastructure.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will design and operate scalable, fault-tolerant infrastructure for web services and machine-learning workloads. You will improve monitoring, alerting, incident response, CI/CD, orchestration, and automation; support production operations and on-call response; collaborate on reproducible training environments; and document operational processes.
Requirements
- Master’s degree in Computer Science, Engineering, or a related field.
- 7+ years of experience in a DevOps or SRE role.
- Expertise in cloud computing and distributed systems.
- Experience with root cause analysis, production troubleshooting, and on-call rotations.
- Experience with observability, alerting, and SLAs.
- Experience with CI/CD, Docker, and Kubernetes.
- Knowledge of Prometheus, Grafana, ELK Stack, Datadog, monitoring, logging, alerting, and observability.
- Familiarity with Terraform or CloudFormation.
- Proficiency in Python, Go, Bash, and software development best practices.
- Knowledge of networking, security, and system administration.
Responsibilities
- Design, build, and maintain scalable, highly available, fault-tolerant infrastructure.
- Ensure availability for platform, inference, and model-training environments across HPC clusters.
- Operate production systems and troubleshoot incidents.
- Implement monitoring, alerting, and incident-response systems.
- Develop and maintain CI/CD, containerization, orchestration, monitoring, and logging workflows.
- Participate in on-call rotations and perform root cause analysis.
- Improve infrastructure automation, deployment, and orchestration using Kubernetes, Flux, and Terraform.
- Enable safe and reproducible model-training experiments.
- Build cloud-agnostic platform abstractions.
- Develop workflows, tooling, and automation to improve reliability, availability, and performance.
- Ensure infrastructure meets security and compliance requirements.
- Document processes and procedures.
Benefits
- Healthcare coverage
- Parental leave
- Retirement plans
- Relocation support
- Wellness programs
- Meal allowances
- Transportation allowances
