Senior Site Reliability Engineer AI Infrastructure
Wayve is a London-headquartered embodied-AI company developing and licensing mapless, vehicle-agnostic driving software for assisted, automated, and robotaxi applications.
About Wayve
Wayve Technologies Ltd. develops the Wayve AI Driver, an end-to-end, data-trained software platform that runs on onboard vehicle compute and native sensors. It is designed for OEM integration across L1 driver assistance through L4 automated driving, without HD maps.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will establish reliability practices for model-development and GPU-compute platforms. You will operate cloud clusters, define service objectives, participate in on-call response, improve observability and deployment safety, automate operations and recovery, and strengthen infrastructure-as-code guardrails.
Requirements
- Experience in SRE, Production Engineering, or Cloud Reliability for large-scale cloud systems
- Kubernetes production-cluster experience
- Experience running production workloads in AWS, GCP, or Azure
- Experience operating complex distributed systems
- Experience with large compute clusters
- Linux fundamentals
- Proficiency in Python, Go, C++, or another scripting or systems language
- Troubleshooting skills across networking, storage, distributed systems, and performance
- Experience with Datadog, Prometheus, Grafana, or OpenTelemetry
- Incident leadership, postmortem writing, and reliability influence skills
- Experience with GPU environments, ML infrastructure, MLOps, Terraform, secure cloud environments, and reliability programs is desirable
Responsibilities
- Own reliability, availability, and performance for model-development and GPU-compute environments
- Define and operationalize SLOs, SLIs, and error budgets
- Improve capacity planning, scaling, and resource efficiency for GPU-backed clusters
- Establish production-readiness standards with engineering partners
- Participate in a 24/7 on-call rotation
- Lead incident triage, escalation, communication, and root-cause analysis
- Turn post-incident learning into architectural and automation improvements
- Design and operate monitoring, logging, tracing, alerting, and health dashboards
- Improve deployment safety, change management, validation, and rollback
- Build automation for cluster operations, workflows, remediation, and scaling
- Implement self-healing and recovery workflows
- Harden CI/CD, infrastructure-as-code, and policy-driven guardrails
Benefits
- Hybrid working policy
- Work-from-home time
