Staff SRE AI Infrastructure

Wayve is a London-headquartered embodied-AI company developing and licensing mapless, vehicle-agnostic driving software for assisted, automated, and robotaxi applications.

London, United Kingdom
About Wayve

Wayve Technologies Ltd. develops the Wayve AI Driver, an end-to-end, data-trained software platform that runs on onboard vehicle compute and native sensors. It is designed for OEM integration across L1 driver assistance through L4 automated driving, without HD maps.

View jobs by Wayve

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will establish reliability foundations for cloud platforms and GPU compute environments. You will own availability and performance, define service-level objectives, lead incident response, build observability and automation, and improve deployment, recovery, and infrastructure operations.

Requirements

  • Experience in SRE, production engineering, or cloud reliability roles supporting large-scale cloud systems
  • Experience operating GPU-backed environments or large-scale machine-learning infrastructure
  • Experience running production model-training or inference pipelines
  • Strong Kubernetes experience operating production clusters
  • Experience running production workloads in AWS, GCP, or Azure
  • Experience operating complex distributed systems in production
  • Linux fundamentals and proficiency in Python, Go, C++, or another scripting or systems language
  • Troubleshooting skills across networking, storage, distributed systems, and performance
  • Experience designing and operating observability stacks such as Datadog, Prometheus, Grafana, or OpenTelemetry
  • Experience with Terraform and secure cloud production environments
  • Experience defining SLOs and SLIs
  • Experience establishing SRE processes

Responsibilities

  • Own the reliability, availability, and performance of model-development and GPU compute platforms
  • Define and operationalize SLOs, SLIs, and error budgets
  • Improve capacity planning, scaling strategies, and GPU resource efficiency
  • Participate in a 24/7 on-call rotation
  • Lead incident triage, escalation, communications, and root-cause analysis
  • Design and operate monitoring, logging, tracing, and alerting systems
  • Build automation for cluster operations, training workflows, remediation, and scaling
  • Implement self-healing and resilient recovery workflows
  • Harden CI/CD and release processes
  • Support infrastructure-as-code and policy-driven guardrails

Benefits

  • Hybrid working arrangement