Site Reliability Engineering Lead

Graphcore is a SoftBank-owned AI-compute company developing AI processors, systems, and software for machine-learning workloads.

Series ERecently funded0 current maintainers0 active leadsTeam intelligence

Maintainer signals as of 9/25/2026

Bristol, United Kingdom
About Graphcore

Graphcore develops Intelligence Processing Units (IPUs) and the Poplar SDK for building and running machine-learning applications. It continues operating under the Graphcore name as a wholly owned SoftBank Group subsidiary.

View jobs by Graphcore

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will establish and lead an SRE organization through production launch, stabilization, and scale. You will define operating practices, hire and mentor engineers, embed reliability requirements into the platform, lead incident response, and build the automation, observability, and operational discipline needed for reliable production services.

Requirements

  • Site reliability engineering
  • Production engineering
  • Infrastructure reliability
  • Production readiness
  • 24x7x365 operations
  • SLO
  • Incident management
  • Observability
  • Automation
  • Toil reduction
  • Capacity management
  • Linux
  • Networking
  • Storage
  • Distributed system
  • Incident response
  • InfiniBand
  • RDMA
  • RoCE
  • Ethernet
  • Kubernetes
  • Slurm
  • Datacenter operations

Responsibilities

  • Build and develop the SRE organization
  • Define staffing, on-call, escalation, incident management, access, change management, and readiness practices
  • Establish operational interfaces with datacenter operations, engineering teams, vendors, and service owners
  • Develop runbooks, procedures, failure-mode documentation, and incident-response practices
  • Develop training, simulations, and incident-response exercises
  • Embed reliability, serviceability, observability, and operational requirements into platform development
  • Lead development of SLOs, health indicators, alerting standards, severity definitions, and reliability reporting
  • Engineer recurring operational problems out through automation and improved observability
  • Lead or participate in major production incidents
  • Establish blameless post-incident reviews and corrective actions
  • Represent production reliability concerns in engineering and leadership discussions

Benefits

  • Flexible working
  • Medical coverage
  • Dental coverage
  • Vision coverage
  • Flexible Spending Account
  • Health Savings Account
  • Disability insurance
  • Life insurance
  • 401(k) retirement plan
  • Commuter benefits
  • Wellness services
  • Employee Assistance Programme