Site Reliability Engineering Lead
Graphcore is a SoftBank-owned AI-compute company developing AI processors, systems, and software for machine-learning workloads.
Maintainer signals as of 9/25/2026
Funding history
About Graphcore
Graphcore develops Intelligence Processing Units (IPUs) and the Poplar SDK for building and running machine-learning applications. It continues operating under the Graphcore name as a wholly owned SoftBank Group subsidiary.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will establish and lead an SRE organization through production launch, stabilization, and scale. You will define operating practices, hire and mentor engineers, embed reliability requirements into the platform, lead incident response, and build the automation, observability, and operational discipline needed for reliable production services.
Requirements
- Site reliability engineering
- Production engineering
- Infrastructure reliability
- Production readiness
- 24x7x365 operations
- SLO
- Incident management
- Observability
- Automation
- Toil reduction
- Capacity management
- Linux
- Networking
- Storage
- Distributed system
- Incident response
- InfiniBand
- RDMA
- RoCE
- Ethernet
- Kubernetes
- Slurm
- Datacenter operations
Responsibilities
- Build and develop the SRE organization
- Define staffing, on-call, escalation, incident management, access, change management, and readiness practices
- Establish operational interfaces with datacenter operations, engineering teams, vendors, and service owners
- Develop runbooks, procedures, failure-mode documentation, and incident-response practices
- Develop training, simulations, and incident-response exercises
- Embed reliability, serviceability, observability, and operational requirements into platform development
- Lead development of SLOs, health indicators, alerting standards, severity definitions, and reliability reporting
- Engineer recurring operational problems out through automation and improved observability
- Lead or participate in major production incidents
- Establish blameless post-incident reviews and corrective actions
- Represent production reliability concerns in engineering and leadership discussions
Benefits
- Flexible working
- Medical coverage
- Dental coverage
- Vision coverage
- Flexible Spending Account
- Health Savings Account
- Disability insurance
- Life insurance
- 401(k) retirement plan
- Commuter benefits
- Wellness services
- Employee Assistance Programme
