Staff Site Reliability Engineer Automation and Platform
Cerebras Systems, Inc.Visit Cerebras Systems, Inc. website
Cerebras builds wafer-scale AI computing systems and a cloud inference platform for training, fine-tuning, and serving AI models.
Sunnyvale, California, United States
About Cerebras Systems, Inc.
Cerebras Systems is an AI-infrastructure company founded in 2015. It sells rack-scale wafer-scale computing systems and provides cloud-based, API-accessible AI inference alongside on-premises deployments.
Skills
About the Role
You will lead the design of reliable, scalable software delivery and operational platforms. You will build self-service workflows, evolve reliability practices, automate operational toil, support incident escalations, mentor SREs, and measure improvements in deployment velocity, reliability, and service ownership.
Requirements
- 8+ years of SRE, infrastructure engineering, or platform engineering experience
- Experience improving automation and reliability at large scale
- Expertise operating large-scale heterogeneous clusters with a proprietary cloud control plane
- Experience designing CI/CD or GitOps systems using Argo CD or similar tools
- Experience with Loki, Tempo, Mimir, and Prometheus
- Ability to lead complex projects and influence cross-functional stakeholders
Responsibilities
- Define and implement a strategy for reliably delivering and operating software across multiple datacenters and cloud solutions
- Architect self-service platforms and internal tooling for critical workflows
- Define and evolve SLOs, SLIs, error budgets, postmortems, chaos testing, and capacity forecasting
- Mentor SREs and prioritize automation based on production pain points
- Measure toil reduction, deployment velocity, SLO compliance, MTTR, and self-service adoption
