Site Reliability Engineer Ops and Automation
Cerebras Systems, Inc.Visit Cerebras Systems, Inc. website
Cerebras builds wafer-scale AI computing systems and a cloud inference platform for training, fine-tuning, and serving AI models.
Sunnyvale, California, United States
About Cerebras Systems, Inc.
Cerebras Systems is an AI-infrastructure company founded in 2015. It sells rack-scale wafer-scale computing systems and provides cloud-based, API-accessible AI inference alongside on-premises deployments.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will operate production systems, manage releases, capacity changes, and cluster upgrades. You will build self-service delivery pipelines, automation, developer tools, telemetry, observability, and alerting. You will collaborate on reliability practices including SLOs, post-mortems, and capacity planning.
Requirements
- 2-4+ years of SRE experience with operations or automation focus
- Production Kubernetes experience
- Python or Go
- Prometheus and Grafana
- Observability-driven workflows
- Ability to measure and communicate reliability, operational-toil, and velocity impact
Responsibilities
- Operate releases, capacity changes, and cluster upgrades
- Develop self-service continuous-delivery pipelines
- Build reusable automation and internal developer tools
- Develop telemetry, observability, and alerting solutions
- Identify and implement automation opportunities
- Contribute to SLOs, post-mortems, and capacity planning
Benefits
- No 24/7 on-call rotations
