Staff Site Reliability Engineer Automation and Platform

Cerebras builds wafer-scale AI computing systems and a cloud inference platform for training, fine-tuning, and serving AI models.

Sunnyvale, California, United States
About Cerebras Systems, Inc.

Cerebras Systems is an AI-infrastructure company founded in 2015. It sells rack-scale wafer-scale computing systems and provides cloud-based, API-accessible AI inference alongside on-premises deployments.

View jobs by Cerebras Systems, Inc.

Skills

About the Role

You will lead the design of reliable, scalable software delivery and operational platforms. You will build self-service workflows, evolve reliability practices, automate operational toil, support incident escalations, mentor SREs, and measure improvements in deployment velocity, reliability, and service ownership.

Requirements

  • 8+ years of SRE, infrastructure engineering, or platform engineering experience
  • Experience improving automation and reliability at large scale
  • Expertise operating large-scale heterogeneous clusters with a proprietary cloud control plane
  • Experience designing CI/CD or GitOps systems using Argo CD or similar tools
  • Experience with Loki, Tempo, Mimir, and Prometheus
  • Ability to lead complex projects and influence cross-functional stakeholders

Responsibilities

  • Define and implement a strategy for reliably delivering and operating software across multiple datacenters and cloud solutions
  • Architect self-service platforms and internal tooling for critical workflows
  • Define and evolve SLOs, SLIs, error budgets, postmortems, chaos testing, and capacity forecasting
  • Mentor SREs and prioritize automation based on production pain points
  • Measure toil reduction, deployment velocity, SLO compliance, MTTR, and self-service adoption
Staff Site Reliability Engineer Automation and Platform at Cerebras Systems, Inc. | JobStash