Site Reliability Engineer Ops and Automation

Cerebras builds wafer-scale AI computing systems and a cloud inference platform for training, fine-tuning, and serving AI models.

Sunnyvale, California, United States
About Cerebras Systems, Inc.

Cerebras Systems is an AI-infrastructure company founded in 2015. It sells rack-scale wafer-scale computing systems and provides cloud-based, API-accessible AI inference alongside on-premises deployments.

View jobs by Cerebras Systems, Inc.

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will operate production systems, manage releases, capacity changes, and cluster upgrades. You will build self-service delivery pipelines, automation, developer tools, telemetry, observability, and alerting. You will collaborate on reliability practices including SLOs, post-mortems, and capacity planning.

Requirements

  • 2-4+ years of SRE experience with operations or automation focus
  • Production Kubernetes experience
  • Python or Go
  • Prometheus and Grafana
  • Observability-driven workflows
  • Ability to measure and communicate reliability, operational-toil, and velocity impact

Responsibilities

  • Operate releases, capacity changes, and cluster upgrades
  • Develop self-service continuous-delivery pipelines
  • Build reusable automation and internal developer tools
  • Develop telemetry, observability, and alerting solutions
  • Identify and implement automation opportunities
  • Contribute to SLOs, post-mortems, and capacity planning

Benefits

  • No 24/7 on-call rotations
Site Reliability Engineer Ops and Automation at Cerebras Systems, Inc. | JobStash