Principal SRE AI Inference

Cerebras builds wafer-scale AI computing systems and a cloud inference platform for training, fine-tuning, and serving AI models.

Sunnyvale, California, United States
About Cerebras Systems, Inc.

Cerebras Systems is an AI-infrastructure company founded in 2015. It sells rack-scale wafer-scale computing systems and provides cloud-based, API-accessible AI inference alongside on-premises deployments.

View jobs by Cerebras Systems, Inc.

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will shape the architecture for reliable inference infrastructure across data centers and cloud environments. You will build self-service operational platforms, evolve SRE practices, lead incident escalations, mentor senior SREs, and measure improvements in reliability and operational efficiency.

Requirements

  • 15+ years of SRE, infrastructure engineering, or platform engineering experience
  • Experience with large-scale compute fleets, control planes, schedulers, orchestration systems, capacity management, and reliability automation
  • Experience driving cross-team architecture for production control planes, fleet management, or self-service infrastructure
  • Experience with observability, incident response, and SLO-based reliability management
  • Ability to lead ambiguous technical programs and mentor senior engineers

Responsibilities

  • Define and implement reliability strategy across data centers and cloud solutions
  • Architect self-service platforms and internal tooling for critical workflows
  • Define SLOs, SLIs, error budgets, postmortems, chaos testing, and capacity forecasting
  • Mentor senior SREs and support critical incident escalations
  • Prioritize automation based on production pain points
  • Measure toil reduction, deployment velocity, SLO compliance, MTTR, and self-service adoption
Principal SRE AI Inference at Cerebras Systems, Inc. | JobStash