Principal Engineer AI Inference Reliability
Cerebras Systems, Inc.Visit Cerebras Systems, Inc. website
Cerebras builds wafer-scale AI computing systems and a cloud inference platform for training, fine-tuning, and serving AI models.
Sunnyvale, California, United States
About Cerebras Systems, Inc.
Cerebras Systems is an AI-infrastructure company founded in 2015. It sells rack-scale wafer-scale computing systems and provides cloud-based, API-accessible AI inference alongside on-premises deployments.
Skills
About the Role
You will lead reliability strategy across the inference stack. You will define SLOs and incident-response frameworks, build fault-tolerance mechanisms and reliability tooling, lead incident reviews, and guide engineers in operating large-scale distributed systems.
Requirements
- Bachelor's or master's degree in computer science or a related field
- 7+ years of backend, infrastructure, or reliability engineering experience for large-scale distributed systems
- Programming proficiency in Python, C++, Go, Rust, or another backend language
- Experience with SLO, SLI, and SLA design, incident response, and postmortems
- Communication and cross-functional leadership skills
Responsibilities
- Define and drive reliability strategy and establish SLOs
- Design fault detection, graceful degradation, failover, throttling, and recovery mechanisms
- Lead incident management, postmortems, root-cause analysis, and prevention loops
- Influence system architecture for redundancy, durability, observability, and debuggability
- Develop tooling for chaos testing, load simulation, and distributed fault injection
- Build dashboards and alerts for service-health metrics
- Mentor engineers on reliable system design, testing, and operations
