Principal Engineer AI Inference Reliability

Cerebras builds wafer-scale AI computing systems and a cloud inference platform for training, fine-tuning, and serving AI models.

Sunnyvale, California, United States
About Cerebras Systems, Inc.

Cerebras Systems is an AI-infrastructure company founded in 2015. It sells rack-scale wafer-scale computing systems and provides cloud-based, API-accessible AI inference alongside on-premises deployments.

View jobs by Cerebras Systems, Inc.

Skills

About the Role

You will lead reliability strategy across the inference stack. You will define SLOs and incident-response frameworks, build fault-tolerance mechanisms and reliability tooling, lead incident reviews, and guide engineers in operating large-scale distributed systems.

Requirements

  • Bachelor's or master's degree in computer science or a related field
  • 7+ years of backend, infrastructure, or reliability engineering experience for large-scale distributed systems
  • Programming proficiency in Python, C++, Go, Rust, or another backend language
  • Experience with SLO, SLI, and SLA design, incident response, and postmortems
  • Communication and cross-functional leadership skills

Responsibilities

  • Define and drive reliability strategy and establish SLOs
  • Design fault detection, graceful degradation, failover, throttling, and recovery mechanisms
  • Lead incident management, postmortems, root-cause analysis, and prevention loops
  • Influence system architecture for redundancy, durability, observability, and debuggability
  • Develop tooling for chaos testing, load simulation, and distributed fault injection
  • Build dashboards and alerts for service-health metrics
  • Mentor engineers on reliable system design, testing, and operations
Principal Engineer AI Inference Reliability at Cerebras Systems, Inc. | JobStash