IT SRE Team Lead
Cerebras builds wafer-scale AI computing systems and a cloud inference platform for training, fine-tuning, and serving AI models.
About Cerebras Systems, Inc.
Cerebras Systems is an AI-infrastructure company founded in 2015. It sells rack-scale wafer-scale computing systems and provides cloud-based, API-accessible AI inference alongside on-premises deployments.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will define the reliability strategy for internal IT systems and lead engineers responsible for automation, observability, and incident response. You will automate provisioning, access management, patching, and lifecycle operations; establish SLOs and operational reporting; manage outages and remediation; and drive infrastructure-as-code and GitOps practices.
Requirements
- At least 8 years of SRE, DevOps, or IT engineering experience, including 2 years in leadership
- Experience building and deploying AI agents for triage and bug fixes
- Software engineering experience with Python, Go, or similar languages
- Experience with Okta, Entra, Jamf, Intune, and SaaS integrations
- Experience with Terraform and CI/CD pipelines
- Experience running on-call rotations, defining SLOs, and improving operational maturity
- Experience supporting technical engineering populations
Responsibilities
- Define and own the reliability strategy for internal IT systems, including SLOs, error budgets, and health reporting
- Build and lead IT SRE engineers focused on automation, observability, and incident response
- Automate provisioning, access management, patching, and lifecycle operations
- Instrument internal services and SaaS integrations with monitoring, alerting, and on-call workflows
- Run incident response, root-cause analysis, and durable remediation for IT outages
- Drive infrastructure-as-code and GitOps practices
- Partner on identity, access, and network reliability
