Cluster Operations Software Engineer
Cerebras builds wafer-scale AI computing systems and a cloud inference platform for training, fine-tuning, and serving AI models.
About Cerebras Systems, Inc.
Cerebras Systems is an AI-infrastructure company founded in 2015. It sells rack-scale wafer-scale computing systems and provides cloud-based, API-accessible AI inference alongside on-premises deployments.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will operate advanced AI compute clusters, maintain their health and availability, and optimize capacity. You will build operational software, monitoring platforms, automation services, APIs, dashboards, and reliability tooling. You will troubleshoot incidents, handle engineering escalations, and participate in a 24/7 on-call rotation.
Requirements
- 6-8 years of relevant experience managing complex compute infrastructure
- Python
- Go
- Distributed systems
- Linux compute systems and command-line tools
- Docker
- Kubernetes
- Monitoring and alerting systems
- Troubleshooting complex technical issues
- Willingness to participate in a 24/7 on-call rotation
Responsibilities
- Deploy, configure, and debug container-based services using Docker
- Build software solutions for cluster operations, monitoring, workflow automation, dashboards, and reliability
- Develop APIs, automation services, and integrations for operational visibility and fleet management
- Manage and operate AI compute infrastructure clusters
- Monitor cluster health and resolve potential issues
- Optimize compute capacity and resource allocation
- Provide 24/7 monitoring and troubleshooting support
- Handle engineering escalations and resolve technical challenges
