AI Inference Core Senior Software Engineer for Platform and DevOps
Cerebras Systems, Inc.Visit Cerebras Systems, Inc. website
Cerebras builds wafer-scale AI computing systems and a cloud inference platform for training, fine-tuning, and serving AI models.
Sunnyvale, California, United States
About Cerebras Systems, Inc.
Cerebras Systems is an AI-infrastructure company founded in 2015. It sells rack-scale wafer-scale computing systems and provides cloud-based, API-accessible AI inference alongside on-premises deployments.
Skills
About the Role
You will build and operate the platform layer for engineering infrastructure. You will maintain CI/CD and Kubernetes systems, automate deployments and self-service workflows, improve reliability and observability, debug cross-system failures, perform root-cause analysis, and implement durable platform improvements.
Requirements
- 5+ years of professional experience in platform engineering, DevOps, infrastructure engineering, site reliability engineering, or software engineering
- Experience building or maintaining CI/CD pipelines and automated delivery workflows
- Experience operating Kubernetes and containerized services
- Experience with a major cloud platform and programmatic infrastructure provisioning
- Understanding of Linux or Unix fundamentals
- Understanding of DNS, routing, load balancing, proxies, ports, TLS, and service connectivity
- Proficiency in Python, Shell, or another infrastructure-automation language
- Experience with monitoring, logging, alerting, dashboards, and incident investigation
Responsibilities
- Design, build, and maintain CI/CD systems for build, test, integration, qualification, and release workflows
- Build and operate Kubernetes-based platforms and services
- Develop deployment systems, internal tools, and self-service workflows
- Improve infrastructure reliability, capacity, performance, cost efficiency, monitoring, and operational readiness
- Debug issues across CI pipelines, Kubernetes, networking, storage, authentication, operating systems, and distributed applications
- Perform root-cause analysis and implement lasting fixes
- Deliver scalable infrastructure solutions with partner teams
