Senior SDET Inference Platform
Cerebras builds wafer-scale AI computing systems and a cloud inference platform for training, fine-tuning, and serving AI models.
About Cerebras Systems, Inc.
Cerebras Systems is an AI-infrastructure company founded in 2015. It sells rack-scale wafer-scale computing systems and provides cloud-based, API-accessible AI inference alongside on-premises deployments.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will build and maintain automated test infrastructure for an inference platform. You will validate Kubernetes deployments in cloud and hardware environments, test CI/CD, ingress, service discovery, NGINX, and load balancing, and investigate distributed-system issues. You will develop testbeds, identify reliability risks, improve observability and debugging, and create validation strategies for platform releases.
Requirements
- 3+ years of experience in software engineering, quality engineering, systems engineering, or infrastructure development
- Strong programming skills in Python or Go
- Experience building automation tools, testing frameworks, or internal developer tooling
- Hands-on experience with CI/CD systems such as Jenkins
- Experience debugging complex systems, distributed services, or networked infrastructure
- Familiarity with systems-level development, infrastructure tooling, or platform integration
- Problem-solving ability across multiple system and infrastructure layers
- Communication and collaboration skills
- Experience mentoring junior engineers
Responsibilities
- Design, build, and maintain test infrastructure and automation
- Validate the platform across cloud-managed Kubernetes and hardware deployments
- Test Kubernetes workloads, CI/CD pipelines, ingress, service discovery, NGINX, and load balancing
- Collaborate to ensure platform features ship reliably
- Investigate and debug networking, orchestration, deployment, and distributed-service issues
- Develop and maintain testbeds for performance, scalability, and reliability validation
- Identify failure points, bottlenecks, and edge cases affecting platform stability
- Contribute test plans and validation strategies for platform releases
- Improve observability, diagnostics, and debugging workflows
- Partner to deliver high-quality production-ready releases
