Staff GPU Inference SDET

Cerebras builds wafer-scale AI computing systems and a cloud inference platform for training, fine-tuning, and serving AI models.

Sunnyvale, California, United States
About Cerebras Systems, Inc.

Cerebras Systems is an AI-infrastructure company founded in 2015. It sells rack-scale wafer-scale computing systems and provides cloud-based, API-accessible AI inference alongside on-premises deployments.

View jobs by Cerebras Systems, Inc.

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will build automated test frameworks, regression gates, and release qualification pipelines for GPU inference systems. You will validate distributed serving workloads, performance models, numerical correctness, resilience, and recovery; integrate observability and CI/CD; and investigate software and hardware failures.

Requirements

  • Eight or more years of software engineering experience as an SDET, infrastructure quality lead, or systems test engineer
  • Experience bringing up, provisioning, and validating multi-node GPU clusters
  • Knowledge of LLM serving engines, prefill and decode disaggregation, KV-cache management, and dynamic batching
  • Expert Python programming and test automation framework experience
  • Kubernetes, Slurm, or Ray
  • InfiniBand, RoCE, or NCCL
  • Root-cause analysis, stress testing, and node failure simulation

Responsibilities

  • Design automated test frameworks, regression gates, and release qualification pipelines
  • Benchmark and stress-test distributed LLM serving frameworks
  • Build workload replay and benchmarking tools
  • Validate model accuracy, precision stability, determinism, and output correctness
  • Engineer fault-injection and recovery tests for GPU clusters
  • Integrate test pipelines with telemetry and continuous monitoring
Staff GPU Inference SDET at Cerebras Systems, Inc. | JobStash