Staff Software Engineer GPU Inference
Cerebras Systems, Inc.Visit Cerebras Systems, Inc. website
Cerebras builds wafer-scale AI computing systems and a cloud inference platform for training, fine-tuning, and serving AI models.
Sunnyvale, California, United States
About Cerebras Systems, Inc.
Cerebras Systems is an AI-infrastructure company founded in 2015. It sells rack-scale wafer-scale computing systems and provides cloud-based, API-accessible AI inference alongside on-premises deployments.
Skills
About the Role
You will build and operate the GPU prefill path across APIs, model-serving workers, runtimes, GPU nodes, networking, and rack-scale infrastructure. You will improve operational readiness, reliability, numerical correctness, performance, and capacity efficiency while developing benchmarks, profiling automation, dashboards, and regression gates.
Requirements
- 8+ years of software engineering experience
- Experience building, operating, or optimizing production inference systems for demanding GPU workloads
- Programming experience in C++ and Python
- Experience with a model-serving framework such as vLLM, SGLang, TensorRT-LLM, or Triton Inference Server
- Understanding of GPU execution and performance
- Experience debugging distributed systems across multiple layers
- Experience with Linux, containers, Kubernetes, observability, CI/CD, and latency-sensitive production services
- Ability to design benchmarks and translate performance findings into production improvements
- Communication and technical leadership skills
Responsibilities
- Design, build, deploy, and maintain the GPU inference stack
- Establish deployment, upgrade, rollback, health-checking, capacity-management, and failure-recovery practices
- Define service-level indicators and objectives for GPU-backed inference
- Optimize time to first token, throughput, latency, GPU utilization, memory efficiency, and capacity
- Improve scheduling, batching, caching, parallelism, admission, quantization, and graph execution
- Debug failures and performance regressions across application, runtime, distributed systems, and hardware layers
- Build validation and regression infrastructure for model quality and numerical correctness
- Develop benchmarks, workload replay tools, profiling automation, dashboards, and regression gates
