Software Engineer GPU Inference
Cerebras Systems, Inc.Visit Cerebras Systems, Inc. website
Cerebras builds wafer-scale AI computing systems and a cloud inference platform for training, fine-tuning, and serving AI models.
Sunnyvale, California, United States
About Cerebras Systems, Inc.
Cerebras Systems is an AI-infrastructure company founded in 2015. It sells rack-scale wafer-scale computing systems and provides cloud-based, API-accessible AI inference alongside on-premises deployments.
Skills
About the Role
You will build, deploy, and operate the GPU prefill path across API services, serving workers, vLLM, PyTorch, ROCm, GPU nodes, networking, and rack-scale infrastructure. You will improve reliability, numerical correctness, observability, latency, throughput, and capacity efficiency through debugging, benchmarking, automation, and release validation.
Requirements
- 5+ years of software engineering experience
- Production inference systems experience for large language models, multimodal models, or demanding GPU workloads
- C++
- Python
- Multithreading
- Concurrency
- Memory management
- vLLM, SGLang, TensorRT-LLM, Triton Inference Server, or an equivalent serving framework
- GPU execution and performance optimization
- Distributed-system debugging
- Linux
- Containerization
- Kubernetes or comparable orchestration
- Observability
- CI/CD
- Benchmarking
- Technical leadership
- Computer Science, Computer Engineering, Electrical Engineering, related degree, or equivalent practical experience
Responsibilities
- Design, build, deploy, and maintain the GPU prefill path
- Establish deployment, upgrade, rollback, health-checking, capacity-management, and recovery practices
- Define service-level indicators and objectives for GPU-backed inference
- Profile and optimize inference latency, throughput, utilization, memory efficiency, and capacity
- Tune model-serving scheduling, batching, caching, parallelism, admission, quantization, and graph execution
- Diagnose failures and regressions across application, runtime, distributed-system, and hardware layers
- Build validation infrastructure for model quality, numerical accuracy, determinism, and compatibility
- Develop benchmarks, workload replay tools, profiling automation, dashboards, and regression gates
