Staff Software Engineer GPU Inference

Cerebras builds wafer-scale AI computing systems and a cloud inference platform for training, fine-tuning, and serving AI models.

Sunnyvale, California, United States
About Cerebras Systems, Inc.

Cerebras Systems is an AI-infrastructure company founded in 2015. It sells rack-scale wafer-scale computing systems and provides cloud-based, API-accessible AI inference alongside on-premises deployments.

View jobs by Cerebras Systems, Inc.

Skills

About the Role

You will build and operate the GPU prefill path across APIs, model-serving workers, runtimes, GPU nodes, networking, and rack-scale infrastructure. You will improve operational readiness, reliability, numerical correctness, performance, and capacity efficiency while developing benchmarks, profiling automation, dashboards, and regression gates.

Requirements

  • 8+ years of software engineering experience
  • Experience building, operating, or optimizing production inference systems for demanding GPU workloads
  • Programming experience in C++ and Python
  • Experience with a model-serving framework such as vLLM, SGLang, TensorRT-LLM, or Triton Inference Server
  • Understanding of GPU execution and performance
  • Experience debugging distributed systems across multiple layers
  • Experience with Linux, containers, Kubernetes, observability, CI/CD, and latency-sensitive production services
  • Ability to design benchmarks and translate performance findings into production improvements
  • Communication and technical leadership skills

Responsibilities

  • Design, build, deploy, and maintain the GPU inference stack
  • Establish deployment, upgrade, rollback, health-checking, capacity-management, and failure-recovery practices
  • Define service-level indicators and objectives for GPU-backed inference
  • Optimize time to first token, throughput, latency, GPU utilization, memory efficiency, and capacity
  • Improve scheduling, batching, caching, parallelism, admission, quantization, and graph execution
  • Debug failures and performance regressions across application, runtime, distributed systems, and hardware layers
  • Build validation and regression infrastructure for model quality and numerical correctness
  • Develop benchmarks, workload replay tools, profiling automation, dashboards, and regression gates
Staff Software Engineer GPU Inference at Cerebras Systems, Inc. | JobStash