Software Engineer GPU Inference

Cerebras builds wafer-scale AI computing systems and a cloud inference platform for training, fine-tuning, and serving AI models.

Sunnyvale, California, United States
About Cerebras Systems, Inc.

Cerebras Systems is an AI-infrastructure company founded in 2015. It sells rack-scale wafer-scale computing systems and provides cloud-based, API-accessible AI inference alongside on-premises deployments.

View jobs by Cerebras Systems, Inc.

Skills

About the Role

You will build, deploy, and operate the GPU prefill path across API services, serving workers, vLLM, PyTorch, ROCm, GPU nodes, networking, and rack-scale infrastructure. You will improve reliability, numerical correctness, observability, latency, throughput, and capacity efficiency through debugging, benchmarking, automation, and release validation.

Requirements

  • 5+ years of software engineering experience
  • Production inference systems experience for large language models, multimodal models, or demanding GPU workloads
  • C++
  • Python
  • Multithreading
  • Concurrency
  • Memory management
  • vLLM, SGLang, TensorRT-LLM, Triton Inference Server, or an equivalent serving framework
  • GPU execution and performance optimization
  • Distributed-system debugging
  • Linux
  • Containerization
  • Kubernetes or comparable orchestration
  • Observability
  • CI/CD
  • Benchmarking
  • Technical leadership
  • Computer Science, Computer Engineering, Electrical Engineering, related degree, or equivalent practical experience

Responsibilities

  • Design, build, deploy, and maintain the GPU prefill path
  • Establish deployment, upgrade, rollback, health-checking, capacity-management, and recovery practices
  • Define service-level indicators and objectives for GPU-backed inference
  • Profile and optimize inference latency, throughput, utilization, memory efficiency, and capacity
  • Tune model-serving scheduling, batching, caching, parallelism, admission, quantization, and graph execution
  • Diagnose failures and regressions across application, runtime, distributed-system, and hardware layers
  • Build validation infrastructure for model quality, numerical accuracy, determinism, and compatibility
  • Develop benchmarks, workload replay tools, profiling automation, dashboards, and regression gates
Software Engineer GPU Inference at Cerebras Systems, Inc. | JobStash