Staff Software Engineer Inference Cloud

Cerebras builds wafer-scale AI computing systems and a cloud inference platform for training, fine-tuning, and serving AI models.

Sunnyvale, California, United States
About Cerebras Systems, Inc.

Cerebras Systems is an AI-infrastructure company founded in 2015. It sells rack-scale wafer-scale computing systems and provides cloud-based, API-accessible AI inference alongside on-premises deployments.

View jobs by Cerebras Systems, Inc.

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will own major areas of cloud-platform architecture for inference services. You will build distributed platform components, improve traffic control and reliability, write and review production code, resolve critical production issues, drive observability and capacity planning, and mentor senior engineers.

Requirements

  • 8+ years of software engineering experience
  • Experience building and operating large-scale distributed systems or cloud infrastructure
  • Expertise in distributed systems architecture, networking, compute orchestration, container platforms, and multi-region production services
  • Experience making architectural decisions for highly available, latency-sensitive systems
  • Experience optimizing latency, throughput, and efficiency in high-QPS systems
  • Proficiency in Go, C++, or Python
  • Experience with metrics, logging, tracing, alerting, incident response, and SLO-driven operations
  • Ability to influence senior engineers and cross-functional partners

Responsibilities

  • Shape platform direction for multi-region topology, failure domains, service boundaries, and system evolution
  • Design and build service discovery, request routing, load balancing, caching, batching, and traffic-management components
  • Architect active-active systems with failover, graceful degradation, and SLOs
  • Define admission control, quota management, rate limiting, and quality-of-service mechanisms
  • Write and review production code on critical platform paths
  • Lead production issue resolution, observability, incident response, capacity planning, and post-incident improvement
  • Translate requirements into scalable system designs with adjacent teams
  • Mentor senior engineers through design feedback, pairing, and technical standards