Principal Engineer Inference Cloud
Cerebras Systems, Inc.Visit Cerebras Systems, Inc. website
Cerebras builds wafer-scale AI computing systems and a cloud inference platform for training, fine-tuning, and serving AI models.
Sunnyvale, California, United States
About Cerebras Systems, Inc.
Cerebras Systems is an AI-infrastructure company founded in 2015. It sells rack-scale wafer-scale computing systems and provides cloud-based, API-accessible AI inference alongside on-premises deployments.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will set the technical direction for the inference cloud platform and solve high-leverage distributed-systems problems. You will design multi-region architecture, contribute production code, improve reliability and performance, lead production incident work, and mentor engineers.
Requirements
- 10+ years of software engineering experience with substantial individual-contributor experience in large-scale distributed systems or cloud infrastructure
- Expertise in cloud distributed-systems architecture, networking, compute orchestration, container platforms, and multi-region production services
- Experience designing highly available, latency-sensitive systems
- Experience optimizing high-QPS latency, throughput, and efficiency
- Proficiency in Go, C++, Python, or another backend or systems language
- Experience with metrics, logging, tracing, alerting, incident response, and SLI, SLO, and SLA operations
- Ability to influence senior engineers and cross-functional partners
Responsibilities
- Identify and prioritize high-leverage platform problems
- Set long-term direction for multi-region topology, failure domains, service boundaries, and system evolution
- Architect active-active systems with failover, graceful degradation, and SLOs
- Improve latency, throughput, capacity efficiency, and resilience
- Contribute production code and review designs and implementations
- Drive observability, incident response, capacity planning, and post-incident improvements
- Drive platform-wide decisions on reliability, API design, and deployment strategy
- Mentor engineers through design feedback and engineering standards
