ML Systems Performance Engineer
Cerebras Systems, Inc.Visit Cerebras Systems, Inc. website
Cerebras builds wafer-scale AI computing systems and a cloud inference platform for training, fine-tuning, and serving AI models.
Sunnyvale, California, United States
About Cerebras Systems, Inc.
Cerebras Systems is an AI-infrastructure company founded in 2015. It sells rack-scale wafer-scale computing systems and provides cloud-based, API-accessible AI inference alongside on-premises deployments.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will build kernel-level and end-to-end performance models for ML models. You will optimize kernel microcode and compiler algorithms, debug runtime and cluster performance, and develop infrastructure that visualizes performance data from the Wafer Scale Engine and compute cluster.
Requirements
- Computer architecture
- Low-level deep learning and LLM mathematics
- Analytical problem-solving
- 3+ years of relevant experience in computer architecture, CPU or GPU performance, kernel optimization, or HPC
- Experience with CPU or GPU simulators
- Performance profiling and debugging
- C++
- Python
Responsibilities
- Build kernel-level and end-to-end performance models for ML models
- Optimize and debug kernel microcode and compiler algorithms
- Debug runtime performance on systems and clusters
- Develop tools and infrastructure to visualize performance data
