CoDesign and NextGen Performance Engineer
Cerebras Systems, Inc.Visit Cerebras Systems, Inc. website
Cerebras builds wafer-scale AI computing systems and a cloud inference platform for training, fine-tuning, and serving AI models.
Sunnyvale, California, United States
About Cerebras Systems, Inc.
Cerebras Systems is an AI-infrastructure company founded in 2015. It sells rack-scale wafer-scale computing systems and provides cloud-based, API-accessible AI inference alongside on-premises deployments.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will characterize, analyze, and optimize AI model performance across the hardware and software stack. You will bring up new hardware generations, build performance models, optimize kernels and compiler algorithms, debug runtime performance, and develop tools to visualize cluster performance data.
Requirements
- Computer architecture
- Low-level deep learning and LLM mathematics
- 3+ years of relevant experience in computer architecture, CPU or GPU performance, kernel optimization, or HPC
- CPU or GPU simulators
- Performance profiling
- System-pipeline debugging
- C++
- Python
Responsibilities
- Bring up and optimize performance on new Wafer-Scale Engine generations
- Build kernel-level and end-to-end performance models
- Optimize and debug kernel microcode and compiler algorithms
- Debug runtime performance on systems and clusters
- Develop tools to visualize performance data
