CoDesign and NextGen Performance Engineer

Cerebras builds wafer-scale AI computing systems and a cloud inference platform for training, fine-tuning, and serving AI models.

Sunnyvale, California, United States
About Cerebras Systems, Inc.

Cerebras Systems is an AI-infrastructure company founded in 2015. It sells rack-scale wafer-scale computing systems and provides cloud-based, API-accessible AI inference alongside on-premises deployments.

View jobs by Cerebras Systems, Inc.

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will characterize, analyze, and optimize AI model performance across the hardware and software stack. You will bring up new hardware generations, build performance models, optimize kernels and compiler algorithms, debug runtime performance, and develop tools to visualize cluster performance data.

Requirements

  • Computer architecture
  • Low-level deep learning and LLM mathematics
  • 3+ years of relevant experience in computer architecture, CPU or GPU performance, kernel optimization, or HPC
  • CPU or GPU simulators
  • Performance profiling
  • System-pipeline debugging
  • C++
  • Python

Responsibilities

  • Bring up and optimize performance on new Wafer-Scale Engine generations
  • Build kernel-level and end-to-end performance models
  • Optimize and debug kernel microcode and compiler algorithms
  • Debug runtime performance on systems and clusters
  • Develop tools to visualize performance data
CoDesign and NextGen Performance Engineer at Cerebras Systems, Inc. | JobStash