ML Systems Performance Engineer

Cerebras builds wafer-scale AI computing systems and a cloud inference platform for training, fine-tuning, and serving AI models.

Sunnyvale, California, United States
About Cerebras Systems, Inc.

Cerebras Systems is an AI-infrastructure company founded in 2015. It sells rack-scale wafer-scale computing systems and provides cloud-based, API-accessible AI inference alongside on-premises deployments.

View jobs by Cerebras Systems, Inc.

Skills

About the Role

You will build kernel-level and end-to-end performance models for ML models. You will optimize kernel microcode and compiler algorithms, debug runtime and cluster performance, and develop infrastructure that visualizes performance data from the Wafer Scale Engine and compute cluster.

Requirements

  • Computer architecture
  • Low-level deep learning and LLM mathematics
  • Analytical problem-solving
  • 3+ years of relevant experience in computer architecture, CPU or GPU performance, kernel optimization, or HPC
  • Experience with CPU or GPU simulators
  • Performance profiling and debugging
  • C++
  • Python

Responsibilities

  • Build kernel-level and end-to-end performance models for ML models
  • Optimize and debug kernel microcode and compiler algorithms
  • Debug runtime performance on systems and clusters
  • Develop tools and infrastructure to visualize performance data