ML Algorithm Mapping and Performance Engineer Core ML

Cerebras builds wafer-scale AI computing systems and a cloud inference platform for training, fine-tuning, and serving AI models.

Sunnyvale, California, United States
About Cerebras Systems, Inc.

Cerebras Systems is an AI-infrastructure company founded in 2015. It sells rack-scale wafer-scale computing systems and provides cloud-based, API-accessible AI inference alongside on-premises deployments.

View jobs by Cerebras Systems, Inc.

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will build analytical and empirical performance models for machine learning training and inference algorithms. You will benchmark and prototype implementations, analyze bottlenecks across the system stack, evaluate emerging ML techniques, and communicate performance trade-offs and recommendations through tools, reports, presentations, and design reviews.

Requirements

  • Foundation in computer architecture, parallel computing, and systems performance
  • Understanding of machine learning fundamentals and ML systems
  • Experience with performance modeling, complexity analysis, benchmarking, or system simulation
  • Python
  • C++
  • Experience profiling and debugging performance in ML, HPC, CPU, GPU, or accelerator-based systems

Responsibilities

  • Build analytical and empirical performance models for machine learning training and inference algorithms
  • Characterize asymptotic behavior and algorithmic trade-offs across model and hardware scale
  • Construct Pareto frontiers for model quality, latency, throughput, memory, communication, and compute cost
  • Develop prototype implementations and benchmarks for the Cerebras WSE and relevant GPU or software baselines
  • Analyze kernel, compiler, runtime, communication, and algorithmic bottlenecks
  • Evaluate emerging ML techniques including parallel generation, diffusion, speculative decoding, attention, sparsity, mixture-of-experts, low-precision computation, and distributed training
  • Recommend high-value implementation and hardware-software co-design directions
  • Develop tools and visualizations for performance projections and design trade-offs
  • Communicate conclusions and recommendations through technical reports, presentations, and design reviews