ML Runtime and Kernel Engineer Core ML

Cerebras builds wafer-scale AI computing systems and a cloud inference platform for training, fine-tuning, and serving AI models.

Sunnyvale, California, United States
About Cerebras Systems, Inc.

Cerebras Systems is an AI-infrastructure company founded in 2015. It sells rack-scale wafer-scale computing systems and provides cloud-based, API-accessible AI inference alongside on-premises deployments.

View jobs by Cerebras Systems, Inc.

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will translate machine learning research prototypes into efficient production implementations. You will develop runtime components and kernels, profile and debug performance across software layers, optimize large-scale training and low-latency inference, build validation tooling, and help guide platform architecture decisions.

Requirements

  • Experience developing high-performance systems software, ML systems, runtimes, compilers, or computational kernels
  • C++
  • Python
  • Parallel programming
  • Memory management
  • Concurrency
  • Data structures
  • Performance optimization
  • Debugging and profiling complex software
  • PyTorch or JAX

Responsibilities

  • Design and implement runtime components and high-performance kernels
  • Translate research prototypes into efficient implementations and GPU comparisons
  • Profile and debug performance across framework, compiler, runtime, communication, and kernel layers
  • Optimize computation, memory movement, communication, and concurrency
  • Develop benchmarks, instrumentation, and automated tests
  • Collaborate to evaluate design alternatives and deliver end-to-end capabilities
  • Contribute to software architecture and roadmap decisions