ML Software Tool Development Engineer
Cerebras Systems, Inc.Visit Cerebras Systems, Inc. website
Cerebras builds wafer-scale AI computing systems and a cloud inference platform for training, fine-tuning, and serving AI models.
Sunnyvale, California, United States
About Cerebras Systems, Inc.
Cerebras Systems is an AI-infrastructure company founded in 2015. It sells rack-scale wafer-scale computing systems and provides cloud-based, API-accessible AI inference alongside on-premises deployments.
Skills
About the Role
You will develop system-level debugging, validation, and observability platforms. You will build automated anomaly analysis, visualization, failure-classification, regression-detection, profiling, and instrumentation tools. You will improve bring-up and validation workflows, support incident response, and lead corrective actions.
Requirements
- Proficiency in C++ and Python
- Experience building reliable, high-performance systems and tooling
- Experience debugging complex hardware and software systems to root cause
- Experience analyzing system-level data structures, execution graphs, or dependency networks
- Experience designing visualization and analysis tools for technical data
- Experience with compiler internals, custom hardware interfaces, or low-level protocol design
- Written and verbal communication skills
- Ability to independently lead complex technical projects end to end
- Familiarity with machine learning training and inference pipelines
- Knowledge of distributed training and large-model scaling
- Experience with high-performance clusters, HPC systems, or hardware and software co-design
Responsibilities
- Lead the design and implementation of system-level debugging, validation, and observability platforms
- Develop automated systems for collecting and analyzing numerical and execution anomalies
- Create visualization and analysis tools for root-cause investigation
- Build frameworks for failure classification, regression detection, and anomaly monitoring
- Extend compilers, runtimes, and programming interfaces for profiling and instrumentation
- Improve system bring-up, low-level debugging, and validation workflows
- Partner with compiler, hardware, firmware, runtime, and infrastructure teams
- Establish best practices for debuggability, reliability, and operational excellence
- Lead high-impact initiatives
- Support incident response and drive long-term corrective actions
