ML Runtime and Kernel Engineer Core ML
Cerebras builds wafer-scale AI computing systems and a cloud inference platform for training, fine-tuning, and serving AI models.
About Cerebras Systems, Inc.
Cerebras Systems is an AI-infrastructure company founded in 2015. It sells rack-scale wafer-scale computing systems and provides cloud-based, API-accessible AI inference alongside on-premises deployments.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will translate machine learning research prototypes into efficient production implementations. You will develop runtime components and kernels, profile and debug performance across software layers, optimize large-scale training and low-latency inference, build validation tooling, and help guide platform architecture decisions.
Requirements
- Experience developing high-performance systems software, ML systems, runtimes, compilers, or computational kernels
- C++
- Python
- Parallel programming
- Memory management
- Concurrency
- Data structures
- Performance optimization
- Debugging and profiling complex software
- PyTorch or JAX
Responsibilities
- Design and implement runtime components and high-performance kernels
- Translate research prototypes into efficient implementations and GPU comparisons
- Profile and debug performance across framework, compiler, runtime, communication, and kernel layers
- Optimize computation, memory movement, communication, and concurrency
- Develop benchmarks, instrumentation, and automated tests
- Collaborate to evaluate design alternatives and deliver end-to-end capabilities
- Contribute to software architecture and roadmap decisions
