Full Stack LLM Engineer
Cerebras Systems, Inc.Visit Cerebras Systems, Inc. website
Cerebras builds wafer-scale AI computing systems and a cloud inference platform for training, fine-tuning, and serving AI models.
Sunnyvale, California, United States
About Cerebras Systems, Inc.
Cerebras Systems is an AI-infrastructure company founded in 2015. It sells rack-scale wafer-scale computing systems and provides cloud-based, API-accessible AI inference alongside on-premises deployments.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will bring up machine-learning models on CSX systems and work across model translation, graph lowering, compiler optimization, runtime integration, and performance tuning. You will debug performance and correctness issues and prototype improvements to tools, APIs, and automation flows.
Requirements
- Python modeling experience
- Compiler IR knowledge
- Performance profiling experience
- Debugging skills for performance, numerical accuracy, and runtime integration
- PyTorch or TensorFlow experience
- Knowledge of attention, mixture-of-experts, or diffusion models
- C and C++ proficiency
- Low-level optimization experience
- LLVM or MLIR compiler development experience
- Optimization techniques involving NP-hard problems
Responsibilities
- Bring up machine-learning models on CSX systems
- Translate model architectures and lower graphs
- Optimize compilers and tune performance
- Integrate models with runtimes
- Debug performance and correctness issues across model code, compiler IRs, runtimes, and hardware utilization
- Prototype improvements to tools, APIs, and automation flows
