Senior Research Engineer - Inference ML
Cerebras builds wafer-scale AI computing systems and a cloud inference platform for training, fine-tuning, and serving AI models.
About Cerebras Systems, Inc.
Cerebras Systems is an AI-infrastructure company founded in 2015. It sells rack-scale wafer-scale computing systems and provides cloud-based, API-accessible AI inference alongside on-premises deployments.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will adapt advanced language and vision models for efficient execution on Cerebras hardware. You will research, prototype, validate, and optimize inference architectures and algorithms, train and evaluate models, troubleshoot integrations, and develop performance diagnostics. You will collaborate across software, hardware, and product to deliver low-latency, high-throughput inference.
Requirements
- Bachelor’s degree plus 7+ years of ML software development experience, Master’s degree plus 4+ years of software development experience, PhD plus 2+ years of relevant experience, or equivalent practical experience.
- 4+ years of experience testing, maintaining, or launching software products, including 2+ years in software design and architecture.
- 3+ years of machine-learning-focused software development experience.
- Strong Python and/or C++ programming skills.
- Experience with generative AI and machine learning systems.
- Evidence of machine learning research impact through top-conference publications, open-source contributions, or high-quality preprints.
- Proficiency with PyTorch, Transformers, vLLM, or SGLang.
- Understanding of transformer-based language and/or vision models and experience implementing and optimizing them.
- Experience translating research into model variants, training strategies, and evaluation workflows.
- Performance optimization experience on specialized hardware.
Responsibilities
- Design, implement, and optimize transformer architectures for NLP and computer vision on Cerebras hardware.
- Research and prototype inference algorithms and model architectures focused on speculative decoding, pruning, compression, sparse attention, and sparsity.
- Train models to convergence, run hyperparameter sweeps, and analyze results.
- Bring up and validate new models on the Cerebras system and troubleshoot integration issues.
- Profile and optimize model code to maximize throughput and minimize latency.
- Develop diagnostic tooling and scripts to identify performance bottlenecks.
- Collaborate with software, hardware, and product stakeholders to deliver projects.
