Full Stack LLM Engineer

Cerebras builds wafer-scale AI computing systems and a cloud inference platform for training, fine-tuning, and serving AI models.

Sunnyvale, California, United States
About Cerebras Systems, Inc.

Cerebras Systems is an AI-infrastructure company founded in 2015. It sells rack-scale wafer-scale computing systems and provides cloud-based, API-accessible AI inference alongside on-premises deployments.

View jobs by Cerebras Systems, Inc.

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will bring up machine-learning models on CSX systems and work across model translation, graph lowering, compiler optimization, runtime integration, and performance tuning. You will debug performance and correctness issues and prototype improvements to tools, APIs, and automation flows.

Requirements

  • Python modeling experience
  • Compiler IR knowledge
  • Performance profiling experience
  • Debugging skills for performance, numerical accuracy, and runtime integration
  • PyTorch or TensorFlow experience
  • Knowledge of attention, mixture-of-experts, or diffusion models
  • C and C++ proficiency
  • Low-level optimization experience
  • LLVM or MLIR compiler development experience
  • Optimization techniques involving NP-hard problems

Responsibilities

  • Bring up machine-learning models on CSX systems
  • Translate model architectures and lower graphs
  • Optimize compilers and tune performance
  • Integrate models with runtimes
  • Debug performance and correctness issues across model code, compiler IRs, runtimes, and hardware utilization
  • Prototype improvements to tools, APIs, and automation flows
Full Stack LLM Engineer at Cerebras Systems, Inc. | JobStash