Engineering Manager Kernel Reliability

Cerebras builds wafer-scale AI computing systems and a cloud inference platform for training, fine-tuning, and serving AI models.

Sunnyvale, California, United States
About Cerebras Systems, Inc.

Cerebras Systems is an AI-infrastructure company founded in 2015. It sells rack-scale wafer-scale computing systems and provides cloud-based, API-accessible AI inference alongside on-premises deployments.

View jobs by Cerebras Systems, Inc.

Skills

About the Role

You will provide hands-on technical leadership for kernel-centric reliability across internal and customer-facing systems. You will set the technical vision and roadmap, develop diagnostic capabilities, improve failure analysis and debugging, collaborate on software and hardware design, and lead and mentor engineers.

Requirements

  • 6+ years of software engineering experience
  • 3+ years leading teams in software or hardware reliability, debugging, diagnostics, or failure analysis
  • Parallel and distributed programming expertise
  • Debugging and diagnostic tool development or expert usage
  • Experience debugging distributed and parallel applications
  • Deep understanding of computer architectures
  • Monitoring and reliability engineering background
  • Ability to recruit, retain, mentor, and partner cross-functionally

Responsibilities

  • Own the technical vision and roadmap for kernel-centric reliability
  • Provide tooling and manual intervention for failure analysis and diagnostics
  • Enhance debug tools to accelerate failure analysis
  • Improve kernels and the software stack for on-field debugging
  • Co-design next-generation architectures for reliability and debuggability
  • Lead, mentor, and grow engineers