Senior Software Development Engineer in Test AI Cluster

Cerebras builds wafer-scale AI computing systems and a cloud inference platform for training, fine-tuning, and serving AI models.

Sunnyvale, California, United States
About Cerebras Systems, Inc.

Cerebras Systems is an AI-infrastructure company founded in 2015. It sells rack-scale wafer-scale computing systems and provides cloud-based, API-accessible AI inference alongside on-premises deployments.

View jobs by Cerebras Systems, Inc.

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will define test strategies and automate validation for large-scale AI infrastructure. You will break distributed ML training and inference systems into testable components, test cluster availability, failures, performance, stress, and security, and improve observability and reliability. You will validate cluster software, hardware, networking, monitoring, and deployment components while debugging distributed hardware and software issues.

Requirements

  • Bachelor's or master's degree in computer science, electrical engineering, AI, data science, or a related field
  • 5+ years of experience testing enterprise software, distributed systems, or datacenter hardware and software
  • Strong coding skills in Python, Go, or C/C++
  • Strong debugging skills for large distributed hardware and software systems
  • Experience with debugging tools such as pdb, gdb, strace, and network monitors
  • Understanding of operating-system internals, memory management, file systems, security, and performance
  • Understanding of datacenter layout and server, memory, BIOS, PCIe, networking, and storage characteristics

Responsibilities

  • Define and execute optimized test strategies and methodologies for AI infrastructure
  • Adapt testing approaches to new technologies and AI models
  • Break large distributed ML training and inference systems into unit-testable components
  • Automate tests for cluster availability, failure scenarios, performance, stress, and security
  • Champion cluster security, reliability, uptime, and observability
  • Test AI cluster software, hardware, interconnects, and monitoring components