AI Fleet Platform Software Engineer

Cerebras builds wafer-scale AI computing systems and a cloud inference platform for training, fine-tuning, and serving AI models.

Sunnyvale, California, United States
About Cerebras Systems, Inc.

Cerebras Systems is an AI-infrastructure company founded in 2015. It sells rack-scale wafer-scale computing systems and provides cloud-based, API-accessible AI inference alongside on-premises deployments.

View jobs by Cerebras Systems, Inc.

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will build and operate software that manages large fleets of AI clusters. You will create services, integrations, operational tools, and user-facing applications that help operators monitor health, capacity, performance, and incidents. You will lead projects from design through production, automate operational workflows, and improve reliability as the fleet grows.

Requirements

  • 12+ years of industry experience building and operating production software for distributed systems or large-scale infrastructure
  • Strong Go or Python skills
  • Experience designing services and APIs
  • Expertise in control planes, fleet management systems, or operational platforms
  • Experience with Linux, containers, Kubernetes, and distributed-system failures
  • Experience designing for asynchronous work, retries, and partial failures
  • Experience with event streaming, workflow automation, or time-series telemetry
  • Strong judgment in reliability, security, and observability
  • Ability to lead ambiguous projects and collaborate across engineering and operations teams

Responsibilities

  • Build and operate software for managing large fleets of AI clusters
  • Provide operators with actionable views of cluster health, capacity, performance, and issues
  • Develop services and integrations across infrastructure systems
  • Automate incident investigation and service-restoration workflows
  • Design reliable systems that withstand component and site failures
  • Gather platform-user needs and make practical product and engineering decisions
  • Lead projects from design through production and use operational feedback to improve them