AI Inference Core Senior Software Engineer for Platform and DevOps

Cerebras builds wafer-scale AI computing systems and a cloud inference platform for training, fine-tuning, and serving AI models.

Sunnyvale, California, United States
About Cerebras Systems, Inc.

Cerebras Systems is an AI-infrastructure company founded in 2015. It sells rack-scale wafer-scale computing systems and provides cloud-based, API-accessible AI inference alongside on-premises deployments.

View jobs by Cerebras Systems, Inc.

Skills

About the Role

You will build and operate the platform layer for engineering infrastructure. You will maintain CI/CD and Kubernetes systems, automate deployments and self-service workflows, improve reliability and observability, debug cross-system failures, perform root-cause analysis, and implement durable platform improvements.

Requirements

  • 5+ years of professional experience in platform engineering, DevOps, infrastructure engineering, site reliability engineering, or software engineering
  • Experience building or maintaining CI/CD pipelines and automated delivery workflows
  • Experience operating Kubernetes and containerized services
  • Experience with a major cloud platform and programmatic infrastructure provisioning
  • Understanding of Linux or Unix fundamentals
  • Understanding of DNS, routing, load balancing, proxies, ports, TLS, and service connectivity
  • Proficiency in Python, Shell, or another infrastructure-automation language
  • Experience with monitoring, logging, alerting, dashboards, and incident investigation

Responsibilities

  • Design, build, and maintain CI/CD systems for build, test, integration, qualification, and release workflows
  • Build and operate Kubernetes-based platforms and services
  • Develop deployment systems, internal tools, and self-service workflows
  • Improve infrastructure reliability, capacity, performance, cost efficiency, monitoring, and operational readiness
  • Debug issues across CI pipelines, Kubernetes, networking, storage, authentication, operating systems, and distributed applications
  • Perform root-cause analysis and implement lasting fixes
  • Deliver scalable infrastructure solutions with partner teams