Senior Member of Technical Staff

Cerebras builds wafer-scale AI computing systems and a cloud inference platform for training, fine-tuning, and serving AI models.

Sunnyvale, California, United States
About Cerebras Systems, Inc.

Cerebras Systems is an AI-infrastructure company founded in 2015. It sells rack-scale wafer-scale computing systems and provides cloud-based, API-accessible AI inference alongside on-premises deployments.

View jobs by Cerebras Systems, Inc.

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will design and develop resilient, highly available software for distributed inference environments. You will create cloud deployment workflows, automation, Python APIs, containerized services, Kubernetes strategies, monitoring tools, documentation, and defect-resolution processes.

Requirements

  • Master's degree or foreign equivalent in Computer Science or a related field
  • 18 months of relevant experience
  • Terraform
  • AWS CloudFormation
  • AWS CDK
  • Ansible
  • Docker
  • Kubernetes
  • AWS EKS
  • AWS ECS
  • AWS Fargate
  • Helm
  • AWS EC2
  • AWS Lambda
  • AWS Auto Scaling
  • AWS CloudWatch
  • AWS X-Ray
  • ELK
  • Prometheus
  • Grafana
  • Python
  • Node.js
  • JavaScript
  • Flask
  • PostgreSQL
  • Redis
  • NFS
  • Jenkins
  • Git

Responsibilities

  • Design and develop resilient and highly available distributed software
  • Develop cloud-based deployment workflows for AI inference software
  • Develop Python scripts and APIs for data preprocessing and inference workflows
  • Use parallel programming to improve AWS compute-resource efficiency
  • Develop performance visualization and analysis components
  • Develop containerized inference software and Kubernetes orchestration strategies
  • Automate failure detection and mitigation
  • Debug model deployment, orchestration, and networking issues
  • Triage and resolve defects using logs, metrics, and distributed traces
  • Define inference service interface requirements
  • Author technical documentation for infrastructure, workflows, and APIs
  • Track defects, enhancements, and release notes