Staff Engineer Datacenter Server Lifecycle

AI safety and research company building reliable, interpretable, and steerable AI systems, including the Claude product family and developer platform.

San Francisco, United States
About Anthropic

Anthropic PBC develops frontier AI systems and deploys them through Claude products and the Claude Platform, with a stated focus on safety, interpretability, and steerability.

View jobs by Anthropic

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will own server lifecycle strategy from provisioning through decommissioning. You will build fleet automation and operational tooling, maintain lifecycle procedures, enforce trusted-compute standards, coordinate networking connectivity, and track machine health, configuration, and operational status.

Requirements

  • Server hardware experience
  • Rack deployment
  • Cabling
  • Hardware troubleshooting
  • Hardware lifecycle management
  • Asset tracking
  • Provisioning
  • Maintenance scheduling
  • Decommissioning
  • Python, Rust, Go, or Java
  • Cloud infrastructure
  • Kubernetes
  • AWS, Azure, or GCP
  • Datacenter infrastructure management
  • GPU hardware
  • AI accelerator
  • LinuxBoot
  • NixOS
  • Fleet management
  • Capacity planning
  • Secure boot
  • TPM
  • Hardware attestation
  • Firmware verification

Responsibilities

  • Build automation for datacenter fleets at scale
  • Define and own server lifecycle strategy from provisioning through decommissioning
  • Maintain automation and procedures for lifecycle events
  • Design and enforce trusted compute standards across the server lifecycle
  • Ensure end-to-end connectivity across sites
  • Build and maintain fleet health, configuration, and status tooling

Benefits

  • Optional equity donation matching
  • Generous vacation
  • Parental leave
  • Flexible working hours
Staff Engineer Datacenter Server Lifecycle at Anthropic | JobStash