Staff Infrastructure Engineer Cluster Infrastructure
AI safety and research company building reliable, interpretable, and steerable AI systems, including the Claude product family and developer platform.
Maintainer signals as of 9/23/2026
Funding history
About Anthropic
Anthropic PBC develops frontier AI systems and deploys them through Claude products and the Claude Platform, with a stated focus on safety, interpretability, and steerability.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will set technical direction for compute-cluster lifecycle management. You will lead provisioning, updates, and decommissioning strategy; coordinate capacity and connectivity; establish secure-by-default clusters; improve scalability and fault tolerance; and advance incident response, on-call practices, and technical mentorship.
Requirements
- Expertise in distributed systems, reliability, cloud platforms, Kubernetes, infrastructure as code, and AWS, GCP, or Azure
- Proficiency in Rust, Go, Python, or another systems language
- Proficiency with Terraform
- Experience leading complex multi-quarter initiatives across teams or systems
- Ability to align senior stakeholders and communicate effectively
Responsibilities
- Own technical strategy and roadmap for agent-driven cluster lifecycle management
- Ensure new compute capacity is ingested on time
- Coordinate physical build-out and cloud solutions for inter-cluster connectivity
- Ensure clusters are provisioned secure-by-default
- Drive cluster scalability, homogeneity, and fault-tolerance strategy
- Shape long-term compute, data, and infrastructure strategy with partners
- Establish and evolve incident response, postmortem, and on-call practices
- Provide technical mentorship and coaching
Benefits
- Equity donation matching
- Generous vacation
- Parental leave
- Flexible working hours
