Staff Software Engineer Node Infra
AI safety and research company building reliable, interpretable, and steerable AI systems, including the Claude product family and developer platform.
Maintainer signals as of 9/23/2026
Funding history
About Anthropic
Anthropic PBC develops frontier AI systems and deploys them through Claude products and the Claude Platform, with a stated focus on safety, interpretability, and steerability.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will lead the technical strategy for accelerator node lifecycle management, from provisioning and cluster bring-up through health checking and automated repair. You will design reliable infrastructure, drive cross-team cluster initiatives, improve operational practices, and mentor engineers.
Requirements
- Distributed systems
- Reliability engineering
- Cloud platform
- Kubernetes
- Infrastructure as code
- AWS
- GCP
- Azure
- Rust
- Go
- Python
- Terraform
- Machine learning accelerator
- Technical leadership
- Stakeholder communication
- Software engineering experience
Responsibilities
- Own the technical strategy and roadmap for node lifecycle management
- Drive initiatives to build and scale AI clusters across clouds and accelerator families
- Design and operate automated hardware health detection, isolation, and remediation systems
- Define infrastructure architecture and solve complex technical problems
- Shape compute, data, and infrastructure strategy with providers and internal stakeholders
- Establish operational excellence practices, including incident response, postmortems, and on-call
- Mentor and coach engineers
Benefits
- Visa sponsorship efforts
- Equity donation matching
- Generous vacation
- Parental leave
- Flexible working hours
