Engineering Manager, Inference Infrastructure
AI safety and research company building reliable, interpretable, and steerable AI systems, including the Claude product family and developer platform.
Funding history
About Anthropic
Anthropic PBC develops frontier AI systems and deploys them through Claude products and the Claude Platform, with a stated focus on safety, interpretability, and steerability.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will lead the engineering group responsible for inference fleet coordination and the request path to models. You will set technical strategy for placement, load balancing, capacity, caching, and control-plane systems; improve throughput, latency, utilization, and cost; and run operational practices including on-call and incident response.
Requirements
- Engineering-management experience for critical-path production infrastructure
- Large-scale systems knowledge including load balancing, scheduling, orchestration, autoscaling, distributed state, or networking
- Performance and efficiency optimization experience
- Production operations experience including on-call, incident response, capacity events, and deployment discipline
- Cross-functional relationship building
- Interest in machine-learning systems and transformer inference
- 5+ years of engineering-management experience
- LLM inference serving knowledge
- Cluster scheduler, autoscaler, load balancer, service-mesh, or fleet-control-plane experience
- Multi-cloud workload experience
- Heterogeneous accelerator-fleet knowledge
- Bachelor’s degree or equivalent education, training, or experience
Responsibilities
- Own the technical roadmap for inference-fleet coordination, traffic routing, capacity, caching, and control-plane protocols
- Partner with product, inference-engine, performance, and capacity teams to ship measurable improvements
- Build quantitative modeling practices for system changes
- Set control-plane strategy across heterogeneous hardware, cloud providers, and serving surfaces
- Run on-call rotations, incident response, postmortem review, and deployment safety
- Create alignment between API, inference-engine, capacity-planning, and cloud-deployment teams
- Develop, retain, hire, and coach engineers
- Shape team structure and grow technical leads
- Unblock critical initiatives and synthesize design debates
Benefits
- Visa sponsorship assistance
- Optional equity donation matching
- Generous vacation
- Parental leave
- Flexible working hours
