Engineering Manager, Inference Infrastructure

AI safety and research company building reliable, interpretable, and steerable AI systems, including the Claude product family and developer platform.

San Francisco, United States
About Anthropic

Anthropic PBC develops frontier AI systems and deploys them through Claude products and the Claude Platform, with a stated focus on safety, interpretability, and steerability.

View jobs by Anthropic

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will lead the engineering group responsible for inference fleet coordination and the request path to models. You will set technical strategy for placement, load balancing, capacity, caching, and control-plane systems; improve throughput, latency, utilization, and cost; and run operational practices including on-call and incident response.

Requirements

  • Engineering-management experience for critical-path production infrastructure
  • Large-scale systems knowledge including load balancing, scheduling, orchestration, autoscaling, distributed state, or networking
  • Performance and efficiency optimization experience
  • Production operations experience including on-call, incident response, capacity events, and deployment discipline
  • Cross-functional relationship building
  • Interest in machine-learning systems and transformer inference
  • 5+ years of engineering-management experience
  • LLM inference serving knowledge
  • Cluster scheduler, autoscaler, load balancer, service-mesh, or fleet-control-plane experience
  • Multi-cloud workload experience
  • Heterogeneous accelerator-fleet knowledge
  • Bachelor’s degree or equivalent education, training, or experience

Responsibilities

  • Own the technical roadmap for inference-fleet coordination, traffic routing, capacity, caching, and control-plane protocols
  • Partner with product, inference-engine, performance, and capacity teams to ship measurable improvements
  • Build quantitative modeling practices for system changes
  • Set control-plane strategy across heterogeneous hardware, cloud providers, and serving surfaces
  • Run on-call rotations, incident response, postmortem review, and deployment safety
  • Create alignment between API, inference-engine, capacity-planning, and cloud-deployment teams
  • Develop, retain, hire, and coach engineers
  • Shape team structure and grow technical leads
  • Unblock critical initiatives and synthesize design debates

Benefits

  • Visa sponsorship assistance
  • Optional equity donation matching
  • Generous vacation
  • Parental leave
  • Flexible working hours
Engineering Manager, Inference Infrastructure at Anthropic | JobStash