Performance Engineer

AI safety and research company building reliable, interpretable, and steerable AI systems, including the Claude product family and developer platform.

San Francisco, United States
About Anthropic

Anthropic PBC develops frontier AI systems and deploys them through Claude products and the Claude Platform, with a stated focus on safety, interpretability, and steerability.

View jobs by Anthropic

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will identify systems bottlenecks and develop performance solutions for large-scale machine learning workloads. You may implement high-throughput sampling, GPU kernels, load balancing, performance models, and fault-tolerant distributed systems, and debug kernel-level network latency in containerized environments.

Requirements

  • Significant software engineering or machine learning experience at supercomputing scale
  • High-performance large-scale machine learning systems
  • GPU or accelerator programming
  • Machine learning framework internals
  • Operating system internals
  • Transformer language modeling
  • Bachelor’s degree or equivalent education, training, or experience

Responsibilities

  • Identify performance problems in large-scale machine learning systems
  • Develop systems that improve throughput and robustness
  • Implement low-latency, high-throughput language-model sampling
  • Implement GPU kernels for low-precision inference
  • Develop load-balancing algorithms for serving efficiency
  • Build quantitative system-performance models
  • Design fault-tolerant distributed systems
  • Debug kernel-level network latency in containerized environments

Benefits

  • Visa sponsorship support
  • Equity donation matching
  • Generous vacation
  • Parental leave
  • Flexible working hours
Performance Engineer at Anthropic | JobStash