Performance Engineer
AI safety and research company building reliable, interpretable, and steerable AI systems, including the Claude product family and developer platform.
Funding history
About Anthropic
Anthropic PBC develops frontier AI systems and deploys them through Claude products and the Claude Platform, with a stated focus on safety, interpretability, and steerability.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will identify systems bottlenecks and develop performance solutions for large-scale machine learning workloads. You may implement high-throughput sampling, GPU kernels, load balancing, performance models, and fault-tolerant distributed systems, and debug kernel-level network latency in containerized environments.
Requirements
- Significant software engineering or machine learning experience at supercomputing scale
- High-performance large-scale machine learning systems
- GPU or accelerator programming
- Machine learning framework internals
- Operating system internals
- Transformer language modeling
- Bachelor’s degree or equivalent education, training, or experience
Responsibilities
- Identify performance problems in large-scale machine learning systems
- Develop systems that improve throughput and robustness
- Implement low-latency, high-throughput language-model sampling
- Implement GPU kernels for low-precision inference
- Develop load-balancing algorithms for serving efficiency
- Build quantitative system-performance models
- Design fault-tolerant distributed systems
- Debug kernel-level network latency in containerized environments
Benefits
- Visa sponsorship support
- Equity donation matching
- Generous vacation
- Parental leave
- Flexible working hours
