Research Scientist Interpretability

AI safety and research company building reliable, interpretable, and steerable AI systems, including the Claude product family and developer platform.

Series F+Recently funded0 current maintainers0 active leadsTeam intelligence

Maintainer signals as of 9/24/2026

San Francisco, United States
About Anthropic

Anthropic PBC develops frontier AI systems and deploys them through Claude products and the Claude Platform, with a stated focus on safety, interpretability, and steerability.

View jobs by Anthropic

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will develop methods for understanding language models by reverse engineering learned algorithms. You will design experiments, analyze interpretability features and circuits, build experiment and visualization infrastructure, and communicate findings internally and publicly.

Requirements

  • Scientific research
  • Interpretability
  • Python
  • Experimentation
  • Neural network
  • Language model
  • Research infrastructure
  • Data visualization

Responsibilities

  • Develop methods for understanding language models
  • Design and run experiments in toy scenarios and large models
  • Create and analyze interpretability features and circuits
  • Build infrastructure for experiments and result visualization
  • Communicate results internally and publicly

Benefits

  • Equity donation matching
  • Generous vacation
  • Parental leave
  • Flexible working hours