Safeguards Enforcement Analyst Safety Evaluations

AI safety and research company building reliable, interpretable, and steerable AI systems, including the Claude product family and developer platform.

Series F+Recently funded0 current maintainers0 active leadsTeam intelligence

Maintainer signals as of 9/23/2026

San Francisco, United States
About Anthropic

Anthropic PBC develops frontier AI systems and deploys them through Claude products and the Claude Platform, with a stated focus on safety, interpretability, and steerability.

View jobs by Anthropic

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will run and monitor model evaluations, interpret findings, surface regressions, and help drive mitigations. You will coordinate new and existing evaluations with policy and domain experts, improve evaluation quality and tooling, develop product-specific evaluation processes, and maintain rigorous documentation.

Requirements

  • Trust and safety
  • Content operations
  • Policy enforcement
  • Process development
  • Program management
  • Project management
  • AI-assisted workflow
  • Prioritization
  • Documentation
  • Stakeholder management

Responsibilities

  • Run evaluations, monitor results, and surface regressions or unexpected behavior changes
  • Coordinate evaluation approaches and creation with policy and domain experts
  • Interpret evaluation outcomes and drive mitigations
  • Build processes and evaluation paradigms that maintain high-quality evaluations
  • Develop product-specific evaluation frameworks
  • Help design and scope evaluation-tooling improvements
  • Write and maintain evaluation documentation

Benefits

  • Visa sponsorship
  • Equity donation matching
  • Generous vacation
  • Parental leave
  • Flexible working hours
Safeguards Enforcement Analyst Safety Evaluations at Anthropic | JobStash