Red Team Engineer Safeguards

AI safety and research company building reliable, interpretable, and steerable AI systems, including the Claude product family and developer platform.

San Francisco, United States
About Anthropic

Anthropic PBC develops frontier AI systems and deploys them through Claude products and the Claude Platform, with a stated focus on safety, interpretability, and steerability.

View jobs by Anthropic

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will conduct adversarial testing across product surfaces and simulate sophisticated threat actors. You will research new testing methods, execute full kill-chain attacks, build automated testing frameworks, and turn findings into concrete product and safety improvements.

Requirements

  • Experience in penetration testing, red teaming, or application security
  • Experience with model jailbreaking
  • Experience testing agentic workflows for prompt injection vectors
  • Web application security expertise
  • Experience with security testing tools
  • Experience building custom automation and LLM-specific testing frameworks
  • Track record of discovering and chaining novel attack vectors
  • Public security work such as CVEs, blog posts, or disclosed bug bounty reports
  • Written and verbal communication skills

Responsibilities

  • Conduct adversarial testing across product surfaces
  • Develop attack scenarios that combine exploitation techniques
  • Research and implement testing approaches for agent systems and tool use
  • Design and execute full kill-chain attacks
  • Build and maintain systematic testing methodologies
  • Develop automated testing frameworks
  • Collaborate to translate findings into improvements
  • Establish metrics for detecting novel abuse

Benefits

  • Optional equity donation matching
  • Generous vacation
  • Generous parental leave
  • Flexible working hours
Red Team Engineer Safeguards at Anthropic | JobStash