Cyber Evaluations Engineer
7 hours agoSalary: 300K - 405KRemote-Friendly, United States; San Francisco, CA | Washington, DCHybridCybersecurityJobs by Anthropic
AnthropicVisit Anthropic website
AI safety and research company building reliable, interpretable, and steerable AI systems, including the Claude product family and developer platform.
San Francisco, United States
Funding history
About Anthropic
Anthropic PBC develops frontier AI systems and deploys them through Claude products and the Claude Platform, with a stated focus on safety, interpretability, and steerability.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will design and run cyber capability, safety, and robustness evaluations for new models. You will analyze jailbreak and prompt-bypass data, build detection probes and evaluation tooling, measure detection performance, and turn findings into safeguard improvements with policy and engineering partners.
Requirements
- Experience building or running software or ML evaluations, benchmarks, or test suites
- Hands-on cybersecurity experience such as CTF participation, vulnerability research, exploit development, or security research
- Python proficiency
- Ability to communicate evaluation results to cross-functional and policy stakeholders
Responsibilities
- Design and run capability, uplift, and safety evaluations for cyber-relevant model risk
- Execute safeguard-robustness testing before major model launches
- Analyze evaluation results and communicate findings
- Design, prototype, and tune cyber-misuse detection probes
- Develop a layered abuse-detection architecture with the cyber policy team
- Build and maintain internal evaluation and scoring tooling
- Translate evaluation findings into safeguard improvements with policy and engineering partners
Benefits
- Visa sponsorship support
- Optional equity donation matching
- Generous vacation
- Parental leave
- Flexible working hours
