ML Research Engineer Safeguards

AI safety and research company building reliable, interpretable, and steerable AI systems, including the Claude product family and developer platform.

Series F+Recently funded0 current maintainers0 active leadsTeam intelligence

Maintainer signals as of 9/24/2026

San Francisco, United States
About Anthropic

Anthropic PBC develops frontier AI systems and deploys them through Claude products and the Claude Platform, with a stated focus on safety, interpretability, and steerability.

View jobs by Anthropic

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will develop classifiers and synthetic-data pipelines to identify misuse and anomalous behavior at scale. You will monitor coordinated harms, evaluate agentic-product safety, test prompt-injection risks, deploy mitigations, and research automated red-teaming and adversarial robustness.

Requirements

  • Machine learning engineering
  • Research engineering
  • Applied research
  • Python
  • Machine learning system
  • Research-to-deployment
  • AI safety
  • Misuse mitigation

Responsibilities

  • Develop classifiers to detect misuse and anomalous behavior at scale
  • Develop synthetic-data pipelines and representative evaluations for classifier training
  • Build systems to monitor coordinated harms across multiple exchanges
  • Develop methods to aggregate and analyze signals across contexts
  • Evaluate and improve the safety of agentic products
  • Develop threat models and test environments for agentic risks
  • Develop and deploy mitigations for prompt injection attacks
  • Conduct research on automated red-teaming and adversarial robustness

Benefits

  • Optional equity donation matching
  • Generous vacation
  • Parental leave
  • Flexible working hours