Staff Software Engineer Safeguards
1 month agoLeadSalary: 320K - 485KSan Francisco, CA | New York City, NYHybridCybersecurityJobs by Anthropic
AnthropicVisit Anthropic website
AI safety and research company building reliable, interpretable, and steerable AI systems, including the Claude product family and developer platform.
Maintainer signals as of 9/24/2026
San Francisco, United States
Funding history
About Anthropic
Anthropic PBC develops frontier AI systems and deploys them through Claude products and the Claude Platform, with a stated focus on safety, interpretability, and steerability.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will build monitoring, abuse-detection, and enforcement systems for AI products. You will surface abuse patterns for model hardening, create analyst dashboards, and develop reliable multilayered safety defenses that operate at scale across product and API surfaces.
Requirements
- Use Python and TypeScript
- Work across the stack
- Explain complex technical concepts to non-technical stakeholders
Responsibilities
- Develop monitoring systems to detect unwanted behavior and support automated enforcement
- Build internal dashboards for analyst review
- Build abuse-detection mechanisms and infrastructure
- Surface abuse patterns to research teams for model hardening
- Build reliable multilayered safety defenses that operate at scale
Benefits
- Optional equity donation matching
- Generous vacation
- Parental leave
- Flexible working hours
