Product Designer Evals & Prompts
AnthropicVisit Anthropic website
AI safety and research company building reliable, interpretable, and steerable AI systems, including the Claude product family and developer platform.
San Francisco, United States
Funding history
About Anthropic
Anthropic PBC develops frontier AI systems and deploys them through Claude products and the Claude Platform, with a stated focus on safety, interpretability, and steerability.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will write and test product prompts, build automated graders and evaluation tools, and improve evaluation workflows for designers. You will support model releases, maintain a reliable evaluation harness, investigate regressions, and package evaluation results into training signals.
Requirements
- Write production-quality Python
- Build and maintain LLM evaluation pipelines
- Build internal tools for nontechnical users
- Build test harnesses and sandbox tool calls
- Ship prompts or work closely with prompt practitioners
- Analyze model transcripts
Responsibilities
- Write, revise, test, and ship product prompts
- Build automated graders and evaluations from design rubrics
- Build low-code evaluation tools for designers
- Observe tool use and simplify evaluation workflows
- Test product surfaces and create prompt migrations for model releases
- Build and scale the evaluation harness
- Investigate whether regressions arise from the harness or model
- Package evaluation results into training signals
Benefits
- Equity donation matching
- Generous vacation
- Parental leave
- Flexible working hours
