Member of Technical Staff Embedded Assessments
METR (Model Evaluation & Threat Research) is an independent nonprofit research organization that evaluates frontier AI systems’ autonomous capabilities and catastrophic-risk potential.
Funding history
Investors
About METR
METR develops scientific methods and runs empirical evaluations of frontier AI systems, including autonomous-task capability, evaluation-integrity behavior, and mitigations, to inform public and policy decision-making about AI risks.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will conduct embedded exercises at frontier AI labs for several weeks at a time. Between exercises, you will develop assessment methods, build tooling, coordinate future exercises, write results, and help scale the work. You may work on AI R&D acceleration assessment, compute allocation, monitorability red-teaming, and incident investigations.
Requirements
- At least two of the listed skills
- LLM prompting skills
- Familiarity with relevant research or ability to quickly learn it
- Security experience
- Ability to quickly understand large codebases
- Loss-of-control threat modeling and safety-case analysis
- Verbal communication, writing, and stakeholder management skills
Responsibilities
- Conduct embedded exercises at frontier AI labs
- Develop assessment methodologies
- Build tooling for future exercises
- Write results and coordinate future exercises
- Contribute to AI R&D acceleration assessments, compute allocation, monitorability red-teaming, and incident investigations
