Member of Technical Staff Evaluation Execution
METR (Model Evaluation & Threat Research) is an independent nonprofit research organization that evaluates frontier AI systems’ autonomous capabilities and catastrophic-risk potential.
Maintainer signals as of 9/23/2026
Funding history
Investors
About METR
METR develops scientific methods and runs empirical evaluations of frontier AI systems, including autonomous-task capability, evaluation-integrity behavior, and mitigations, to inform public and policy decision-making about AI risks.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will run models on evaluation tasks, integrate them into agent scaffolds, inspect results, and improve evaluation software and infrastructure. You will communicate findings through graphs and written reports, coordinate sensitive evaluation work with stakeholders, and manage multiple concurrent evaluation projects.
Requirements
- Software engineering and infrastructure fundamentals
- Ability to debug unfamiliar systems from logs and fix performance bottlenecks
- Ability to work quickly and prioritize high-impact work
- High attention to detail
Responsibilities
- Run models on evaluation tasks and review results carefully
- Integrate models into agent scaffolds and evaluation infrastructure
- Build software that improves evaluation speed and informativeness
- Design graphs and communicate conclusions to varied audiences
- Manage concurrent evaluation projects and stakeholders
- Coordinate sensitive evaluations with leadership, lab contacts, and regulators
