Member of Technical Staff Evaluation Execution

METR (Model Evaluation & Threat Research) is an independent nonprofit research organization that evaluates frontier AI systems’ autonomous capabilities and catastrophic-risk potential.

Recently funded0 current maintainers0 active leadsTeam intelligence

Maintainer signals as of 9/23/2026

Berkeley, United States
About METR

METR develops scientific methods and runs empirical evaluations of frontier AI systems, including autonomous-task capability, evaluation-integrity behavior, and mitigations, to inform public and policy decision-making about AI risks.

View jobs by METR

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will run models on evaluation tasks, integrate them into agent scaffolds, inspect results, and improve evaluation software and infrastructure. You will communicate findings through graphs and written reports, coordinate sensitive evaluation work with stakeholders, and manage multiple concurrent evaluation projects.

Requirements

  • Software engineering and infrastructure fundamentals
  • Ability to debug unfamiliar systems from logs and fix performance bottlenecks
  • Ability to work quickly and prioritize high-impact work
  • High attention to detail

Responsibilities

  • Run models on evaluation tasks and review results carefully
  • Integrate models into agent scaffolds and evaluation infrastructure
  • Build software that improves evaluation speed and informativeness
  • Design graphs and communicate conclusions to varied audiences
  • Manage concurrent evaluation projects and stakeholders
  • Coordinate sensitive evaluations with leadership, lab contacts, and regulators
Member of Technical Staff Evaluation Execution at METR | JobStash