Task Development Engineer

METR (Model Evaluation & Threat Research) is an independent nonprofit research organization that evaluates frontier AI systems’ autonomous capabilities and catastrophic-risk potential.

Recently fundedCompany intelligence
Berkeley, United States
About METR

METR develops scientific methods and runs empirical evaluations of frontier AI systems, including autonomous-task capability, evaluation-integrity behavior, and mitigations, to inform public and policy decision-making about AI risks.

View jobs by METR

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will create difficult, well-scoped evaluation tasks that remain challenging as model capabilities grow. You will verify task solvability and specifications, baseline and score AI or human completions where useful, and identify and improve inefficient or low-quality task-development workflows.

Requirements

  • Software engineering experience with complex projects and codebases
  • Experience building difficult AI evaluations
  • High attention to detail
  • Ability to identify misspecifications and ambiguity

Responsibilities

  • Develop difficult and novel tasks for models
  • Build well-scoped tasks that remain challenging as model time horizons grow
  • Verify that existing tasks are solvable and correctly specified
  • Baseline tasks and score AI or human task completions when helpful
  • Improve task-development processes and infrastructure