Task Development Engineer
METR (Model Evaluation & Threat Research) is an independent nonprofit research organization that evaluates frontier AI systems’ autonomous capabilities and catastrophic-risk potential.
Funding history
Investors
About METR
METR develops scientific methods and runs empirical evaluations of frontier AI systems, including autonomous-task capability, evaluation-integrity behavior, and mitigations, to inform public and policy decision-making about AI risks.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will create difficult, well-scoped evaluation tasks that remain challenging as model capabilities grow. You will verify task solvability and specifications, baseline and score AI or human completions where useful, and identify and improve inefficient or low-quality task-development workflows.
Requirements
- Software engineering experience with complex projects and codebases
- Experience building difficult AI evaluations
- High attention to detail
- Ability to identify misspecifications and ambiguity
Responsibilities
- Develop difficult and novel tasks for models
- Build well-scoped tasks that remain challenging as model time horizons grow
- Verify that existing tasks are solvable and correctly specified
- Baseline tasks and score AI or human task completions when helpful
- Improve task-development processes and infrastructure
