Frontline Evaluator - Frontier AI Systems

Transluce is an independent San Francisco 501(c)(3) nonprofit research lab building open technology and research infrastructure for scalable oversight and understanding of AI systems.

0 current maintainers0 active leadsTeam intelligence

Maintainer signals as of 9/25/2026

San Francisco, United States
About Transluce

Transluce develops research, platforms, and open-source tools intended to help evaluators and other stakeholders understand, measure, and steer advanced AI behavior in the public interest. Its active product, Docent, analyzes AI-agent transcripts using traceable behavior rubrics, qualitative review, and quantitative analysis.

View jobs by Transluce

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will investigate the honesty, alignment, and unexpected behaviors of frontier AI systems. You will build and run automated evaluations, analyze large datasets and agent transcripts, identify notable patterns, validate conclusions under constraints, and turn findings into rigorous reports for technical, policy, and lab audiences. You will collaborate with internal and external technical stakeholders on evaluations, incident investigations, and model-behavior reports.

Requirements

  • Demonstrate strong empirical judgment to extract findings from data, design follow-up experiments, and identify anomalies
  • Use Python for data analysis, experiments, and evaluation tooling
  • Turn ambiguous concerns into testable questions
  • Prioritize effectively and adapt to incomplete information
  • Iterate quickly while balancing scrappiness and thoroughness
  • Communicate technical findings clearly to different audiences
  • Collaborate openly and give and receive feedback

Responsibilities

  • Develop and conduct evaluations to investigate misalignment and unexpected behaviors in AI systems
  • Write code to build and run evaluation and analysis workflows
  • Analyze agent transcripts, datasets, and evaluation results to identify notable patterns and behaviors
  • Perform investigations under time and access constraints while validating conclusions
  • Translate findings into clear, rigorous written reports
  • Collaborate with teammates and external technical stakeholders to conduct evaluations, communicate progress, and share findings

Benefits

  • International visa sponsorship