AI Behavior Researcher - Agent Alignment

Transluce is an independent San Francisco 501(c)(3) nonprofit research lab building open technology and research infrastructure for scalable oversight and understanding of AI systems.

San Francisco, United States
About Transluce

Transluce develops research, platforms, and open-source tools intended to help evaluators and other stakeholders understand, measure, and steer advanced AI behavior in the public interest. Its active product, Docent, analyzes AI-agent transcripts using traceable behavior rubrics, qualitative review, and quantitative analysis.

View jobs by Transluce

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will lead projects that design and develop scientifically valid automated evaluations of frontier AI systems. You will analyze agentic honesty and alignment behaviors, write evaluation code, improve evaluation validity, and collaborate on policy-focused and production-ready evaluation practices.

Requirements

  • Expertise in quantitative generative AI evaluation and measurement
  • Experience designing and validating automated AI evaluation methods
  • Python proficiency
  • Experimental design
  • Ability to iterate quickly while balancing scrappiness and thoroughness
  • Communication skills

Responsibilities

  • Develop automated evaluations of AI agent honesty and alignment
  • Write code for automated evaluations, including environment simulators and LLM-as-a-judge pipelines
  • Design methods to improve the ecological validity of automated evaluations
  • Measure how effects change as models become more capable
  • Collaborate on high-impact evaluations for public policy
  • Productionize best practices in AI behavior evaluation

Benefits

  • International visa sponsorship