AI Behavior Researcher - Human Impacts
Transluce is an independent San Francisco 501(c)(3) nonprofit research lab building open technology and research infrastructure for scalable oversight and understanding of AI systems.
Funding history
About Transluce
Transluce develops research, platforms, and open-source tools intended to help evaluators and other stakeholders understand, measure, and steer advanced AI behavior in the public interest. Its active product, Docent, analyzes AI-agent transcripts using traceable behavior rubrics, qualitative review, and quantitative analysis.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will lead projects to build scientifically valid automated evaluations of frontier AI systems. You will study AI behaviors affecting user autonomy and wellbeing, develop user simulations and judge rubrics, improve evaluation realism for specific populations, and productionize evaluation practices.
Requirements
- Expertise in quantitative generative AI evaluation and measurement
- Experience designing and validating automated AI evaluation methods
- Python proficiency
- Experimental design
- Ability to balance the needs of AI researchers, domain experts, and decision makers
- Communication skills
Responsibilities
- Develop automated evaluations of AI impacts on users
- Write code for automated evaluations, including user simulators and LLM-as-a-judge pipelines
- Improve the ecological validity and realism of evaluations for specific populations
- Write and revise judge rubrics for model behaviors affecting wellbeing and decision making
- Productionize best practices in AI behavioral evaluation
Benefits
- International visa sponsorship
