Member of Technical Staff - Evaluations

Reflection is an AI research lab building open frontier models and a full AI stack for developers, enterprises, and public-sector users.

New York, United States
About Reflection

Reflection develops open-weight AI models, open-source software for customizing and running agents, AI-factory infrastructure, and related solutions. Its current research emphasizes large language models, reinforcement learning, and agentic reasoning.

View jobs by Reflection

Skills

About the Role

You will conduct comparative analyses of model capabilities and build evaluation systems that link data, evaluations, and model behavior. You will develop frameworks for reasoning, alignment, and usefulness, collaborate with research partners on model improvements, and expand measurement through synthetic evaluations, human feedback, and real-world interaction data.

Requirements

  • Statistical analysis
  • Experimental design
  • LLM evaluation methodology
  • Static benchmark evaluation
  • Human preference evaluation
  • Agentic task evaluation

Responsibilities

  • Conduct comparative analysis of model capabilities
  • Build and refine evaluation systems and processes
  • Develop generalizable evaluation frameworks for reasoning alignment and usefulness
  • Collaborate with research teams to translate insights into model improvements
  • Develop measurements using synthetic evaluations human feedback and real-world interaction data

Benefits

  • Stock options
  • Medical insurance
  • Dental insurance
  • Vision insurance
  • Life insurance
  • Annual wellness allowance
  • Daily office lunch and dinner
  • 22 weeks of paid parental leave
  • Unlimited paid time off in the U.S.
  • 30 days of vacation in the U.K.
  • Visa sponsorship
  • Regular off-sites
  • Happy hours
  • Team celebrations
Member of Technical Staff - Evaluations at Reflection | JobStash