Member of Technical Staff - Evaluations
Reflection is an AI research lab building open frontier models and a full AI stack for developers, enterprises, and public-sector users.
Funding history
About Reflection
Reflection develops open-weight AI models, open-source software for customizing and running agents, AI-factory infrastructure, and related solutions. Its current research emphasizes large language models, reinforcement learning, and agentic reasoning.
Skills
About the Role
You will conduct comparative analyses of model capabilities and build evaluation systems that link data, evaluations, and model behavior. You will develop frameworks for reasoning, alignment, and usefulness, collaborate with research partners on model improvements, and expand measurement through synthetic evaluations, human feedback, and real-world interaction data.
Requirements
- Statistical analysis
- Experimental design
- LLM evaluation methodology
- Static benchmark evaluation
- Human preference evaluation
- Agentic task evaluation
Responsibilities
- Conduct comparative analysis of model capabilities
- Build and refine evaluation systems and processes
- Develop generalizable evaluation frameworks for reasoning alignment and usefulness
- Collaborate with research teams to translate insights into model improvements
- Develop measurements using synthetic evaluations human feedback and real-world interaction data
Benefits
- Stock options
- Medical insurance
- Dental insurance
- Vision insurance
- Life insurance
- Annual wellness allowance
- Daily office lunch and dinner
- 22 weeks of paid parental leave
- Unlimited paid time off in the U.S.
- 30 days of vacation in the U.K.
- Visa sponsorship
- Regular off-sites
- Happy hours
- Team celebrations
