Research Scientist for Manipulation Evals
Non-profit AI safety research lab conducting frontier-AI safety research and operating remote research sprints and fellowships.
Funding history
Investors
About Apart Research
Apart Research researches manipulation and other safety issues in frontier AI systems, collaborates with AI providers and regulatory bodies, and develops AI-safety talent through research sprints, Studio, and fellowship programs.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will own the evaluation pipeline from prioritizing risk scenarios through building and running evaluations against frontier models. You will investigate model behavior, analyze datasets and transcripts, develop defensible measures, publish results, and communicate technical findings clearly for regulatory and academic audiences.
Requirements
- Use coding agents such as Claude Code and Codex to build evaluations
- Demonstrate hands-on evaluation experience and strong red-teaming instincts
- Have at least one co-authored paper at a top AI or machine-learning venue, or impactful workshop papers or preprints
- Demonstrate hands-on safety-relevant empirical work in evaluations, benchmarks, adversarial robustness, or alignment
- Demonstrate strong Python, software engineering, evaluation-framework, LLM API, and data-analysis skills
- Write clear technical material for a regulator audience
- Work independently on open-ended research mandates
Responsibilities
- Develop and prioritize harmful-manipulation risk scenarios
- Build and run end-to-end evaluations against frontier models using coding agents
- Design rigorous evaluation methods for AI manipulation
- Review relevant literature and track the state of the art
- Investigate model behavior and red-team evaluation results
- Analyze transcripts and datasets to identify manipulation patterns and quantitative measures
- Deliver new research results at least weekly
- Write technical findings for regulatory and academic audiences
- Co-author public papers, benchmark releases, and technical reports
- Engage directly with the EU AI Office
Benefits
- Fully remote work
- Travel budget for conferences, EU institutional meetings, and coworking across Europe
- Learning and development budget
- Equipment budget
- Home-office budget
