Machine Learning Engineer Evals
Nous ResearchVisit Nous Research website
Open-source AI company that trains models, builds AI agents, and develops distributed-training infrastructure.
Nous Research on X (Twitter)Nous Research on DiscordNous Research on GitHubNous Research on Documentation
Maintainer signals as of 9/24/2026
Distributed
Funding history
About Nous Research
Nous Research develops open AI models, the Hermes Agent, Nous Portal, and distributed-training technology. Its current site presents an installable, open-source agent and a paid cloud/model-access service.
Skills
About the Role
You will run and improve end-to-end model evaluation pipelines. You will create judge-calibration protocols, extend benchmarks with tasks and automated graders, analyze model failures, and document recommendations. You will own recurring evaluation workflows and deliver tools that researchers use.
Requirements
- 3+ years in software engineering, ML engineering, data science, or a research-adjacent role
- Experience with an LLM evaluation framework
- Hands-on LLM experience, including prompting and few-shot design
- Python
- Git
- CI/CD
- Docker
- Linux command line
- Understanding of evaluation statistics, including accuracy, Cohen's kappa, and confidence intervals
- Experience with at least three specified evaluation competencies
- Communicate clearly with researchers and engineers
Responsibilities
- Run end-to-end evaluation pipelines and reproduce known results during onboarding
- Build and document judge-calibration protocols
- Extend benchmarks with tasks, prompts, environments, rubrics, automated graders, and quality assurance
- Analyze model-output failures, quantify failure modes, and recommend improvements
- Own recurring evaluation workflows and ship evaluation tooling
