Machine Learning Engineer Evals
Nous Research is an applied AI research group that develops and releases open-source AI models, datasets, and tools to democratize artificial intelligence.
Maintainer signals as of 9/2/2026
Funding history
About Nous Research
Nous Research is a leader in the development of human-centric language models and simulators, with a primary focus on model architecture, data synthesis, fine-tuning, and reasoning to align AI with real-world user experiences. As an applied research group, they aim to democratize AI development by releasing open-source resources like datasets (e.g., Hermes 3 Dataset), AI models (e.g., Hermes 3, DeepHermes), and frameworks. A key project is the Psyche Network, an open infrastructure for AI development, and they also engage in on-chain evaluations, indicating an intersection with blockchain technology. Additionally, Nous Research provides an OpenAI-compatible API for its models, operating on a pay-per-use basis with pricing determined by token consumption.
Skills
About the Role
You will operate end-to-end evaluation pipelines and reproduce known results. You will calibrate LLM judges, extend benchmarks with tasks and graders, analyze model failures, document findings, and own recurring evaluation workflows and tooling used by researchers.
Requirements
- 3+ years of experience in software engineering, ML engineering, data science, or a research-adjacent role
- Experience with an LLM evaluation framework
- Hands-on LLM experience, including prompting and few-shot design
- Python
- Git
- CI/CD
- Docker
- Linux command line
- Understanding of evaluation statistics, including Cohen's kappa and confidence intervals
- Experience with agent benchmarks, evaluation datasets, failure analysis, or non-deterministic evaluation
- Clear communication with researchers and engineers
Responsibilities
- Run the end-to-end evaluation pipeline and reproduce known results
- Build and document judge calibration protocols
- Extend benchmarks with tasks, environments, rubrics, automated graders, and quality assurance
- Analyze model-output failures and recommend training, prompt, or benchmark improvements
- Own recurring evaluation workflows and ship researcher-facing tooling
