Machine Learning Engineer Evals

Open-source AI company that trains models, builds AI agents, and develops distributed-training infrastructure.

Series A14 current maintainers11 active leadsTeam intelligence

Maintainer signals as of 9/24/2026

Distributed
About Nous Research

Nous Research develops open AI models, the Hermes Agent, Nous Portal, and distributed-training technology. Its current site presents an installable, open-source agent and a paid cloud/model-access service.

View jobs by Nous Research

Skills

About the Role

You will run and improve end-to-end model evaluation pipelines. You will create judge-calibration protocols, extend benchmarks with tasks and automated graders, analyze model failures, and document recommendations. You will own recurring evaluation workflows and deliver tools that researchers use.

Requirements

  • 3+ years in software engineering, ML engineering, data science, or a research-adjacent role
  • Experience with an LLM evaluation framework
  • Hands-on LLM experience, including prompting and few-shot design
  • Python
  • Git
  • CI/CD
  • Docker
  • Linux command line
  • Understanding of evaluation statistics, including accuracy, Cohen's kappa, and confidence intervals
  • Experience with at least three specified evaluation competencies
  • Communicate clearly with researchers and engineers

Responsibilities

  • Run end-to-end evaluation pipelines and reproduce known results during onboarding
  • Build and document judge-calibration protocols
  • Extend benchmarks with tasks, prompts, environments, rubrics, automated graders, and quality assurance
  • Analyze model-output failures, quantify failure modes, and recommend improvements
  • Own recurring evaluation workflows and ship evaluation tooling