Senior Research Scientist Model Evaluation

Toronto-founded enterprise AI company building foundation models, agentic AI, enterprise search, and private deployment products.

Toronto, Canada
About Cohere

Cohere provides enterprise-focused AI models and end-to-end products, including North for agentic workplace AI, Compass for enterprise search, and models for generation and retrieval. It supports private, VPC, and on-premises deployment options.

View jobs by Cohere

Skills

About the Role

You will create next-generation evaluation methods and infrastructure for measuring language-model progress. You will build benchmarks, translate model feedback into repeatable evaluations, research LLM judges and data-synthesis pipelines, and develop scalable tools for investigating model performance.

Requirements

  • LLM
  • LLM evaluation
  • Prototype development
  • Data quality
  • Data review
  • Software engineering

Responsibilities

  • Create evaluation benchmarks that test model capabilities
  • Translate model feedback into trustworthy and repeatable evaluations
  • Conduct research on LLM evaluation methods, LLM judges, data synthesis, and evaluation efficiency
  • Build scalable and reusable tools to investigate model performance

Benefits

  • Weekly lunch stipend
  • Health and dental benefits
  • Mental health budget
  • RRSP matching
  • 401K
  • Pension Scheme
  • Parental leave top-up for up to six months
  • Arts and culture credit
  • Fitness and wellness credit
  • Quality-time credit
  • Workspace improvement credit
  • Six weeks of paid vacation
  • Remote-office travel budget
  • Annual company offsite
  • Co-working benefit
  • Home office stipend
Senior Research Scientist Model Evaluation at Cohere | JobStash