Staff Research Engineer Model Efficiency

Toronto-founded enterprise AI company building foundation models, agentic AI, enterprise search, and private deployment products.

Toronto, Canada
About Cohere

Cohere provides enterprise-focused AI models and end-to-end products, including North for agentic workplace AI, Compass for enterprise search, and models for generation and retrieval. It supports private, VPC, and on-premises deployment options.

View jobs by Cohere

Skills

About the Role

You will develop, prototype, and deploy techniques that improve production language-model speed and efficiency. You will work across model architecture, routing optimization, decoding, inference algorithms, software and hardware co-design, GPU acceleration, and performance optimization while preserving model quality.

Requirements

  • Machine Learning PhD
  • LLM architecture
  • LLM inference optimization
  • Model efficiency
  • Software engineering
  • ICLR
  • ACL
  • NeurIPS
  • Mentoring

Responsibilities

  • Develop, prototype, and deploy techniques that improve production model speed and efficiency

Benefits

  • Weekly lunch stipend
  • Health and dental benefits
  • Mental health budget
  • RRSP matching
  • 401K
  • Pension Scheme
  • Parental leave top-up for up to six months
  • Arts and culture credit
  • Fitness and wellness credit
  • Quality-time credit
  • Workspace improvement credit
  • Six weeks of paid vacation
  • Remote-office travel budget
  • Annual company offsite
  • Co-working benefit
  • Home office stipend
Staff Research Engineer Model Efficiency at Cohere | JobStash