Member of Technical Staff Model Efficiency

Toronto-founded enterprise AI company building foundation models, agentic AI, enterprise search, and private deployment products.

Toronto, Canada
About Cohere

Cohere provides enterprise-focused AI models and end-to-end products, including North for agentic workplace AI, Compass for enterprise search, and models for generation and retrieval. It supports private, VPC, and on-premises deployment options.

View jobs by Cohere

Skills

About the Role

You will improve LLM inference performance across the inference stack. You will identify execution bottlenecks, develop and measure optimizations, and ship improvements to latency, throughput, and quality. You will collaborate on GPU, CUDA, kernel-level, MoE, and large-scale model execution techniques.

Requirements

  • 5+ years of experience writing high-performance production-quality code
  • Programming skills in C++ or Python
  • Experience with large language models and the LLM inference ecosystem
  • Ability to diagnose and resolve performance bottlenecks across the model execution stack

Responsibilities

  • Improve core inference performance metrics across the model execution stack
  • Identify performance bottlenecks and develop optimizations
  • Experiment, measure, and ship inference improvements
  • Develop expertise in GPU, CUDA, kernel-level, MoE, and model execution optimization

Benefits

  • Weekly lunch stipend of $75/£75 or local-currency equivalent
  • Health and dental benefits
  • Mental health budget
  • RRSP matching, 401K, and pension scheme
  • Parental leave top-up for up to 6 months
  • Arts, culture, fitness, wellness, quality-time, and workspace-improvement benefits
  • Education and learning stipend
  • 6 weeks of paid vacation
  • Remote travel budget and annual company offsite
  • Co-working benefit
  • $500 home office stipend