Audio Inference Engineer Model Efficiency

Toronto-founded enterprise AI company building foundation models, agentic AI, enterprise search, and private deployment products.

Toronto, Canada
About Cohere

Cohere provides enterprise-focused AI models and end-to-end products, including North for agentic workplace AI, Compass for enterprise search, and models for generation and retrieval. It supports private, VPC, and on-premises deployment options.

View jobs by Cohere

Skills

About the Role

You will improve audio-inference latency, throughput, and quality. You will investigate system bottlenecks, deliver efficient solutions for audio processing and streaming workloads, and work with training and serving infrastructure teams to integrate real-time audio-model development and deployment.

Requirements

  • Experience developing high-performance audio or machine-learning inference systems
  • Proficiency in C++ and Python
  • Experience with deep-learning models for audio, speech, or language applications

Responsibilities

  • Optimize audio-inference latency, throughput, and quality
  • Identify system bottlenecks and deliver performance solutions
  • Improve audio-processing and streaming workloads
  • Collaborate with training and serving infrastructure teams
  • Support integration between model development and deployment for real-time audio inference

Benefits

  • Equity
  • Weekly lunch stipend of $75/£75 or equivalent local currency
  • Health and dental benefits
  • Mental health budget
  • RRSP matching, 401K, and pension scheme
  • Parental leave top-up for up to 6 months
  • Arts, culture, fitness, wellness, quality-time, and workspace-improvement benefits
  • Education and learning stipend
  • 6 weeks of paid vacation
  • Remote travel budget and annual company offsite
  • Co-working benefit
  • $500 home office stipend
Audio Inference Engineer Model Efficiency at Cohere | JobStash