Research Engineer - Data Infrastructure

AI voice technology company building lifelike speech synthesis and audio AI models.

About ElevenLabs

ElevenLabs is an AI voice company (elevenlabs.io) developing lifelike text-to-speech, voice cloning, dubbing and audio AI models used by millions of developers and enterprises.

View jobs by ElevenLabs

Skills

About the Role

You will build large scale pipelines to collect, process, filter, and transform datasets for model training. You will train classifiers and quality filters, design data curation strategies, evaluate how data affects model outcomes, and create reliable tooling for working with massive datasets.

Requirements

  • No formal certifications or degrees required
  • Experience building data intensive systems for machine learning training pipelines
  • Strong engineering skills in distributed data processing at scale
  • Experience with Kubernetes or custom pipelines over large datasets
  • Ability to evaluate data quality composition and curation independently
  • Ability to build tooling to measure data effects on model outcomes
  • Web crawler experience is a bonus

Responsibilities

  • Build large scale data pipelines for dataset collection processing filtering and transformation
  • Train classifiers quality filters and labeling models for data processing
  • Design deduplication quality scoring labeling and augmentation strategies
  • Evaluate how data quality composition and curation affect model outcomes
  • Create tooling and infrastructure for exploring and training on massive datasets
  • Build or operate web crawlers

Benefits

  • Annual discretionary professional development stipend
  • Annual discretionary social travel stipend
  • Annual company offsite
  • Monthly co-working stipend