Research Engineer Web Crawlers
ElevenLabsVisit ElevenLabs website
AI voice technology company building lifelike speech synthesis and audio AI models.
About ElevenLabs
ElevenLabs is an AI voice company (elevenlabs.io) developing lifelike text-to-speech, voice cloning, dubbing and audio AI models used by millions of developers and enterprises.
Skills
About the Role
You will build and operate large scale distributed web crawlers that discover fetch and extract data across billions of pages. You will solve content extraction deduplication freshness recrawling politeness and rate limit challenges while creating targeted pipelines for audio video and multilingual data. You will also develop tooling to monitor and explore crawled datasets and evaluate their quality coverage and compliance.
Requirements
- Hands on experience building and scaling web crawlers or scraping systems
- Strong distributed systems engineering skills at scale
- Experience with Kubernetes queue based architectures or custom pipelines processing billions of documents
- Ability to autonomously evaluate data quality coverage and compliance
- Evidence of solving difficult engineering problems through projects designs or GitHub contributions
Responsibilities
- Build and operate large scale distributed web crawlers
- Solve content extraction deduplication freshness recrawling politeness and rate limit challenges
- Create targeted crawling pipelines for audio video and multilingual content
- Develop tooling to request monitor and explore crawled data
- Evaluate the quality coverage and compliance of crawled data
Benefits
- Annual discretionary professional development stipend
- Annual discretionary social travel stipend
- Annual company offsite
- Monthly co-working stipend if not located near a main hub
