Member of Technical Staff - Web Crawl Engineer
Reflection is an AI research lab building open frontier models and a full AI stack for developers, enterprises, and public-sector users.
Funding history
About Reflection
Reflection develops open-weight AI models, open-source software for customizing and running agents, AI-factory infrastructure, and related solutions. Its current research emphasizes large language models, reinforcement learning, and agentic reasoning.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will build and operate systems for web-scale crawling, URL discovery, scheduling, content extraction, and dataset delivery. You will optimize crawl quality and coverage, develop specialized crawlers, create monitoring and reliability systems, analyze performance, and resolve production issues.
Requirements
- Experience building large-scale web crawling, search indexing, content acquisition, or internet-scale data collection systems
- Understanding of crawling architectures, URL frontier management, scheduling, and distributed crawl coordination
- Experience with Ray, Spark, Beam, Flink, or similar distributed systems
- Familiarity with content extraction, HTML parsing, browser automation, rendering systems, and web technologies
- Experience operating petabyte-scale data-processing systems
- Systems engineering skills in reliability, observability, performance optimization, and debugging
- Experience designing experiments to improve crawl quality, coverage, and efficiency
- Communication skills
Responsibilities
- Build and operate web-scale crawling infrastructure
- Design and optimize URL discovery, prioritization, scheduling, and crawl orchestration
- Develop distributed crawlers that respect site constraints and operational requirements
- Build content extraction, rendering, parsing, and normalization systems
- Improve crawl coverage, freshness, efficiency, and quality through experimentation
- Design recrawling, change-detection, and incremental-update infrastructure
- Develop specialized crawlers for high-value domains and dynamic websites
- Analyze crawl performance and web coverage
- Build observability, monitoring, and reliability systems
- Debug production issues and improve crawling infrastructure
Benefits
- Stock options
- Medical, dental, vision, and life insurance
- Annual wellness allowance
- Daily office lunch and dinner
- 22 weeks of paid parental leave
- Unlimited paid time off in the U.S.
- 30 vacation days in the U.K.
- Visa sponsorship support
- Regular off-sites, happy hours, and team celebrations
