Research Crawling Engineer
Wynd LabsVisit Wynd Labs website
Wynd Labs provides internet-scale web data products for AI labs, business intelligence, and multimodal research.
Distributed
Funding history
About Wynd Labs
Wynd Labs delivers structured web data at scale for AI development and research. Its product suite includes a Search API for real-time JS-rendered search engine results, multimodal data pipelines spanning video, text, images, and audio, and curated datasets for training and evaluation. Data can be enriched, filtered, and delivered through APIs, direct download, or cloud storage.
Skills
Bare-Metal InfrastructureBenchmarkingBrowser BehaviorCChrome Devtools ProtocolCloudData-PipelineData QualityDataset CurationDeduplicationDistributed SystemsFilteringGoHeadless BrowserHttpIp RotationJavaLlm PretrainingNetworkingNlpNormalizationParallel ProcessingPlaywrightProxy SystemPuppeteerPythonRequest OrchestrationRetrievalRustWeb Crawler
About the Role
Design, build, and operate large-scale web data acquisition systems, including distributed crawlers, anti-bot handling, dynamic-site extraction, data normalization, and pipelines for cleaning, deduplication, and dataset construction.
Requirements
- Strong programming experience in Go, Rust, Python, Java, or C++
- Experience building web crawlers or large-scale data pipelines
- Solid understanding of HTTP, networking, and browser behavior
- Familiarity with distributed systems and parallel processing
- Experience with large datasets at TB–PB scale preferred
- Ability to debug unstable or adversarial environments
Responsibilities
- Build and maintain large-scale web crawlers across diverse domains
- Design high-throughput, fault-tolerant systems for data collection
- Handle anti-bot systems, rate limits, and dynamic JavaScript-heavy sites
- Develop pipelines for cleaning, deduplication, filtering, and normalization
- Construct and maintain datasets for research and model training
- Monitor crawl performance, coverage, and data quality
- Collaborate with research teams on data collection and modeling needs
- Optimize infrastructure for cost, latency, and reliability
- Own end-to-end data acquisition pipelines
Benefits
- Benefits package
- Equity package
- Fully remote work
