Web Scraping Specialist
Wynd Labs provides internet-scale web data products for AI labs, business intelligence, and multimodal research.
Funding history
About Wynd Labs
Wynd Labs delivers structured web data at scale for AI development and research. Its product suite includes a Search API for real-time JS-rendered search engine results, multimodal data pipelines spanning video, text, images, and audio, and curated datasets for training and evaluation. Data can be enriched, filtered, and delivered through APIs, direct download, or cloud storage.
Skills
About the Role
Design, build, and maintain robust web scraping pipelines for extracting data from complex websites. The role involves writing and refining scraping code, handling dynamic content and pagination, cleaning and storing extracted data, monitoring scraping runs, and scaling jobs with cloud infrastructure and distributed techniques.
Requirements
- Demonstrated ability to extract data from complex websites
- Proficiency in Python or JavaScript
- Experience with BeautifulSoup, Scrapy, or Selenium
- Knowledge of asynchronous programming, multithreading, and distributed scraping
- In-depth knowledge of HTML, CSS, JavaScript, and the DOM
- Experience with NoSQL databases such as MongoDB or Cassandra
- Experience with cloud services including AWS, Google Cloud, and Azure
- Ability to apply machine learning for data cleaning or categorization
- Participation in open-source projects related to web scraping or data processing
Responsibilities
- Lead data gathering and analysis from online sources
- Write, test and refine code to extract data from websites
- Handle pagination and dynamic content loaded via AJAX
- Clean and format extracted data to meet quality standards
- Store and manage scraped data in appropriate databases
- Monitor scraping processes and resolve operational issues
- Optimize scraping processes for reliability and scale
Benefits
- Remote work
- Equity package
- Competitive salary
- Benefits package
