Large Language Model Data Engineer
Baichuan AI (Baichuan Intelligent Technology)Visit Baichuan AI (Baichuan Intelligent Technology) website
Beijing-based AI company that develops large language models and AI-native products, including hosted model APIs and the Baixiaoying assistant.
Baichuan AI (Baichuan Intelligent Technology) on GitHubBaichuan AI (Baichuan Intelligent Technology) on Documentation
Beijing, China
Funding history
About Baichuan AI (Baichuan Intelligent Technology)
Baichuan Intelligent Technology builds general-purpose and medical large-language models, provides developer-facing API services, and operates AI-native applications.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will build and optimize distributed crawler workflows covering scheduling, collection, parsing, and storage. You will develop targeted collection processes and domain knowledge bases, contribute to retrieval-augmented generation and search algorithms, and improve the stability and timeliness of data throughout the pipeline.
Requirements
- Bachelor's degree or above
- At least 2 years of web crawling or big-data processing experience
- HTTP and TCP knowledge
- Web crawling technology and tools including Fiddler and Scrapy
- Hadoop, Spark, Hive, and HBase
- Python, Java, or Go
- Web crawler anti-bot technology
- APK unpacking and reverse engineering is preferred
- APK or mini-program data collection experience is preferred
- PyTorch or TensorFlow is preferred
- Deep learning, machine learning, and natural language processing knowledge is preferred
- Big-data framework internals and data-processing bottleneck optimization are preferred
- Big-data component operations and maintenance ability is preferred
Responsibilities
- Build and optimize distributed crawler systems
- Optimize data scheduling, collection, parsing, and storage workflows
- Develop targeted data collection processes and domain knowledge bases
- Contribute to content understanding, retrieval, and ranking algorithms for RAG systems
- Clean, process, and analyze data to improve end-to-end data stability and timeliness
