Data Scientist (AI Data & LLM Specialist)
Eclipse Collective supports Eclipse, an Ethereum SVM layer-2 network combining Solana’s parallelized execution environment with Ethereum’s liquidity and settlement. It provides a bridge, developer tools, documentation, and resources for users and developers building or using decentralized applications.
About Eclipse Collective
Eclipse Collective operates services around Eclipse, described as Ethereum’s first SVM L2. The network uses the Solana Virtual Machine for high-throughput parallel execution, Ethereum for settlement and security, and Celestia for data availability. Its website provides a wallet-connected bridge, developer tools and documentation, and an ecosystem for users and developers to build, deploy, and use decentralized applications and nodes on the Eclipse protocol.
Skills
About the Role
You will establish and document scalable data annotation strategies and enforce quality metrics including inter-annotator agreement. You will research and prototype optimal data formats and preprocessing for fine-tuning LLMs, prepare datasets for tokenization, embedding generation, and NER, and build automated quality analysis and feedback loops. You will use APIs/SDKs to automate annotation and active learning and collaborate with engineers to implement data processing pipelines, producing clear documentation for technical and non-technical audiences.
Requirements
- Proven experience as a Data Scientist or Machine Learning Engineer focused on data quality and preparation
- Strong understanding of data labeling methodologies and experience with data annotation platforms and workflows
- Experience preparing datasets for training and fine-tuning LLMs, including tokenization, embedding generation, and NER
- Proficiency in Python and data science libraries such as Pandas, NumPy, and Scikit-learn
- Experience with spaCy and Hugging Face
- Experience using APIs and SDKs to automate data annotation and active learning loops
- Excellent communication skills and ability to create clear documentation for technical and non-technical audiences
- Nice-to-have: experience with audio data processing
- Nice-to-have: familiarity with data annotation platforms and tools
- Nice-to-have: knowledge of modern MLOps principles and practices
- Nice-to-have: experience with RLHF and large language model data curation
Responsibilities
- Design and document formal data annotation strategies and guidelines
- Define and enforce quality metrics including inter-annotator agreement
- Research, define, and prototype data formats and preprocessing for LLM training and fine-tuning
- Establish automated processes and metrics for data quality analysis and feedback
- Use APIs and SDKs to automate data annotation and active learning loops
- Collaborate with engineering to guide implementation of data processing pipelines
Benefits
- Flexible work schedule with synchronous and asynchronous collaboration
- Quarterly in-person meetups
- Equity
- Benefits package
