Data Scientist (AI Data & LLM Specialist)

Eclipse Collective operates Eclipse, an Ethereum Layer 2 using the Solana Virtual Machine for high-throughput execution.

Grand Cayman, KY
About Eclipse Collective

Eclipse Collective operates Eclipse, described as Ethereum’s first SVM Layer 2. Eclipse combines Solana’s parallelized execution environment with Ethereum settlement and security, and uses Celestia for data availability. The organization provides a wallet-connected bridge, developer documentation and tools, ecosystem resources, and services for users, dApp developers, and node operators.

View jobs by Eclipse Collective

Skills

About the Role

You will establish and document scalable data annotation strategies and enforce quality metrics including inter-annotator agreement. You will research and prototype optimal data formats and preprocessing for fine-tuning LLMs, prepare datasets for tokenization, embedding generation, and NER, and build automated quality analysis and feedback loops. You will use APIs and SDKs to automate annotation and active learning and collaborate with engineers to implement data processing pipelines, producing clear documentation for technical and non-technical audiences.

Requirements

  • Proven experience as a Data Scientist or Machine Learning Engineer focused on data quality and preparation
  • Strong understanding of data labeling methodologies and experience with data annotation platforms and workflows
  • Experience preparing datasets for training and fine-tuning LLMs, including tokenization, embedding generation, and NER
  • Proficiency in Python and data science libraries such as Pandas, NumPy, and Scikit-learn
  • Experience with spaCy and Hugging Face
  • Experience using APIs and SDKs to automate data annotation and active learning loops
  • Excellent communication skills and ability to create clear documentation for technical and non-technical audiences
  • Nice-to-have: experience with audio data processing
  • Nice-to-have: familiarity with data annotation platforms and tools
  • Nice-to-have: knowledge of modern MLOps principles and practices
  • Nice-to-have: experience with RLHF and large language model data curation

Responsibilities

  • Design and document formal data annotation strategies and guidelines
  • Define and enforce quality metrics including inter-annotator agreement
  • Research, define, and prototype data formats and preprocessing for LLM training and fine-tuning
  • Establish automated processes and metrics for data quality analysis and feedback
  • Use APIs and SDKs to automate data annotation and active learning loops
  • Collaborate with engineering to guide implementation of data processing pipelines

Benefits

  • Flexible work schedule with synchronous and asynchronous collaboration
  • Quarterly in-person meetups
  • Equity
  • Benefits package
Data Scientist (AI Data & LLM Specialist) at Eclipse Collective | JobStash