Member of Technical Staff Data

AI research and product company building diffusion-based language models for production applications.

Palo Alto, United States

Funding history

About Inception

Inception develops and deploys the Mercury family of diffusion LLMs, which generate and refine output in parallel to target lower-latency, lower-cost production AI workloads.

View jobs by Inception

Skills

About the Role

You will develop data mixes for LLM training and build pipelines that process petabyte-scale datasets. You will create systems for web crawling, ingestion, storage, retrieval, and versioning; evaluate data quality and diversity; and ensure data collection follows privacy regulations.

Requirements

  • BS, MS, or PhD in Computer Science, Machine Learning, or a related field, or equivalent experience
  • 3 or more years of experience building data-processing pipelines at scale for AI or ML applications
  • Proficiency in Python and experience with Apache Spark, Beam, and Airflow
  • Familiarity with synthetic data generation and data augmentation
  • Familiarity with web scraping, crawling technologies, and Common Crawl datasets
  • Understanding of machine-learning fundamentals and experience with PyTorch or TensorFlow
  • Experience with SQL and NoSQL databases

Responsibilities

  • Develop data mixes for LLM training using open-source datasets, synthetic data, and curated human feedback
  • Design and implement data pipelines for petabyte-scale datasets
  • Build systems for web crawling, data ingestion, and real-time data processing
  • Develop tools and frameworks for data storage, retrieval, and versioning across distributed systems
  • Create evaluation frameworks for data diversity, quality, and representativeness
  • Ensure data collection adheres to privacy regulations

Benefits

  • Equity
  • Flexible vacation and paid time off
  • Health, dental, and vision insurance
  • 401k match
  • Catered meals
  • Commuter subsidies
Member of Technical Staff Data at Inception | JobStash