Member of Technical Staff Data
InceptionVisit Inception website
AI research and product company building diffusion-based language models for production applications.
Palo Alto, United States
Funding history
About Inception
Inception develops and deploys the Mercury family of diffusion LLMs, which generate and refine output in parallel to target lower-latency, lower-cost production AI workloads.
Skills
About the Role
You will develop data mixes for LLM training and build pipelines that process petabyte-scale datasets. You will create systems for web crawling, ingestion, storage, retrieval, and versioning; evaluate data quality and diversity; and ensure data collection follows privacy regulations.
Requirements
- BS, MS, or PhD in Computer Science, Machine Learning, or a related field, or equivalent experience
- 3 or more years of experience building data-processing pipelines at scale for AI or ML applications
- Proficiency in Python and experience with Apache Spark, Beam, and Airflow
- Familiarity with synthetic data generation and data augmentation
- Familiarity with web scraping, crawling technologies, and Common Crawl datasets
- Understanding of machine-learning fundamentals and experience with PyTorch or TensorFlow
- Experience with SQL and NoSQL databases
Responsibilities
- Develop data mixes for LLM training using open-source datasets, synthetic data, and curated human feedback
- Design and implement data pipelines for petabyte-scale datasets
- Build systems for web crawling, data ingestion, and real-time data processing
- Develop tools and frameworks for data storage, retrieval, and versioning across distributed systems
- Create evaluation frameworks for data diversity, quality, and representativeness
- Ensure data collection adheres to privacy regulations
Benefits
- Equity
- Flexible vacation and paid time off
- Health, dental, and vision insurance
- 401k match
- Catered meals
- Commuter subsidies
