Member of Technical Staff Data Infrastructure
InceptionVisit Inception website
AI research and product company building diffusion-based language models for production applications.
Palo Alto, United States
Funding history
About Inception
Inception develops and deploys the Mercury family of diffusion LLMs, which generate and refine output in parallel to target lower-latency, lower-cost production AI workloads.
Skills
About the Role
You will architect and scale infrastructure for distributed training pipelines and petabyte-scale data catalogs. You will build data ingestion, processing, storage, retrieval, versioning, web-crawling, and real-time processing systems. You will support data quality, privacy compliance, and model-training operations.
Requirements
- 3+ years of experience building data-processing pipelines at scale for AI/ML applications
- Proficiency in Python
- Experience with Apache Spark Beam or Airflow
- Knowledge of synthetic data generation and data augmentation
- Knowledge of web scraping crawling technologies and Common Crawl datasets
- Understanding of machine learning fundamentals
- Experience with PyTorch or TensorFlow
- Experience with SQL and NoSQL databases
Responsibilities
- Design, build, and operate scalable fault-tolerant infrastructure for LLM research
- Develop high-throughput data ingestion processing and transformation systems
- Build web-crawling data-ingestion and real-time processing systems
- Develop distributed data storage retrieval and versioning tools
- Ensure data collection complies with privacy regulations
Benefits
- Equity
- Flexible vacation and paid time off
- Health dental and vision insurance
- 401k match
- Catered meals
- Commuter subsidies
