Member of Technical Staff Data Infrastructure

AI research and product company building diffusion-based language models for production applications.

Palo Alto, United States

Funding history

About Inception

Inception develops and deploys the Mercury family of diffusion LLMs, which generate and refine output in parallel to target lower-latency, lower-cost production AI workloads.

View jobs by Inception

Skills

About the Role

You will architect and scale infrastructure for distributed training pipelines and petabyte-scale data catalogs. You will build data ingestion, processing, storage, retrieval, versioning, web-crawling, and real-time processing systems. You will support data quality, privacy compliance, and model-training operations.

Requirements

  • 3+ years of experience building data-processing pipelines at scale for AI/ML applications
  • Proficiency in Python
  • Experience with Apache Spark Beam or Airflow
  • Knowledge of synthetic data generation and data augmentation
  • Knowledge of web scraping crawling technologies and Common Crawl datasets
  • Understanding of machine learning fundamentals
  • Experience with PyTorch or TensorFlow
  • Experience with SQL and NoSQL databases

Responsibilities

  • Design, build, and operate scalable fault-tolerant infrastructure for LLM research
  • Develop high-throughput data ingestion processing and transformation systems
  • Build web-crawling data-ingestion and real-time processing systems
  • Develop distributed data storage retrieval and versioning tools
  • Ensure data collection complies with privacy regulations

Benefits

  • Equity
  • Flexible vacation and paid time off
  • Health dental and vision insurance
  • 401k match
  • Catered meals
  • Commuter subsidies
Member of Technical Staff Data Infrastructure at Inception | JobStash