Research Scientist Data

Pika is an active AI creative platform for generating and editing video, images, and audio, with a creator-facing app and an API offering.

Palo Alto, CA, United States

Funding history

About Pika

Pika, operated by Mellis, Inc., provides a generative-AI platform through its websites, mobile apps, APIs, and third-party platforms. Its current product covers video, image, and audio creation/editing and offers access to first-party Pika and third-party models through an aggregated API platform.

View jobs by Pika

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will own the architecture and implementation of large-scale data pipelines for multimodal model training. You will curate and manage text, image, audio, and video datasets; build ingestion, labeling, filtering, augmentation, and storage tools; ensure data quality and compliance; and optimize distributed data processing and delivery.

Requirements

  • 5+ years of experience building and scaling machine-learning data pipelines at staff or lead engineer level
  • Data engineering and ML data-curation experience for LLMs, VLMs, or multimodal models
  • Expertise in distributed data systems such as Spark, Hadoop, or Ray
  • Large-dataset processing and ETL workflow experience
  • Experience building scalable production-grade ML data infrastructure
  • Experience with data labeling, filtering, deduplication, quality assurance, and dataset management
  • Python, SQL, or PySpark programming skills
  • Familiarity with AWS, GCP, or Azure
  • Knowledge of privacy, compliance, ethics, and data-management practices

Responsibilities

  • Own large-scale data pipeline architecture and implementation for model training and research workflows
  • Curate, clean, and manage multimodal datasets
  • Develop scalable tools for data ingestion, labeling, filtering, augmentation, and storage
  • Ensure data quality, reliability, privacy, ethical considerations, and compliance
  • Optimize data processing, transformation, and delivery for distributed training
  • Prototype and productionize methods for dataset creation, management, and improvement
  • Integrate research-driven data advances into production-ready systems
  • Apply emerging data engineering and ML data-management practices

Benefits

  • Equity
  • Health benefits
  • 401k matching
  • Flexible onsite/remote hybrid work
Research Scientist Data at Pika | JobStash