Principal Data Engineer

Gather AI is a Pittsburgh-based physical-AI company providing warehouse intelligence software using computer vision on drones and material-handling equipment.

Pittsburgh, United States
About Gather AI

Its Gather AI Prana platform continuously observes warehouse operations, reasons over floor and enterprise-system data, and routes workflows for dock-to-dock logistics operations.

View jobs by Gather AI

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will architect a greenfield multi-layer data warehouse and separate analytical workloads from production traffic. You will build governed data access, semantic metrics, ingestion and lineage capabilities, and secure multi-tenant pipelines. You will link structured records with drone imagery and video while establishing data infrastructure for ML use cases.

Requirements

  • 10+ years in data engineering
  • 3+ years architecting data platforms for data products, analytics, or AI-driven products
  • Experience building a greenfield data warehouse and leading an OLTP-to-OLAP transition
  • Expert SQL and dbt
  • Hands-on ELT and orchestration experience
  • Large-scale, distributed, or streaming data experience
  • Production experience with Azure, AWS, or GCP
  • Infrastructure as code and CI/CD experience
  • Experience with data quality, security, governance, and multi-tenancy
  • Data transformation and modeling experience
  • Pipeline orchestration and workflow automation experience
  • Semantic and metrics-layer design experience
  • Serving-layer optimization experience
  • Data observability and lineage experience
  • Multimodal data integration experience

Responsibilities

  • Architect a multi-layer data warehouse
  • Deliver a governed self-service data-access layer
  • Build semantic and metrics layers
  • Own availability, freshness, traceability, pipeline quality, and data-loss standards
  • Design tenant isolation, cost attribution, schemas, and row-level RBAC
  • Own ingestion correctness, data contracts, schema validation, and pipeline quality
  • Establish data catalog and lineage capabilities
  • Prove the data foundation on the drone product and generalize it for future products
  • Link structured records to drone imagery and video
  • Establish infrastructure for feature stores and annotation pipelines