Principal Data Engineer
Gather AI is a Pittsburgh-based physical-AI company providing warehouse intelligence software using computer vision on drones and material-handling equipment.
Funding history
About Gather AI
Its Gather AI Prana platform continuously observes warehouse operations, reasons over floor and enterprise-system data, and routes workflows for dock-to-dock logistics operations.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will architect a greenfield multi-layer data warehouse and separate analytical workloads from production traffic. You will build governed data access, semantic metrics, ingestion and lineage capabilities, and secure multi-tenant pipelines. You will link structured records with drone imagery and video while establishing data infrastructure for ML use cases.
Requirements
- 10+ years in data engineering
- 3+ years architecting data platforms for data products, analytics, or AI-driven products
- Experience building a greenfield data warehouse and leading an OLTP-to-OLAP transition
- Expert SQL and dbt
- Hands-on ELT and orchestration experience
- Large-scale, distributed, or streaming data experience
- Production experience with Azure, AWS, or GCP
- Infrastructure as code and CI/CD experience
- Experience with data quality, security, governance, and multi-tenancy
- Data transformation and modeling experience
- Pipeline orchestration and workflow automation experience
- Semantic and metrics-layer design experience
- Serving-layer optimization experience
- Data observability and lineage experience
- Multimodal data integration experience
Responsibilities
- Architect a multi-layer data warehouse
- Deliver a governed self-service data-access layer
- Build semantic and metrics layers
- Own availability, freshness, traceability, pipeline quality, and data-loss standards
- Design tenant isolation, cost attribution, schemas, and row-level RBAC
- Own ingestion correctness, data contracts, schema validation, and pipeline quality
- Establish data catalog and lineage capabilities
- Prove the data foundation on the drone product and generalize it for future products
- Link structured records to drone imagery and video
- Establish infrastructure for feature stores and annotation pipelines
