Software Development Engineer II Data Engineer
Gather AI is a Pittsburgh-based physical-AI company providing warehouse intelligence software using computer vision on drones and material-handling equipment.
Funding history
About Gather AI
Its Gather AI Prana platform continuously observes warehouse operations, reasons over floor and enterprise-system data, and routes workflows for dock-to-dock logistics operations.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will build a data foundation from scratch, beginning with well-scoped pipelines and models for the Drone product. You will move analytical workloads away from production systems, create shared transformation models and metrics, and ensure data remains accurate, traceable, secure, and available. You will collaborate across functions, document your work, ship through CI/CD, and join on-call for the pipelines you own.
Requirements
- 2–5 years of experience building and running production data pipelines.
- Degree in Computer Science or equivalent practical experience.
- Strong SQL skills, including joins, window functions, and CTEs.
- Understanding of dimensional modelling, including facts, dimensions, and grain.
- Experience with PostgreSQL is preferred.
- Hands-on experience building tested models in dbt, Snowflake Dynamic Tables, Databricks Lakeflow Declarative Pipelines, or equivalent.
- Production Python experience with tests and code review.
- Experience orchestrating pipelines with Airflow, Dagster, Databricks Lakeflow Jobs, Snowflake Tasks, or equivalent.
- Experience handling pipeline failures, re-runs, and backfills.
- Experience loading data from an operational database into a warehouse or lakehouse such as Snowflake or Databricks.
- Production cloud experience on Azure or another major cloud.
- Experience with object storage, Git, CI/CD, and Docker or Kubernetes.
- Experience writing tests alongside code and maintaining data correctness.
- Clear written and spoken English communication skills.
Responsibilities
- Build incremental extraction pipelines from production PostgreSQL into the analytical warehouse.
- Write and maintain shared dbt models, dimensions, metric building blocks, and serving tables.
- Implement consistent metrics in the semantic layer.
- Add tests, freshness checks, alerts, and safe backfills to maintain data quality.
- Apply access-control and tenant-separation patterns to models and pipelines.
- Maintain lineage from metrics to source records, drone images, and video.
- Validate incoming WMS data against agreed contracts with the integration team.
- Document models and register them in the data catalog.
- Deliver through CI/CD, use AI-assisted development, and join on-call for owned pipelines.
- Help onboard additional products onto the shared data foundation.
Benefits
- Fully remote role.
