Software Engineer Data Ingestion
Wayve is a London-headquartered embodied-AI company developing and licensing mapless, vehicle-agnostic driving software for assisted, automated, and robotaxi applications.
About Wayve
Wayve Technologies Ltd. develops the Wayve AI Driver, an end-to-end, data-trained software platform that runs on onboard vehicle compute and native sensors. It is designed for OEM integration across L1 driver assistance through L4 automated driving, without HD maps.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will improve the reliability, efficiency, and throughput of pipelines that process real-world driving data. You will diagnose production failures, make pipelines resilient to poor-quality inputs, optimize Spark workloads, support workflow orchestration, and deliver scalable improvements for downstream data users.
Requirements
- Production experience with Apache Spark
- Python engineering experience
- Experience building, debugging, or operating large-scale data-ingestion, ETL, or data-processing pipelines
- Experience with distributed data-processing systems
- Ability to optimize jobs for throughput, compute efficiency, and reliability
- Experience debugging production pipeline failures
- Experience working with corrupt, incomplete, or inconsistent data
- Understanding of multi-step orchestration and downstream dependencies
- Experience at PB-scale or similarly high-throughput data scale
- Experience with Airflow, Flyte, Databricks Workflows, Databricks, Delta Lake, Scala, Java, queue-based processing, or batch processing is advantageous
- Experience with automotive, robotics, autonomy, mapping, ML data platforms, embodied AI, or high-performance distributed systems is advantageous
Responsibilities
- Debug and resolve failing or blocked ingestion pipelines
- Investigate corrupt, malformed, and unexpected data
- Design resilient pipelines that prevent bad segments from blocking workflows
- Improve handling of varied partner and third-party data formats
- Support multi-step workflow orchestration, dependencies, retries, and queue management
- Optimize Spark jobs and data-processing pipelines
- Reduce operational toil and manual interventions
- Prioritize and unblock datasets for downstream users
- Partner with Data Platform and downstream engineering teams
- Improve the ingestion platform's technical direction, maintainability, and operational excellence
Benefits
- Competitive equity package
- Hybrid working policy
- Work-from-home time
