Staff ML Performance Engineer Inference Optimisation
Wayve is a London-headquartered embodied-AI company developing and licensing mapless, vehicle-agnostic driving software for assisted, automated, and robotaxi applications.
About Wayve
Wayve Technologies Ltd. develops the Wayve AI Driver, an end-to-end, data-trained software platform that runs on onboard vehicle compute and native sensors. It is designed for OEM integration across L1 driver assistance through L4 automated driving, without HD maps.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will profile and optimise the inference stack across model graphs, compilers, runtimes, kernels, and memory movement. You will build benchmarking and regression testing, support multiple hardware targets, collaborate on deployment decisions, and help set technical direction for performance engineering.
Requirements
- Production performance optimisation experience under constrained latency, memory, bandwidth, power, thermal, or cost conditions
- TensorRT, CUDA, Qualcomm QNN, Triton, or OpenCL
- Model, kernel, runtime, and compiler understanding
- Debugging
- Profiling
- Testing
- Maintainable code
- Stakeholder communication and collaboration
Responsibilities
- Profile inference-stack bottlenecks and deliver performance improvements
- Implement and validate compiler, runtime, and kernel optimisations
- Build benchmarking and regression testing across models, devices, and releases
- Optimise inference for multiple hardware targets
- Collaborate on model architecture and training or deployment decisions
- Contribute to technical roadmaps and performance-engineering tooling
Benefits
- Hybrid working policy
