Research Scientist Robot Learning VLA and WAM
SpAItialVisit SpAItial website
SpAItial is an AI company building physically grounded world models that generate persistent, explorable 3D worlds from text, images, and panoramas.
London, United Kingdom
Funding history
About SpAItial
SpAItial Ltd operates the Echo model family, a spatial-AI product available through a web app and developer API for generating, editing, sharing, and exporting 3D Gaussian Splat worlds.
Skills
About the Role
You will own vision-language-action and world-action model training from data through robot deployment. You will design action representations, adapt vision-language-model backbones for control, curate robot datasets, build action-conditioned world-model components, and use simulation, fine-tuning, and reinforcement learning to improve policy robustness.
Requirements
- PhD in robotics, machine learning, or computer vision with a robot-learning focus
- Publications, open-source work, or deployed systems
- Deep experience training modern robot-policy designs including VLA, WAM, and diffusion end to end
- Strong imitation-learning fundamentals and familiarity with RL fine-tuning
- Fluency with VLM backbones and control adaptation
- Expert Python and PyTorch skills
- Multi-node distributed-training experience with FSDP or equivalent
Responsibilities
- Own end-to-end VLA and WAM training from data to a robot-running policy
- Contribute to the technical direction for embodied research
- Close the sim-to-real gap through domain randomization, system identification, calibration, and transfer evaluation
- Adapt VLM backbones for control through encoder selection, adapters, and co-training
- Curate and weight heterogeneous robot training datasets
- Design action representation and decoding using tokenization, chunking, diffusion, and flow matching
- Build action-conditioned world-model components that predict future observations
- Run supervised fine-tuning and reinforcement-learning post-training for robustness
