Machine Learning Systems and Infrastructure Engineer
SpAItialVisit SpAItial website
SpAItial is an AI company building physically grounded world models that generate persistent, explorable 3D worlds from text, images, and panoramas.
London, United Kingdom
Funding history
About SpAItial
SpAItial Ltd operates the Echo model family, a spatial-AI product available through a web app and developer API for generating, editing, sharing, and exporting 3D Gaussian Splat worlds.
Skills
About the Role
You will build and operate ML systems for model training, evaluation, data ingestion, experiment orchestration, and serving. You will improve distributed training, develop large-scale data pipelines, operate workflow and inference systems, and maintain deployment, observability, security, and CI/CD capabilities.
Requirements
- 3+ years writing production-quality Python in a large multi-author codebase
- Experience with PyTorch distributed training stacks and debugging multi-GPU jobs
- Experience shipping end-to-end data pipelines at scale
- Experience with GPU compute and performance debugging
- Working knowledge of AWS, GCP, or Azure
- Proficiency with Docker, Kubernetes, and Terraform
- Knowledge of SQL, relational, analytical, embedded, and object storage systems
- Familiarity with workflow orchestration, experiment tracking, observability, and CI/CD tools
Responsibilities
- Own ML systems for training, evaluation, and serving foundation models
- Improve distributed training performance, stability, reproducibility, and checkpointing
- Build Python data ingestion, preprocessing, validation, versioning, and publishing pipelines
- Operate workflow orchestration, GPU scheduling, experiment tracking, and inference systems
- Ship workloads using Docker, Kubernetes, Terraform, and CI/CD pipelines
- Implement monitoring, logging, alerting, SLOs, and incident response
- Manage secrets, IAM, and network boundaries
- Partner with researchers, engineers, and the platform team to improve ML systems
