Machine Learning and Cloud Infrastructure Engineer
SpAItialVisit SpAItial website
SpAItial is an AI company building physically grounded world models that generate persistent, explorable 3D worlds from text, images, and panoramas.
London, United Kingdom
Funding history
About SpAItial
SpAItial Ltd operates the Echo model family, a spatial-AI product available through a web app and developer API for generating, editing, sharing, and exporting 3D Gaussian Splat worlds.
Skills
About the Role
You will build and operate infrastructure for large-scale model training, evaluation, and serving. You will manage GPU clusters, storage, orchestration, observability, security, and production deployment pathways while partnering with researchers and engineers to improve tooling.
Requirements
- 3+ years of professional experience in infrastructure, platform, or cloud engineering
- Experience with GPU compute and performance debugging
- Experience operating AWS, GCP, or Azure environments
- Proficiency with Docker, Kubernetes, and Terraform
- Strong Python and Bash or PowerShell scripting skills
- Familiarity with PyTorch distributed training stacks
- Experience with observability tooling
- Experience building CI/CD for infrastructure and ML workflows
Responsibilities
- Own and evolve ML and cloud infrastructure for training and evaluating foundation models
- Design and operate multi-node, multi-GPU training environments
- Support distributed training stacks and improve performance, stability, and reproducibility
- Build and optimize storage and networking for large-scale datasets
- Package and deploy workloads using containers, orchestration, and infrastructure as code
- Implement monitoring, logging, alerting, SLOs, and incident-response practices
- Manage secrets, IAM, and secure network boundaries
- Support model evaluation and serving infrastructure
