Machine Learning and Cloud Infrastructure Engineer

SpAItial is an AI company building physically grounded world models that generate persistent, explorable 3D worlds from text, images, and panoramas.

London, United Kingdom

Funding history

About SpAItial

SpAItial Ltd operates the Echo model family, a spatial-AI product available through a web app and developer API for generating, editing, sharing, and exporting 3D Gaussian Splat worlds.

View jobs by SpAItial

Skills

About the Role

You will build and operate infrastructure for large-scale model training, evaluation, and serving. You will manage GPU clusters, storage, orchestration, observability, security, and production deployment pathways while partnering with researchers and engineers to improve tooling.

Requirements

  • 3+ years of professional experience in infrastructure, platform, or cloud engineering
  • Experience with GPU compute and performance debugging
  • Experience operating AWS, GCP, or Azure environments
  • Proficiency with Docker, Kubernetes, and Terraform
  • Strong Python and Bash or PowerShell scripting skills
  • Familiarity with PyTorch distributed training stacks
  • Experience with observability tooling
  • Experience building CI/CD for infrastructure and ML workflows

Responsibilities

  • Own and evolve ML and cloud infrastructure for training and evaluating foundation models
  • Design and operate multi-node, multi-GPU training environments
  • Support distributed training stacks and improve performance, stability, and reproducibility
  • Build and optimize storage and networking for large-scale datasets
  • Package and deploy workloads using containers, orchestration, and infrastructure as code
  • Implement monitoring, logging, alerting, SLOs, and incident-response practices
  • Manage secrets, IAM, and secure network boundaries
  • Support model evaluation and serving infrastructure