Head of Infrastructure Operations

Nscale is a London-based, full-stack AI cloud and infrastructure company that provides GPU compute, managed AI services, orchestration software, data centers, and power infrastructure for AI training, fine-tuning, and inference.

London, United Kingdom
About Nscale

Nscale builds and operates vertically integrated AI infrastructure spanning software, GPU compute, networking, storage, purpose-built data centers, and power. Its active cloud platform offers self-service inference endpoints, fine-tuning, managed Kubernetes and Slurm, virtual machines, and GPU clusters.

View jobs by Nscale

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will lead regional data-centre operations, establish operational strategy and standards, manage multi-site teams, oversee physical infrastructure and vendors, lead incident response, ensure compliance, and report on operational health. You will travel regularly to data centres and coordinate capacity, deployment, and service-delivery work.

Requirements

  • 10+ years of data-centre operations, infrastructure management, or facilities-management experience
  • Regional or multi-site operations leadership experience
  • Team-management experience across multiple locations
  • Experience scaling operations and maintaining reliability standards
  • Data-centre infrastructure knowledge including power, cooling, networking, and security
  • ISO 22237 and ISO 27001 Annex A.11 familiarity
  • Monitoring, environmental controls, and infrastructure automation knowledge
  • GPU/HPC infrastructure knowledge
  • SOC 2, ISO 27001, Cyber Essentials Plus, and ISO 22301 familiarity
  • Stakeholder management and senior-leadership communication skills
  • Regular travel to data centres

Responsibilities

  • Own regional data-centre operations strategy and execution
  • Establish operational standards, processes, and roadmaps
  • Drive cost, downtime, and operational-maturity improvements
  • Build, mentor, and lead multi-site operations teams
  • Oversee data-centre operational procedures and ITSM ticket handling
  • Support facility performance across power, cooling, security, and environmental controls
  • Maintain AI infrastructure asset inventory and physical-security controls
  • Establish SLOs and SLIs for availability, performance, and incident response
  • Lead incident response, root-cause analysis, remediation, and prevention
  • Ensure safety, environmental, and compliance requirements
  • Manage vendors, contractors, procurement, SLAs, and contract compliance
  • Coordinate infrastructure deployment, capacity planning, and site commissioning
  • Implement monitoring and alerting and report operational metrics
  • Travel regularly to data centres

Benefits

  • Equity
  • Medical insurance
  • Dental insurance
  • Vision insurance
  • Flexible paid time off
  • Parental leave
  • Retirement plan participation