Infrastructure Operations Deputy Manager

Nscale is a London-based, full-stack AI cloud and infrastructure company that provides GPU compute, managed AI services, orchestration software, data centers, and power infrastructure for AI training, fine-tuning, and inference.

London, United Kingdom
About Nscale

Nscale builds and operates vertically integrated AI infrastructure spanning software, GPU compute, networking, storage, purpose-built data centers, and power. Its active cloud platform offers self-service inference endpoints, fine-tuning, managed Kubernetes and Slurm, virtual machines, and GPU clusters.

View jobs by Nscale

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will take responsibility for data-center devices and infrastructure, ensuring reliable 24/7 operations for AI workloads. You will lead site personnel, manage vendors and inventory, monitor performance, resolve operational issues, report on KPIs, and improve reliability and scalability.

Requirements

  • 5+ years of experience managing data centers in HPC and GPU environments
  • Data-center management
  • HPC
  • GPU deployment
  • Power, cooling, and environmental systems
  • Client-facing communication
  • Reporting
  • Inventory management
  • Spare-parts management
  • On-call participation
  • Occasional travel

Responsibilities

  • Ensure operational reliability and performance for AI workloads
  • Monitor power, cooling, and environmental conditions
  • Oversee installation, configuration, and maintenance of HPC and GPU systems
  • Lead and mentor engineers, technicians, and support staff
  • Manage shift schedules and on-call coverage
  • Serve as the primary client contact for SLA and KPI reporting
  • Manage vendors, contractors, procurement, repairs, and upgrades
  • Manage spare-parts inventory
  • Troubleshoot and resolve technical issues
  • Implement monitoring tools and prepare operational reports
  • Propose and implement reliability, scalability, and cost-effectiveness improvements