Head of Infrastructure Support

Nscale is a London-based, full-stack AI cloud and infrastructure company that provides GPU compute, managed AI services, orchestration software, data centers, and power infrastructure for AI training, fine-tuning, and inference.

London, United Kingdom
About Nscale

Nscale builds and operates vertically integrated AI infrastructure spanning software, GPU compute, networking, storage, purpose-built data centers, and power. Its active cloud platform offers self-service inference endpoints, fine-tuning, managed Kubernetes and Slurm, virtual machines, and GPU clusters.

View jobs by Nscale

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will own regional infrastructure-support outcomes, including service performance, customer impact, capacity planning, and team performance. You will hire and manage engineers, establish operational processes and global handovers, lead escalations and post-incident reviews, and remain technically credible across GPU infrastructure, Linux, high-performance fabrics, automation, and data-centre operations.

Requirements

  • 8+ years of infrastructure, operations, or support engineering experience in production environments
  • 4+ years of direct line management of engineers in an operational support function
  • Experience managing performance and service delivery against SLAs
  • Exposure to GPU, HPC, or large-scale data-centre estates
  • Linux systems-engineering experience across compute, storage, and network layers
  • Knowledge of GPU platforms, driver and firmware stacks, diagnostics, fault isolation, and RMA workflows
  • Knowledge of InfiniBand or RoCE, networking fundamentals, and data-centre operations
  • Hands-on Slurm experience for multi-GPU workloads
  • Scripting skills in Bash, Python, or similar
  • Familiarity with Ansible, Terraform, or similar Infrastructure as Code tools
  • Knowledge of ITIL-aligned processes and SRE practices
  • Availability for out-of-hours escalations, regional on-call, and travel

Responsibilities

  • Own regional service outcomes, customer impact, and team performance
  • Manage regional service KPIs, including SLA adherence, MTTR, first-response time, backlog health, and CSAT
  • Identify and resolve or escalate regional risks
  • Own capacity modelling and headcount planning
  • Establish consistent global operating standards and follow-the-sun handovers
  • Manage engineers through hiring, one-to-ones, reviews, development plans, and performance management
  • Design team structure, develop team leads, and manage shift, rota, and on-call coverage
  • Maintain ticket queue health and escalation flow
  • Ensure ITIL-aligned incident, request, change, and problem-management practices
  • Lead high-impact incident escalations and post-incident reviews
  • Contribute to service readiness for deployments and customer onboarding
  • Guide technical investigations, automation, and operational improvements
  • Travel to sites when needed to lead onsite support activity

Benefits

  • Medical insurance
  • Dental insurance
  • Vision insurance
  • Flexible paid time off
  • Parental leave
  • Retirement plan participation
Head of Infrastructure Support at Nscale | JobStash