Infrastructure Operations Manager
Nscale is a London-based, full-stack AI cloud and infrastructure company that provides GPU compute, managed AI services, orchestration software, data centers, and power infrastructure for AI training, fine-tuning, and inference.
Funding history
About Nscale
Nscale builds and operates vertically integrated AI infrastructure spanning software, GPU compute, networking, storage, purpose-built data centers, and power. Its active cloud platform offers self-service inference endpoints, fine-tuning, managed Kubernetes and Slurm, virtual machines, and GPU clusters.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will lead onsite data center operations, infrastructure, and personnel. You will ensure reliable operation of critical systems, manage HPC and GPU installations, coordinate shifts and on-call support, work with clients and vendors, maintain inventory, report performance, and implement operational improvements.
Requirements
- Bachelor's degree in Computer Science, Engineering, or a related field
- 5+ years of experience managing data centers, particularly in HPC and GPU environments
- Leadership experience managing and developing teams
- Expertise in HPC, GPU deployment, and related hardware and software
- Knowledge of data center power, cooling, and environmental systems
- Client-facing communication and reporting skills
- Experience managing inventory and spare parts
- Ability to work onsite at the data center
- Ability to participate in an on-call rotation
Responsibilities
- Manage all data center devices, systems, and infrastructure
- Ensure continuous operational reliability for AI workloads
- Monitor power, cooling, and environmental conditions
- Oversee installation, configuration, and maintenance of HPC and GPU systems
- Lead and mentor engineers, technicians, and support staff
- Manage shift schedules and on-call coverage
- Serve as the primary client contact and report on SLAs and KPIs
- Manage vendors, contractors, procurement, repairs, and upgrades
- Manage spare-parts inventory
- Troubleshoot and resolve technical issues
- Monitor data center performance and resource utilization
- Prepare operational reports
- Implement reliability, scalability, and cost-effectiveness improvements
