Infrastructure Operations Manager
Nscale is a London-based, full-stack AI cloud and infrastructure company that provides GPU compute, managed AI services, orchestration software, data centers, and power infrastructure for AI training, fine-tuning, and inference.
Funding history
About Nscale
Nscale builds and operates vertically integrated AI infrastructure spanning software, GPU compute, networking, storage, purpose-built data centers, and power. Its active cloud platform offers self-service inference endpoints, fine-tuning, managed Kubernetes and Slurm, virtual machines, and GPU clusters.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will take responsibility for data center devices, systems, and infrastructure. You will lead site operations and staff, monitor critical conditions, manage HPC and GPU systems, work with clients and vendors, maintain inventory, report performance, and improve reliability and efficiency.
Requirements
- Bachelor's degree in Computer Science, Engineering, or a related field
- 5+ years of experience managing data centers, particularly in HPC and GPU environments
- Leadership experience managing and developing teams
- Expertise in HPC, GPU deployment, and related hardware and software
- Knowledge of data center power, cooling, and environmental systems
- Client-facing communication and reporting skills
- Experience managing inventory and spare parts
Responsibilities
- Manage all data center devices, systems, and infrastructure
- Ensure continuous operational reliability for AI workloads
- Monitor power, cooling, and environmental conditions
- Oversee installation, configuration, and maintenance of HPC and GPU systems
- Lead and mentor engineers, technicians, and support staff
- Manage shift schedules and on-call coverage
- Serve as the primary client contact and report on SLAs and KPIs
- Manage vendors, contractors, procurement, repairs, and upgrades
- Manage spare-parts inventory
- Troubleshoot technical issues and escalate complex problems
- Monitor data center performance and resource utilization
- Prepare performance and operational reports
- Implement reliability, scalability, and cost-effectiveness improvements
Benefits
- Equity
- Flexible workplace
