Deputy Manager Infrastructure Operations
Nscale is a London-based, full-stack AI cloud and infrastructure company that provides GPU compute, managed AI services, orchestration software, data centers, and power infrastructure for AI training, fine-tuning, and inference.
Maintainer signals as of 9/25/2026
Funding history
About Nscale
Nscale builds and operates vertically integrated AI infrastructure spanning software, GPU compute, networking, storage, purpose-built data centers, and power. Its active cloud platform offers self-service inference endpoints, fine-tuning, managed Kubernetes and Slurm, virtual machines, and GPU clusters.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will take day-to-day operational ownership of the data centre, coordinate engineers and operators, and maintain reliable site procedures. You will support deployments, vendors, incident response, and training while deputising for the Infrastructure Operations Manager when needed.
Requirements
- Operational experience in a data centre environment
- Working knowledge of HPC and GPU installations, hardware configuration, and network architecture
- Advanced troubleshooting and root cause analysis skills
- Ability to coordinate and influence a team without formal authority
- Ability to relocate to and live in Narvik
- Documentation and standardisation skills
- Willingness to participate in an on-call rotation
Responsibilities
- Own day-to-day site operations and ensure the white space runs effectively
- Act as the operational decision-maker during day shifts and escalate when appropriate
- Oversee installation, configuration, and maintenance of HPC and GPU systems
- Identify operational risks and propose efficiency improvements
- Coordinate engineers, operators, and partner staff
- Train and guide junior team members
- Manage rota and absence cover and deputise for the Infrastructure Operations Manager
- Maintain site documentation and SOPs
- Ensure procedures are followed and feed lessons into standards and training
- Support deployment handovers and coordinate on-site vendors
- Participate in an alternating weekly on-call rotation
- Respond to P1 and P2 call-outs and contribute to root cause analysis
