Infrastructure Operations Operator
Nscale is a London-based, full-stack AI cloud and infrastructure company that provides GPU compute, managed AI services, orchestration software, data centers, and power infrastructure for AI training, fine-tuning, and inference.
Funding history
About Nscale
Nscale builds and operates vertically integrated AI infrastructure spanning software, GPU compute, networking, storage, purpose-built data centers, and power. Its active cloud platform offers self-service inference endpoints, fine-tuning, managed Kubernetes and Slurm, virtual machines, and GPU clusters.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will operate and monitor AI data-centre infrastructure, troubleshoot hardware, software, and networking issues, and resolve operational incidents. You will support customers, coordinate vendors and hardware work, mentor junior operators, maintain runbooks, improve workflows, join on-call rotations, and travel to other data centres when needed.
Requirements
- At least 2 years of data-centre operations, infrastructure operations, technical support, or similar experience
- Understanding of data-centre operations, server hardware, and networking
- Ability to diagnose and resolve complex technical problems
- Experience with service-level agreements and service-management processes
- Ability to mentor less experienced team members
- Relevant technical certifications or equivalent practical experience
- Ability to work shifts, participate in on-call rotations, and support extended hours
- Ability to travel to other data-centre locations
Responsibilities
- Support daily data-centre infrastructure operations
- Monitor operations and meet service-level agreements
- Troubleshoot hardware, software, networking, and infrastructure issues
- Lead incident response and escalate complex problems
- Perform root-cause analysis and prevent recurring issues
- Handle advanced customer support requests and technical escalations
- Mentor and train junior operators
- Coordinate hardware replacements, maintenance, and break-fix work with vendors
- Create and maintain operational documentation and runbooks
- Improve operational processes and infrastructure reliability
Benefits
- Equity
- Flexible work arrangements
