Network Operations Center Specialist
Nscale is a London-based, full-stack AI cloud and infrastructure company that provides GPU compute, managed AI services, orchestration software, data centers, and power infrastructure for AI training, fine-tuning, and inference.
Funding history
About Nscale
Nscale builds and operates vertically integrated AI infrastructure spanning software, GPU compute, networking, storage, purpose-built data centers, and power. Its active cloud platform offers self-service inference endpoints, fine-tuning, managed Kubernetes and Slurm, virtual machines, and GPU clusters.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will monitor network infrastructure, servers, environmental systems, and data centre alarms. You will troubleshoot first-line incidents, document and manage tickets, escalate complex issues, support maintenance and shift handovers, and improve monitoring processes and operational records.
Requirements
- 1–3 years of experience in a NOC, IT helpdesk, technical support, or data center environment.
- Degree or vocational certificate in Network Engineering, Information Technology, or a related field, or relevant hands-on experience or certifications.
- Understanding of TCP/IP, DNS, DHCP, switches, routers, and firewalls.
- Familiarity with monitoring tools and ticketing systems.
- Basic Linux/Unix command-line and Windows Server knowledge.
- Written and verbal English communication skills.
- Ability to work flexible shifts, including nights, weekends, and holidays.
Responsibilities
- Monitor network infrastructure, servers, and critical data centre systems.
- Respond to alerts and perform first-line troubleshooting.
- Escalate complex network, infrastructure, and facility issues.
- Monitor power, cooling, and BMS alarms.
- Log, prioritise, and manage incidents through the ticketing system.
- Perform routine operational checks and follow SOPs.
- Document incidents, maintenance activities, risks, and outstanding actions for shift handovers.
- Support planned maintenance and operational activities.
- Contribute evidence and timelines to root cause investigations.
- Improve monitoring, automation, and operational efficiency.
- Maintain operational records, reports, and system documentation.
