Senior Manager Deployment Engineering GPU Systems and Networking
Nscale is a London-based, full-stack AI cloud and infrastructure company that provides GPU compute, managed AI services, orchestration software, data centers, and power infrastructure for AI training, fine-tuning, and inference.
Funding history
About Nscale
Nscale builds and operates vertically integrated AI infrastructure spanning software, GPU compute, networking, storage, purpose-built data centers, and power. Its active cloud platform offers self-service inference endpoints, fine-tuning, managed Kubernetes and Slurm, virtual machines, and GPU clusters.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will lead and scale an engineering team responsible for core platform systems. You will define platform initiatives, improve delivery velocity and reliability, align interdependent infrastructure work, resolve ambiguous challenges, and raise engineering and operational standards.
Requirements
- Bachelor's or master's degree in Computer Science, Technology, or equivalent
- At least 10 years of network, systems, compute engineering management, or operations people-management experience
- Experience in a large ISP, cloud provider, or hyperscaler data center environment
- Ability to travel across APAC up to 50% of the time
- Experience with InfiniBand or RoCE networking
- Knowledge of MPLS, BGP, OSPF, IS-IS, TCP, IPv4, IPv6, DNS, and DHCP
- Experience with virtualization, containerization, and distributed systems design
- Linux or Unix systems administration experience
- Experience with infrastructure as code, configuration management, scripting, or automation
- Experience with GPU burn-in and validation testing at scale
- Experience with monitoring, observability, and telemetry solutions
Responsibilities
- Lead and scale a high-performing engineering team
- Define and drive multi-quarter platform initiatives
- Improve delivery velocity, system reliability, and quality
- Partner with product and engineering leaders on platform evolution
- Turn ambiguous technical challenges into execution plans
- Drive alignment across interdependent teams
- Raise engineering quality, operational excellence, and execution standards
