Principal Network Engineer
Nscale is a London-based, full-stack AI cloud and infrastructure company that provides GPU compute, managed AI services, orchestration software, data centers, and power infrastructure for AI training, fine-tuning, and inference.
Maintainer signals as of 9/23/2026
Funding history
About Nscale
Nscale builds and operates vertically integrated AI infrastructure spanning software, GPU compute, networking, storage, purpose-built data centers, and power. Its active cloud platform offers self-service inference endpoints, fine-tuning, managed Kubernetes and Slurm, virtual machines, and GPU clusters.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will set the technical direction for AI interconnect networks and evolve large-scale InfiniBand and RoCE fabrics. You will lead complex incident investigations, define network standards, improve reliability and performance, influence system design, and mentor network engineers.
Requirements
- At least 10 years of network-engineering experience focused on HPC, AI, or hyperscale data-centre networking
- Expert operational and architectural experience with InfiniBand or large-scale RoCE fabrics
- Understanding of RDMA internals, congestion management, and fabric failure modes
- Expertise with BGP, OSPF, and ECMP
- Ability to debug cross-layer hardware, firmware, kernel, and application communication-library issues
- Ability to lead complex technical initiatives across teams
- Systems-level understanding of performance, reliability, scalability, and operational cost
Responsibilities
- Own the technical direction and operational strategy for AI interconnect networks
- Design, review, and evolve InfiniBand and RoCE fabric architectures
- Lead complex network incident investigations and systemic fixes
- Improve fabric reliability, performance predictability, and operational maturity
- Define hardware, congestion-control, routing, firmware, and change-safety standards
- Influence end-to-end system design with SRE, compute, and network architecture stakeholders
- Mentor senior and mid-level network engineers
- Improve uptime, latency consistency, capacity efficiency, and incident reduction
Benefits
- Equity
