Senior HPC/GPU Systems Engineer
Nscale is a London-based, full-stack AI cloud and infrastructure company that provides GPU compute, managed AI services, orchestration software, data centers, and power infrastructure for AI training, fine-tuning, and inference.
Funding history
About Nscale
Nscale builds and operates vertically integrated AI infrastructure spanning software, GPU compute, networking, storage, purpose-built data centers, and power. Its active cloud platform offers self-service inference endpoints, fine-tuning, managed Kubernetes and Slurm, virtual machines, and GPU clusters.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will serve as a senior technical escalation point for GPU fleets, high-performance networking, Linux systems, and data-centre operations. You will diagnose hardware and fabric faults, investigate incidents, execute safe production changes, improve observability and runbooks, automate operational work, and mentor engineers.
Requirements
- 6+ years of infrastructure, operations, or support engineering experience in production environments
- 2–3+ years of hands-on GPU, HPC, or large-scale data-centre experience
- GPU driver, firmware, runtime, nvidia-smi, DCGM, and XID troubleshooting experience
- InfiniBand or RoCE RDMA fabric experience
- Slurm, Pyxis/Enroot, and MPI experience
- Linux systems engineering experience
- BMC, Redfish, firmware management, and bare-metal provisioning knowledge
- L2/L3 networking, routing, BGP, VLAN, VXLAN, firewall, and load-balancing knowledge
- Prometheus/Grafana observability and incident-response experience
- Bash or Python scripting skills
- Ansible, Terraform, or similar infrastructure automation experience
- Ability to participate in out-of-hours support and travel for onsite work
Responsibilities
- Join the support duty rotation as a senior escalation point
- Diagnose and remediate GPU node faults across driver, firmware, and hardware layers
- Maintain east-west fabric health and isolate network faults
- Investigate high-performance storage data-path issues
- Conduct root-cause analysis and drive long-term fixes
- Author and execute production changes with risk assessments and backout plans
- Improve dashboards, alerts, and runbooks
- Maintain accurate tickets and stakeholder communications
- Build automation scripts and small tools
- Mentor mid-level engineers
- Respond to critical incidents and participate in on-call
- Travel to Nscale or customer sites for onsite technical expertise
Benefits
- Equity
- Flexible work arrangements
- Remote-first work
- Medical insurance
- Dental insurance
- Vision insurance
- Flexible paid time off
- Parental leave
- Retirement plan participation
