Staff HPC Systems Software Engineer

Nscale is a London-based, full-stack AI cloud and infrastructure company that provides GPU compute, managed AI services, orchestration software, data centers, and power infrastructure for AI training, fine-tuning, and inference.

London, United Kingdom
About Nscale

Nscale builds and operates vertically integrated AI infrastructure spanning software, GPU compute, networking, storage, purpose-built data centers, and power. Its active cloud platform offers self-service inference endpoints, fine-tuning, managed Kubernetes and Slurm, virtual machines, and GPU clusters.

View jobs by Nscale

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will define and evolve technical direction for an HPC systems domain, including Slurm architecture, scheduler integrations, cluster lifecycle, workload environments, and service automation. You will establish shared automation, lifecycle-management, observability, reliability, and deployment patterns. You will lead critical initiatives across multiple teams, resolve complex design issues, create reusable implementations, and contribute hands-on to de-risk important work.

Requirements

  • Experience designing and building production software and automation for HPC systems
  • Experience with Slurm-based environments
  • Go or Python
  • Slurm internals
  • GPU-backed infrastructure
  • HPC networking
  • InfiniBand
  • RoCE
  • RDMA
  • Cloud-native platform integration
  • Technical direction across multiple teams or services
  • Written communication
  • Verbal communication

Responsibilities

  • Own and evolve technical direction for an HPC systems domain
  • Make architectural decisions balancing quality, operations, customer needs, and maintainability
  • Define how Slurm implementations are packaged, automated, and delivered as services
  • Resolve cross-team ambiguity around ownership, interfaces, lifecycle boundaries, and operating models
  • Establish shared automation, observability, reliability, and supportability patterns
  • Drive integrations between Slurm, Kubernetes-adjacent systems, infrastructure APIs, identity systems, and platform tooling
  • Create reusable modules, automation, deployment patterns, and reference implementations
  • Lead critical initiatives across multiple teams
  • Contribute hands-on to de-risk and accelerate critical work
Staff HPC Systems Software Engineer at Nscale | JobStash