Principal Systems Engineer APAC

Nscale is a London-based, full-stack AI cloud and infrastructure company that provides GPU compute, managed AI services, orchestration software, data centers, and power infrastructure for AI training, fine-tuning, and inference.

London, United Kingdom
About Nscale

Nscale builds and operates vertically integrated AI infrastructure spanning software, GPU compute, networking, storage, purpose-built data centers, and power. Its active cloud platform offers self-service inference endpoints, fine-tuning, managed Kubernetes and Slurm, virtual machines, and GPU clusters.

View jobs by Nscale

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will manage bare-metal compute lifecycle states, host inventory, firmware, BIOS settings, operating systems, out-of-band management, and telemetry. You will validate host networking and components, run stress tests, and operate large-scale public-cloud data planes.

Requirements

  • Bachelor's or master's degree in Computer Science, Technology, or equivalent
  • At least 10 years of relevant network and systems experience
  • Coding experience
  • Experience delivering and operating large-scale production systems with 1,000 or more server instances
  • Deep understanding of operating systems, computer networks, and high-performance applications
  • Proficiency in at least three programming languages
  • Experience with the full software development lifecycle
  • Ability to travel up to 50% of the time
  • Linux systems expertise
  • Experience with server and GPU hardware architecture and system management
  • Experience with InfiniBand or RoCE networking
  • Experience designing, developing, and operating public-cloud service data planes
  • Understanding of databases, SQL, caching technologies, and system-level architecture

Responsibilities

  • Manage bare-metal compute lifecycle states
  • Inventory racks, hosts, and host components
  • Apply and validate host firmware
  • Run stress tests on host networking and components
  • Read and apply BIOS settings
  • Install and manage host operating systems remotely
  • Manage host out-of-band management subsystems
  • Read and parse host telemetry