Senior Systems Engineer

Nscale is a London-based, full-stack AI cloud and infrastructure company that provides GPU compute, managed AI services, orchestration software, data centers, and power infrastructure for AI training, fine-tuning, and inference.

London, United Kingdom
About Nscale

Nscale builds and operates vertically integrated AI infrastructure spanning software, GPU compute, networking, storage, purpose-built data centers, and power. Its active cloud platform offers self-service inference endpoints, fine-tuning, managed Kubernetes and Slurm, virtual machines, and GPU clusters.

View jobs by Nscale

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will lead the design and development of distributed systems supporting high-performance infrastructure. You will own critical services, make architecture decisions, resolve system-level issues, improve reliability through automation, mentor engineers, and support production systems through on-call rotation.

Requirements

  • 8–12+ years of experience building and operating large-scale production systems
  • Understanding of operating systems, distributed systems, and computer networking
  • Proficiency in Python, Java, C/C++, Go, or multiple programming languages
  • Experience owning systems and services across the development lifecycle
  • Experience with Linux systems and performance optimization
  • Knowledge of distributed systems architecture, data synchronization, fault tolerance, and state management
  • Experience with compute, storage, and networking environments
  • Knowledge of server and GPU hardware architecture
  • Experience with InfiniBand, RoCE, cloud services, data plane services, MySQL, Redis, or Memcached
  • Ability to participate in an on-call rotation

Responsibilities

  • Lead the design and development of large-scale distributed systems
  • Own critical compute and network infrastructure services
  • Drive architecture decisions for scalability, fault tolerance, state management, and data consistency
  • Partner cross-functionally to deliver solutions through production and operation
  • Mentor and guide engineers
  • Diagnose and resolve complex system-level issues
  • Improve performance, reliability, and operational efficiency through automation and tooling
  • Influence long-term technical strategy and roadmap
  • Participate in an on-call rotation

Benefits

  • Medical coverage
  • Dental coverage
  • Vision coverage
  • Flexible paid time off
  • Parental leave
  • Retirement plan participation