Senior Technical Product Manager GPU Infrastructure

Nscale is a London-based, full-stack AI cloud and infrastructure company that provides GPU compute, managed AI services, orchestration software, data centers, and power infrastructure for AI training, fine-tuning, and inference.

Series CRecently funded0 current maintainers0 active leadsTeam intelligence

Maintainer signals as of 9/23/2026

London, United Kingdom
About Nscale

Nscale builds and operates vertically integrated AI infrastructure spanning software, GPU compute, networking, storage, purpose-built data centers, and power. Its active cloud platform offers self-service inference endpoints, fine-tuning, managed Kubernetes and Slurm, virtual machines, and GPU clusters.

View jobs by Nscale

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will own product strategy and roadmaps for GPU fleet operations software. You will turn operational problems into product capabilities, lead cross-functional delivery, define reliability metrics, guide technical trade-offs, convert incident learnings into commitments, and mentor product managers.

Requirements

  • 5–8 years of product management experience in software or technology.
  • Experience owning significant infrastructure, platform, or operations-facing product areas.
  • Technical fluency in provisioning, orchestration, observability, and control-plane design.
  • Experience building products for SREs, support teams, data center technicians, or similar operators.
  • Ability to deliver products that improve reliability, efficiency, or time to recover.
  • Experience with data center networking, InfiniBand, RoCE, WAN, edge, and global backbone architectures.
  • Experience partnering with network engineering teams on fabric health, congestion monitoring, and link-level failures.

Responsibilities

  • Own strategy and roadmaps for Fleet Operations product areas.
  • Lead cross-functional initiatives from problem framing through rollout on live GPU clusters.
  • Translate operational toil into tooling, automation, and platform capabilities.
  • Define and improve fleet availability, utilization, MTTR, bring-up time, hardware failure, and support metrics.
  • Partner with engineering on architecture and trade-offs across provisioning, orchestration, observability, and control planes.
  • Drive incident reviews and postmortems into product commitments.
  • Mentor junior product managers.
  • Represent Fleet Operations in planning, reviews, and leadership updates.

Benefits

  • Equity package.
  • Medical insurance.
  • Dental insurance.
  • Vision insurance.
  • Flexible paid time off.
  • Parental leave.
  • Retirement plan participation.