Search...

Infrastructure Engineer GPU Fleet HPC

Alpha Compute logo
Alpha Compute

Stealth

Distributed
View jobs by Alpha Compute

Skills

About the Role

You will architect and own the lifecycle of high-density GPU systems supporting production AI training workloads. You will manage fleet health from firmware validation and bare-metal provisioning through decommissioning, oversee liquid-cooled environments, build telemetry and automated remediation, migrate inventory to NetBox DCIM, and manage outage escalations. You will support enterprise technical reviews, capacity planning, and a 24/7 on-call rotation.

Requirements

  • Extensive experience managing large-scale HPC environments or production GPU fleets at a hyperscaler, neocloud, or top-tier research facility.
  • Deep hands-on experience with H200, B200, or B300 systems.
  • Expert knowledge of 400G or 800G InfiniBand, ConnectX-7 NDR, ConnectX-8 XDR, NVLink, and NVSwitch architectures.
  • Strong Linux internals knowledge.
  • Proficiency in infrastructure automation using Python or Go.
  • Deep experience deploying and scaling DCGM-based telemetry and SNMP-based environmental monitoring.
  • Direct experience with Direct-to-Chip systems, coolant chemistry management, or immersion cooling.
  • Familiarity with NVIDIA Mission Control for Blackwell-class cluster management.
  • Expertise in Intel TDX or NVIDIA RIM attestation flows.
  • Prior experience as an initial infrastructure hire building standards from the ground up.

Responsibilities

  • Own the end-to-end lifecycle and health of H200, B200, and B300 nodes from firmware validation and provisioning through decommissioning.
  • Lead operational oversight of high-density liquid-cooled environments, including CDU health, secondary-loop telemetry, and GPU thermals.
  • Architect telemetry and automated remediation using Prometheus, Grafana, and NVIDIA DCGM.
  • Migrate inventory to NetBox DCIM and build API integrations for asset tracking, IPAM, cabling, and compliance audits.
  • Serve as the primary technical interface for third-party facility operators and MSPs.
  • Set SLA and KPI compliance standards, lead technical post-mortems, and manage escalations for cluster-level outages.
  • Support enterprise deal cycles with capacity planning, infrastructure deep-dives, and technical reviews.
  • Participate in a 24/7 on-call rotation and take primary accountability for fleet availability and incident response.