Infrastructure Engineer GPU Fleet HPC
Skills
About the Role
You will architect and own the lifecycle of high-density GPU systems supporting production AI training workloads. You will manage fleet health from firmware validation and bare-metal provisioning through decommissioning, oversee liquid-cooled environments, build telemetry and automated remediation, migrate inventory to NetBox DCIM, and manage outage escalations. You will support enterprise technical reviews, capacity planning, and a 24/7 on-call rotation.
Requirements
- Extensive experience managing large-scale HPC environments or production GPU fleets at a hyperscaler, neocloud, or top-tier research facility.
- Deep hands-on experience with H200, B200, or B300 systems.
- Expert knowledge of 400G or 800G InfiniBand, ConnectX-7 NDR, ConnectX-8 XDR, NVLink, and NVSwitch architectures.
- Strong Linux internals knowledge.
- Proficiency in infrastructure automation using Python or Go.
- Deep experience deploying and scaling DCGM-based telemetry and SNMP-based environmental monitoring.
- Direct experience with Direct-to-Chip systems, coolant chemistry management, or immersion cooling.
- Familiarity with NVIDIA Mission Control for Blackwell-class cluster management.
- Expertise in Intel TDX or NVIDIA RIM attestation flows.
- Prior experience as an initial infrastructure hire building standards from the ground up.
Responsibilities
- Own the end-to-end lifecycle and health of H200, B200, and B300 nodes from firmware validation and provisioning through decommissioning.
- Lead operational oversight of high-density liquid-cooled environments, including CDU health, secondary-loop telemetry, and GPU thermals.
- Architect telemetry and automated remediation using Prometheus, Grafana, and NVIDIA DCGM.
- Migrate inventory to NetBox DCIM and build API integrations for asset tracking, IPAM, cabling, and compliance audits.
- Serve as the primary technical interface for third-party facility operators and MSPs.
- Set SLA and KPI compliance standards, lead technical post-mortems, and manage escalations for cluster-level outages.
- Support enterprise deal cycles with capacity planning, infrastructure deep-dives, and technical reviews.
- Participate in a 24/7 on-call rotation and take primary accountability for fleet availability and incident response.
