Infrastructure Engineer: GPU Fleet (HPC)
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will architect and own the lifecycle of a high-density H200, B200, and B300 GPU fleet. You will manage firmware validation, bare-metal provisioning, decommissioning, liquid-cooled infrastructure, telemetry, automated remediation, NetBox integration, vendor operations, enterprise infrastructure reviews, and 24/7 incident response. You will also lead technical post-mortems, manage escalations, and support capacity planning for major clients.
Requirements
- Extensive experience managing large-scale HPC environments or production GPU fleets at a hyperscaler, neocloud, or top-tier research facility
- Deep hands-on experience with H200, B200, or B300 systems
- Expert knowledge of 400G/800G InfiniBand, ConnectX-7 NDR, ConnectX-8 XDR, NVLink, and NVSwitch architectures
- Strong Linux internals knowledge
- Proven proficiency building infrastructure automation using Python or Go
- Deep experience deploying and scaling DCGM-based telemetry and SNMP-based environmental monitoring
- Direct experience with Direct-to-Chip systems, coolant chemistry management, or immersion cooling
- Familiarity with NVIDIA Mission Control
- Expertise in Intel TDX or NVIDIA RIM attestation flows
- Prior experience as an initial infrastructure hire
Responsibilities
- Own the end-to-end health and lifecycle of H200, B200, and B300 GPU nodes
- Lead operational oversight of high-density liquid-cooled environments
- Monitor CDU health, secondary loop telemetry, and GPU thermals
- Architect telemetry using Prometheus, Grafana, and NVIDIA DCGM
- Trigger automated node draining, reboots, and health validation
- Migrate inventory to NetBox DCIM
- Build API integrations for asset tracking, IPAM, and cabling
- Serve as the primary technical interface for facility operators and MSPs
- Set SLA and KPI compliance standards
- Lead technical post-mortems and manage cluster-level outage escalations
- Support enterprise deal cycles with capacity planning and infrastructure reviews
- Participate in a 24/7 on-call rotation
- Own fleet availability and incident response
