Infrastructure Engineer: GPU Fleet (HPC)

AI infrastructure company providing GPU-as-a-Service and confidential computing, formerly AlphaTON Capital Corp.

Road Town, British Virgin Islands

Projects

About Alpha Compute Corp.

Alpha Compute Corp. acquires, deploys, and leases high-performance NVIDIA GPUs in data centers to provide AI GPU-as-a-Service and hardware-enforced confidential compute for enterprises, technology companies, AI research laboratories, and decentralized AI ecosystems. It operates infrastructure in North America and Europe and retains residual TON exposure from its former digital-asset treasury strategy.

View jobs by Alpha Compute Corp.

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

Architect and own the lifecycle of a high-density GPU fleet supporting production AI training workloads, including firmware validation, bare-metal provisioning, liquid-cooled infrastructure, telemetry, automated remediation, NetBox integration, vendor operations, enterprise infrastructure reviews, and 24/7 incident response.

Requirements

  • Extensive experience managing large-scale HPC environments or production GPU fleets at a hyperscaler, neocloud, or top-tier research facility.
  • Deep hands-on experience with H200, B200, or B300 systems.
  • Expert knowledge of 400G/800G InfiniBand, ConnectX-7 NDR, ConnectX-8 XDR, NVLink, and NVSwitch architectures.
  • Strong Linux internals knowledge.
  • Proven proficiency building infrastructure automation using Python or Go.
  • Deep experience deploying and scaling DCGM-based telemetry and SNMP-based environmental monitoring.
  • Direct experience with Direct-to-Chip systems, coolant chemistry management, or immersion cooling.
  • Familiarity with NVIDIA Mission Control.
  • Expertise in Intel TDX or NVIDIA RIM attestation flows.
  • Prior experience as an initial infrastructure hire.

Responsibilities

  • Own the end-to-end health and lifecycle of H200, B200, and B300 GPU nodes.
  • Lead operational oversight of high-density liquid-cooled environments.
  • Monitor CDU health, secondary loop telemetry, and GPU thermals.
  • Architect telemetry using Prometheus, Grafana, and NVIDIA DCGM.
  • Trigger automated node draining, reboots, and health validation.
  • Migrate inventory to NetBox DCIM.
  • Build API integrations for asset tracking, IPAM, and cabling.
  • Serve as the primary technical interface for facility operators and MSPs.
  • Set SLA and KPI compliance standards.
  • Lead technical post-mortems and manage cluster-level outage escalations.
  • Support enterprise deal cycles with capacity planning and infrastructure reviews.
  • Participate in a 24/7 on-call rotation.
  • Own fleet availability and incident response.