AI Compute Hardware and Platform Lead
Skills
About the Role
You own the selection, qualification, and lifecycle management of GPU compute platforms. You lead server qualification, procurement strategy, platform validation, firmware management, repair and sparing models, acceptance testing, fleet strategy, service logistics, and lifecycle decisions across heterogeneous hardware generations.
Requirements
- Extensive experience building and operating large accelerator fleets
- Experience in hyperscale, cloud, HPC, or advanced systems environments
- Deep knowledge of server architecture, board-level integration, firmware risk, and fleet reliability engineering
- Track record with NVIDIA HGX or DGX, TPU infrastructure, custom accelerator programs, or dense rack-scale compute systems
- Strong understanding of manufacturing quality, service logistics, and infrastructure deployment at scale
- Ability to bridge lab-based engineering with global production operations
- Ability to mentor senior engineers on hardware systems thinking and fleet management
Responsibilities
- Own server qualification for AI compute platforms
- Lead procurement strategy and OEM relationships
- Lead BOM validation and platform bring-up
- Manage firmware lifecycles, repair strategies, sparing models, burn-in, acceptance testing, and end-of-life planning
- Establish fleet strategy across compute, interconnect, power, cooling, and management systems
- Build operating models for FRU inventories, RMA execution, depot workflows, and field replacement
- Create service strategies for board swaps, firmware rollback, and spare-part pooling
- Influence vendor roadmaps through technical engagement
- Build lifecycle decision frameworks for hardware expansion, sustainment, refresh, and retirement
