Senior AI Infrastructure Engineer
Hyperbolic is an open-access AI cloud providing on-demand GPU infrastructure, reserved and private cloud capacity, managed inference, and AI compute services.
Maintainer signals as of 9/2/2026
Funding history
About Hyperbolic
Hyperbolic operates an AI cloud platform for training, fine-tuning, and serving AI models at scale. It provides self-serve GPU instances and clusters, reserved capacity, private cloud infrastructure, OpenAI-compatible inference APIs, dedicated model hosting, and related infrastructure services for developers, researchers, startups, and enterprises.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will design, build, and operate the infrastructure that transforms raw GPUs into a programmable, orchestrated pool for AI workloads. You will implement bare-metal provisioning and lifecycle management, develop GPU scheduling and placement strategies, automate provisioning with infrastructure as code, integrate storage solutions for training data, design APIs and cloud-init workflows for automated configuration, optimize GPU compute with CUDA, and work directly with hardware vendors to troubleshoot and improve integrations.
Requirements
- Bare-metal provisioning and lifecycle management, including IPMI, Redfish, BMC, PXE, and automated OS deployment
- GPU scheduling and orchestration with GPU type awareness, memory, topology, placement, and fragmentation minimization
- Terraform or Pulumi and CI/CD for infrastructure
- Secrets management and configuration management
- Observability stack implementation
- Storage and data infrastructure for AI/ML, including object storage, high-IOPS block storage, and distributed file systems
- API design and cloud-init for automated provisioning
- GPU architecture, CUDA, and GPU compute optimization
- Experience building and scaling cloud infrastructure or distributed systems in production
- Ability to work with hardware vendors and vendor engineering teams
- Strong communication skills
- Preferred: InfiniBand, RoCE, and distributed storage systems such as Ceph, Weka, or VAST Data
Responsibilities
- Build and scale a multi-tenant GPU cloud marketplace
- Design and implement multi-tenancy provisioning and virtualization solutions
- Transform raw GPUs into a programmable, orchestrated resource pool
- Implement bare-metal provisioning and lifecycle management
- Develop GPU scheduling, placement strategies, and fragmentation minimization
- Automate infrastructure using Terraform or Pulumi and CI/CD pipelines
- Implement secrets management, configuration management, and observability
- Design APIs and cloud-init workflows for automated provisioning
- Integrate and operate storage solutions for AI/ML workloads
- Collaborate with hardware vendors to troubleshoot and optimize integrations
