VP of Engineering
Hyperbolic is an open-access AI cloud providing on-demand GPU infrastructure, reserved and private cloud capacity, managed inference, and AI compute services.
Maintainer signals as of 9/2/2026
Funding history
About Hyperbolic
Hyperbolic operates an AI cloud platform for training, fine-tuning, and serving AI models at scale. It provides self-serve GPU instances and clusters, reserved capacity, private cloud infrastructure, OpenAI-compatible inference APIs, dedicated model hosting, and related infrastructure services for developers, researchers, startups, and enterprises.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
Lead the foundational infrastructure powering Hyperbolic's AI cloud platform, including GPU orchestration, compute scheduling, networking, storage, distributed systems, Kubernetes, observability, CI/CD, security, SRE, and platform engineering.
Requirements
- 12+ years building and operating large-scale infrastructure systems.
- Experience leading infrastructure organizations while remaining hands-on technically.
- Experience building or operating a cloud platform at scale.
- Experience building GPU infrastructure or AI/ML compute platforms.
- Track record scaling infrastructure in high-growth startup environments.
- Expert-level Kubernetes knowledge.
- Experience designing and operating multi-region cloud infrastructure.
- Strong understanding of Linux, networking, distributed systems, and storage architecture.
- Experience with Infrastructure-as-Code and automation frameworks.
- Deep expertise in observability, monitoring, and reliability engineering.
- Experience building highly available production systems.
- Experience with GPU scheduling, Slurm, Kubernetes GPU operators, Ray, or distributed training systems.
- Experience managing thousands of GPUs in production environments.
- Background supporting AI training and inference platforms.
Responsibilities
- Lead the design and evolution of a scalable AI cloud platform.
- Define architecture for GPU orchestration, compute scheduling, networking, storage, and distributed systems.
- Make critical decisions on cloud infrastructure, bare-metal deployments, and platform scalability.
- Build and scale large GPU clusters supporting customer workloads.
- Design GPU provisioning, scheduling, utilization optimization, and capacity management systems.
- Drive platform reliability and performance for AI training and inference workloads.
- Establish best practices for Kubernetes, observability, CI/CD, security, and operational excellence.
- Build SRE and Platform Engineering functions from the ground up.
- Define reliability standards including SLOs, SLIs, incident response, and capacity planning.
- Drive infrastructure automation.
- Recruit and develop Infrastructure, Platform, and SRE teams.
- Partner with executive leadership on company strategy, infrastructure investments, budgets, vendors, and capacity planning.
