Forward Deployed Infrastructure Engineer
Hyperbolic is an open-access AI cloud that provides on-demand, reserved, and private GPU infrastructure for AI training, fine-tuning, inference, and production workloads. It serves startups, researchers, AI labs, enterprises, and compute providers.
Projects
About Hyperbolic
Hyperbolic operates an AI cloud platform that gives teams flexible access to high-performance GPUs through a global compute-provider network. Its offerings include self-serve on-demand GPU instances, reserved dedicated clusters, and private cloud infrastructure, supporting training, fine-tuning, inference, batch jobs, and long-running production workloads. The company serves AI-native teams, researchers, AI labs, enterprises, data centers, and infrastructure providers.
Skills
About the Role
You will serve as the technical point of contact during customer trials, running standardized and custom benchmarks to validate performance. You will design, run, and analyze performance tests across customer workloads, diagnose GPU and NCCL issues, optimize container and cluster configurations, and produce clear reports and handoffs. You will maintain benchmarking scripts, containers, and environments and iterate on configurations to close performance gaps.
Requirements
- Experience running infrastructure performance tests or ML model benchmarks (training or inference)
- Strong knowledge of GPU cloud infrastructure and workload bottlenecks
- Clear and fast written communication
- Ability to manage multiple trials and projects concurrently
- Familiarity with AWS Lambda CoreWeave Runpod and similar GPU cloud providers
- Prior customer-facing experience in startup or devtools settings (preferred)
- Background as an ML engineer solutions architect or technical account manager (preferred)
Responsibilities
- Serve as the technical point of contact during customer trials
- Design and run performance and benchmark tests across customer workloads
- Diagnose performance issues and recommend fixes
- Package results into clear reports and handoffs
- Maintain benchmarking scripts containers and environments
- Identify performance gaps and optimize cluster configurations
