Technical Support Engineer

Hyperbolic is an open-access AI cloud providing on-demand GPU infrastructure, reserved and private cloud capacity, managed inference, and AI compute services.

Series A0 current maintainers0 active leadsTeam intelligence

Maintainer signals as of 8/29/2026

San Francisco, California, United States
About Hyperbolic

Hyperbolic operates an AI cloud platform for training, fine-tuning, and serving AI models at scale. It provides self-serve GPU instances and clusters, reserved capacity, private cloud infrastructure, OpenAI-compatible inference APIs, dedicated model hosting, and related infrastructure services for developers, researchers, startups, and enterprises.

View jobs by Hyperbolic

Skills

About the Role

You will be the first technical contact for customers experiencing GPU cluster issues. Own tickets from initial response through resolution, manage SLA commitments, reproduce problems, collect relevant logs and configuration, and determine whether faults originate with the platform or a provider. Work daily in Linux and the command line while troubleshooting SSH access, NFS mounts, storage, quotas, security groups, containers, drivers, billing, and account issues. Execute and improve runbooks, maintain customer documentation, communicate clearly under pressure, escalate with complete context, and participate in an on-call rotation across time zones.

Requirements

  • Very strong Linux and command-line experience
  • Experience owning infrastructure or cloud support tickets against response SLAs
  • Solid networking and storage fundamentals including SSH, NFS, mounts, DNS, firewalls, and security groups
  • Working familiarity with GPU workloads including nvidia-smi, drivers, CUDA, and containers
  • Clear and fast written communication under time pressure
  • Good judgment about knowledge limits and when to escalate
  • Comfort working across time zones and participating in on-call rotations
  • Experience with ticketing and on-call tools such as Zendesk, Linear, or PagerDuty is preferred
  • Bash or Python scripting experience is preferred
  • Exposure to Slurm, Kubernetes, or Docker in multi-tenant environments is preferred
  • Background in GPU cloud, HPC, or hardware-adjacent support is preferred

Responsibilities

  • Own customer tickets from first response through resolution
  • Track SLA commitments and classify issue severity
  • Reproduce problems and collect relevant logs and configuration
  • Troubleshoot SSH access, NFS mounts, storage, quotas, security groups, containers, drivers, billing, and account issues
  • Determine whether faults originate with the platform or a provider
  • Execute and author runbooks
  • Maintain customer-facing documentation and internal knowledge bases
  • Escalate issues with complete technical context
  • Participate in an on-call rotation for critical issues