Director of Infrastructure Engineering

Runpod is an AI developer cloud providing GPU compute, serverless inference, Pods, and clusters for building, training, fine-tuning, deploying, and scaling AI workloads.

Recently fundedCompany intelligence
San Francisco, United States
About Runpod

Runpod Inc. operates a globally distributed GPU cloud platform for AI developers, offering on-demand GPU infrastructure and serverless services across the AI development lifecycle.

View jobs by Runpod

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will lead teams responsible for SRE, networking, and storage. You will establish reliability practices, oversee HPC and global networking, guide distributed-storage performance, develop engineering leaders, forecast capacity, and improve delivery, infrastructure reliability, and availability.

Requirements

  • 7+ years leading software, infrastructure, SRE, or networking teams
  • 8+ years building and operating distributed systems, bare-metal infrastructure, or cloud platforms
  • Experience with InfiniBand or RoCE, spine-leaf architectures, and BGP
  • Experience with distributed storage or parallel file systems
  • Reliability engineering and infrastructure-as-code expertise
  • Experience with Terraform and Ansible
  • Experience with Kubernetes and observability stacks
  • Experience leading distributed technical teams
  • Successful completion of a background check

Responsibilities

  • Lead engineering teams responsible for SRE, networking, and storage
  • Establish SLA and SLO definitions, incident response, observability, and automated remediation
  • Oversee global and HPC network design, scaling, and operation
  • Direct distributed storage architecture and performance tuning
  • Hire, mentor, and grow engineering managers and senior individual contributors
  • Forecast capacity requirements and shape technical roadmaps
  • Improve reliability and delivery metrics
  • Provide architectural oversight for provisioning, virtualization, network fabrics, and storage clusters
  • Coordinate with product delivery and platform teams

Benefits

  • Equity
  • Medical, dental, and vision plans
  • Flexible PTO
  • Remote-first work
  • USD 1,200 home office and equipment stipend