HPC Storage Engineer

Runpod is an AI developer cloud providing GPU compute, serverless inference, Pods, and clusters for building, training, fine-tuning, deploying, and scaling AI workloads.

Recently fundedCompany intelligence
San Francisco, United States
About Runpod

Runpod Inc. operates a globally distributed GPU cloud platform for AI developers, offering on-demand GPU infrastructure and serverless services across the AI development lifecycle.

View jobs by Runpod

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will design, operate, and scale distributed storage systems. You will optimize storage and network performance, automate operational workflows, build observability, lead migrations and capacity expansions, and participate in on-call response.

Requirements

  • 8+ years of infrastructure, storage, or systems engineering experience
  • Experience owning production storage at scale
  • Experience with a distributed storage system
  • Linux internals and storage-stack knowledge
  • Experience building or operating S3-compatible object storage
  • Networking experience for storage workloads
  • Production coding proficiency in Go, Python, Rust, or similar
  • Observability tooling experience
  • Production performance analysis and debugging experience

Responsibilities

  • Own capacity, durability, availability, and performance for network volumes, local NVMe, and S3-compatible object storage
  • Tune the end-to-end storage I/O path
  • Diagnose storage performance problems end to end
  • Lead capacity expansions, hardware refreshes, migrations, and rebalances
  • Design and tune storage network paths with networking teams
  • Optimize RDMA/RoCE and high-speed storage fabrics
  • Write production code and automation for storage services and workflows
  • Build and extend control-plane, S3-compatible, CSI, Kubernetes, vendor, and cloud APIs
  • Automate manual storage operations
  • Participate in code review, testing, and CI
  • Instrument the storage fleet and build dashboards, SLOs, and alerts
  • Participate in storage on-call rotations and post-incident follow-through

Benefits

  • Equity through stock options
  • Medical, dental, and vision plans
  • Flexible PTO
  • Remote-first work
  • Home office and equipment stipend of $1,200