Staff Engineer Distributed Storage and HPC and AI Infrastructure

Together AI operates an AI-native cloud platform for open and custom AI models.

San Francisco, United States
About Together AI

Together AI provides production AI infrastructure spanning inference, accelerated compute, model training and fine-tuning, and secure code sandboxes for AI development.

View jobs by Together AI

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will operate, scale, and optimize multi-petabyte storage systems for AI training and inference workloads. You will develop storage strategy, integrate distributed and object storage technologies, implement caching and tiering architectures, tune storage isolation and data paths, and build Kubernetes storage operators. You will use benchmarking and profiling to improve throughput, performance, and storage infrastructure at cluster scale.

Requirements

  • 8+ years of storage engineering experience at multi-petabyte scale
  • Experience deploying and operating high-performance storage for GPU and HPC clusters
  • Kubernetes and cloud-native storage experience
  • Go and Python programming skills
  • Technical leadership experience
  • Distributed storage expertise with Ceph, WekaFS, Lustre, Vast, GPFS, or similar filesystems
  • Object storage experience with S3, MinIO, Ceph, or R2
  • CSI drivers, StatefulSets, PersistentVolumes, storage operators, and custom controllers
  • GPU storage optimization and RDMA or InfiniBand networking
  • Terraform, Ansible, Helm, GitOps, and ArgoCD
  • Linux filesystems, LVM, NVMe optimization, and RAID
  • Prometheus, Grafana, and Thanos

Responsibilities

  • Architect and implement storage strategy and roadmap
  • Scale multi-petabyte AI and machine learning storage systems
  • Integrate Vast, Weka, and Ceph storage technologies
  • Develop caching and tiered-storage architectures
  • Tune L2 and L3 storage isolation for multi-tenancy
  • Build Kubernetes storage operators and controllers
  • Engineer high-throughput data paths and parallel filesystems
  • Optimize data paths through benchmarking and profiling
  • Contribute code to open-source storage projects and internal tooling
Staff Engineer Distributed Storage and HPC and AI Infrastructure at Together AI | JobStash