Member of Technical Staff Storage Infrastructure

Prime Intellect is an active AI infrastructure company building an open stack for training, deploying, evaluating, and continuously improving agentic models.

Maintainer signals as of 9/2/2026

San Francisco, United States
About Prime Intellect

Prime Intellect, Inc. operates a full-stack AI platform combining hosted reinforcement-learning training, evaluations, inference, secure sandboxes, GPU compute, and open-source research tooling. Its current positioning is the Open Superintelligence Stack, serving AI companies, researchers, and developers.

View jobs by Prime Intellect

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will design and operate storage systems for datasets, checkpoints, artifacts, and research workflows. You will tune storage platforms, benchmark performance, automate operations, test recovery procedures, and improve reliability, security, and observability.

Requirements

  • 3+ years building or operating production distributed storage systems
  • Experience with Lustre, BeeGFS, Ceph, GPFS, or another parallel, distributed filesystem or object storage platform
  • Linux administration and performance troubleshooting skills
  • Experience automating infrastructure operations with Python, Go, Bash, or similar languages
  • Understanding of storage failure modes, data integrity, consistency, replication, and recovery
  • Knowledge of block, file, object storage, NVMe, filesystem tuning, I/O profiling, authentication, authorization, and encryption

Responsibilities

  • Design and operate storage architectures for datasets, checkpoints, inference artifacts, and research workflows
  • Deploy and tune parallel filesystems, object storage, and NVMe caching
  • Benchmark throughput, latency, metadata performance, and concurrent access
  • Build storage provisioning, capacity planning, lifecycle management, and operational automation
  • Design and test replication, recovery, backup, and failure-handling procedures
  • Diagnose performance and reliability issues across applications, clients, networks, filesystems, and devices
  • Implement access controls, tenant separation, quotas, monitoring, and runbooks

Benefits

  • Equity incentives