Senior GPU Cloud Storage Solutions Expert

Bitdeer is a technology company providing Bitcoin mining solutions.

Singapore, SG
About Bitdeer

Bitdeer provides full-spectrum Bitcoin mining and high-performance computing solutions, including SEALMINER mining equipment, Minerbase cooling containers, cloud mining, co-mining, mining management applications, mining rights marketplaces, and large-scale data center operations. The company also offers AI cloud infrastructure with GPU computing, model training and deployment capabilities, and turnkey AI data center solutions for enterprise customers and developers. Bitdeer is headquartered in Singapore and operates globally.

View jobs by Bitdeer

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will deploy and operate distributed storage for AI training and inference workloads. You will optimize storage performance, networking, isolation, capacity, migrations, and disaster recovery. You will instrument storage telemetry, help define fault prediction signals, and turn incident procedures into machine-executable remediation.

Requirements

  • 5+ years of enterprise or HPC storage operations experience
  • At least 2 years supporting AI or ML workloads
  • Experience deploying and operating at least two of WEKA, VAST Data, Ceph, and DDN/Lustre
  • Understanding of AI training I/O patterns
  • Experience with NFS over RDMA and NVMe-oF
  • Knowledge of GPU Direct Storage and RDMA-based data transfer
  • Proficiency with fio, IOR, and mdtest
  • Experience implementing multi-tenant storage isolation and QoS
  • Strong Linux systems knowledge including kernel tuning, filesystem internals, and block device management
  • Storage telemetry anomaly detection experience or aptitude
  • Runbook-as-code mindset

Responsibilities

  • Deploy and operate parallel and distributed storage systems
  • Design storage architectures for AI workload patterns
  • Implement multi-tenant storage isolation, QoS, quotas, and access controls
  • Configure and optimize GPU Direct Storage
  • Deploy and manage NFS over RDMA, NVMe-oF, storage fabrics, and Nvidia CMX
  • Diagnose and tune IOPS, throughput, and latency using storage benchmarks
  • Plan storage capacity for GPU cluster growth and customer workloads
  • Manage firmware, data migration, and disaster recovery procedures
  • Instrument storage telemetry for metrics, logs, and traces
  • Define storage-fault predictor signals and labels
  • Convert incident procedures into automated runbook-as-code remediation