Senior GPU Cloud Storage Solutions Expert SRE SME

Bitdeer is a technology company providing Bitcoin mining solutions.

Singapore, SG
About Bitdeer

Bitdeer provides full-spectrum Bitcoin mining and high-performance computing solutions, including SEALMINER mining equipment, Minerbase cooling containers, cloud mining, co-mining, mining management applications, mining rights marketplaces, and large-scale data center operations. The company also offers AI cloud infrastructure with GPU computing, model training and deployment capabilities, and turnkey AI data center solutions for enterprise customers and developers. Bitdeer is headquartered in Singapore and operates globally.

View jobs by Bitdeer

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will deploy, operate, and optimize distributed storage systems for AI workloads. You will manage storage isolation, QoS, networking, capacity, performance tuning, observability, disaster recovery, and automated incident remediation.

Requirements

  • 5+ years of enterprise or HPC storage operations experience
  • At least 2 years supporting AI or ML workloads
  • Deployment and operations experience with at least two of WEKA, VAST Data, Ceph, and DDN Lustre
  • Understanding of AI training I/O patterns
  • Experience with NFS over RDMA and NVMe-oF
  • Knowledge of GPU Direct Storage and RDMA data transfer
  • Storage benchmarking and tuning experience with fio, IOR, and mdtest
  • Experience implementing multi-tenant storage isolation and QoS
  • Linux systems knowledge including kernel tuning, filesystem internals, and block device management
  • Storage telemetry anomaly detection or feature-label design experience
  • Runbook-as-code mindset

Responsibilities

  • Deploy and operate distributed storage systems including WEKA, VAST Data, Ceph, and DDN Lustre
  • Design storage architectures for AI workload patterns
  • Implement multi-tenant storage isolation, QoS, quotas, and access controls
  • Configure GPU Direct Storage and high-performance storage networking
  • Manage NFS over RDMA, NVMe-oF, storage fabrics, and Nvidia CMX
  • Diagnose and tune storage performance using fio, IOR, and mdtest
  • Plan storage capacity and manage firmware, data migration, and disaster recovery
  • Instrument storage telemetry for observability
  • Convert storage incidents and SOPs into automated runbooks