AI Storage Solutions Expert
Bitdeer Technologies Group is a technology company providing Bitcoin mining solutions, mining hardware, data-center infrastructure, and AI cloud services. It serves individual, institutional, and enterprise customers globally.
Funding history
Investors
Projects
About Bitdeer Technologies Group
Bitdeer provides vertically integrated Bitcoin mining and high-performance computing services. Its operations include mining equipment procurement and manufacturing, datacenter design and construction, equipment management, daily mining operations, cloud mining, and mining-related services. The company also offers AI cloud infrastructure and high-performance computing powered by NVIDIA GPUs for AI and machine-learning workloads, serving customers across global markets.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
Bitdeer Technologies Group is seeking an AI Storage Solutions Expert to own the storage layer for AI training and inference across US data centers. The role covers AI-optimized storage architecture, tenant isolation and QoS, high-performance storage networking, performance tuning, capacity planning, firmware and migration management, telemetry instrumentation, fault prediction, and runbook automation.
Requirements
- 5+ years in enterprise or HPC storage operations, including at least 2 years supporting AI/ML workloads.
- Hands-on deployment and operations experience with at least two of WEKA, VAST Data, Ceph, or DDN/Lustre.
- Strong understanding of AI training I/O patterns.
- Experience with NFS over RDMA and NVMe-oF.
- Knowledge of GPU Direct Storage and RDMA-based data transfer.
- Proficiency with fio, IOR, and mdtest.
- Experience implementing multi-tenant storage isolation and QoS.
- Strong Linux systems knowledge.
- Experience with storage or I/O telemetry anomaly detection, or ability to define required labels and features.
- Runbook-as-code mindset.
Responsibilities
- Deploy and operate parallel and distributed storage systems.
- Design storage architectures for checkpoint I/O bursts, sequential dataset reads, and inference KV caches.
- Implement multi-tenant storage isolation with QoS, quotas, and access controls.
- Configure and optimize GPU Direct Storage.
- Deploy and manage NFS over RDMA, NVMe-oF, high-speed storage fabrics, and Nvidia CMX.
- Diagnose and tune storage performance using fio, IOR, and mdtest.
- Maintain failure-mode runbooks and convert incidents into automated remediation.
- Plan storage capacity for GPU cluster growth and customer workloads.
- Manage firmware, data migration, and disaster recovery procedures.
- Instrument storage telemetry and define signals and labels for storage-fault prediction.
