Search...

Staff Slurm Cluster & HPC Engineer

Bitdeer logo
Bitdeer

Bitdeer is a NASDAQ-listed (BTDR) high-performance computing and Bitcoin mining company headquartered in Singapore. It provides end-to-end Bitcoin mining solutions (mining hardware like SEALMINER, Minerbase containers, cloud mining, hosting, and mining farm/data center operations) as well as AI Cloud services offering GPU compute (NVIDIA GB200 NVL72, B200, H200, H100) for AI training and deployment. Its customers range from individual and institutional Bitcoin miners to enterprises and developers needing scalable AI/ML compute infrastructure.

Singapore, SG
About Bitdeer

Bitdeer Technologies Group is a NASDAQ-listed (ticker: BTDR) technology company headquartered in Singapore that describes itself as a "world-leading" high-performance computing platform and Bitcoin mining services provider. The company is vertically integrated across the value chain, spanning IC design and hardware manufacturing (its own SEALMINER ASIC miners and Minerbase mobile cooling containers), infrastructure construction and cloud mining, and artificial intelligence/high-performance computing. Bitdeer handles the full range of mining-related processes for its customers, including equipment procurement, transport logistics, datacenter design and construction, equipment management, and daily operations, and offers institutional services, a hash rate market, and a miner rights trading marketplace via its mobile apps (Bitdeer App and Minerplus App). Since 2013 Bitdeer has built more than 30 data centers globally and currently operates 9 large-scale data centers (including one of North America's largest) with roughly 3GW of diversified energy capacity and tens of exahashes of managed hash rate, with major operations in the United States, Bhutan, Norway, Canada, Malaysia, and Ethiopia. Since 2023 it has also been expanding a global AI infrastructure business (Bitdeer AI Cloud), powered by thousands of NVIDIA GPUs (including GB200 NVL72 and B200, with GB300 NVL72 and B300 planned), offering turnkey AI datacenter solutions and GPU cloud compute for AI training and deployment starting at around $2/hour. Bitdeer serves both individual/retail Bitcoin miners and institutional clients, as well as AI developers and enterprises seeking scalable, energy-efficient compute.

View jobs by Bitdeer

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You design, deploy, and operate production Slurm clusters across bare-metal and virtualized environments. You implement topology-aware GPU scheduling, multi-tenant policies, Slurm integration with Kubernetes, elastic capacity management, containerized job execution, health checks, automation, observability, accounting, and billing integration. You also document operations, support enterprise customers, handle incidents, and mentor platform engineers.

Requirements

  • 8+ years in HPC, systems, or cloud infrastructure engineering
  • 4+ years operating production Slurm clusters at 100+ GPU-node scale
  • Deep hands-on experience with Slurm configuration, accounting, authentication, and upgrades
  • Experience with NVIDIA drivers, Fabric Manager, DCGM, MIG, InfiniBand, RoCEv2, GPUDirect RDMA, and NCCL
  • Production Kubernetes and operator or CRD experience
  • Experience with bare-metal provisioning, firmware lifecycle management, and virtualized compute
  • Experience with Terraform and Ansible
  • Knowledge of Lustre, GPFS, WEKA, VAST, or NFS
  • Proficiency in Python and Bash
  • Multi-tenant security and fail-closed authorization experience
  • Clear written and verbal communication in English

Responsibilities

  • Design, deploy, and operate production Slurm clusters
  • Manage Slurm high availability, authentication, accounting, and live upgrades
  • Model GPU fabrics for topology-aware scheduling
  • Validate placement quality with NCCL and multi-node training tests
  • Manage tenant accounts, partitions, QOS, fairshare, preemption, reservations, and TRES limits
  • Lead Slinky slurm-operator implementation
  • Evaluate Slurm and Kubernetes co-scheduling with slurm-bridge
  • Shift GPU nodes between Slurm and Kubernetes capacity
  • Operate Pyxis, Enroot, OCI, and containerd job paths
  • Build GPU cluster health checks and automate draining and job requeue
  • Deliver reproducible clusters with Terraform, Ansible, PXE, Redfish, and IPMI
  • Instrument queue, allocation, utilization, and job-failure metrics
  • Reconcile GPU-hours with metering and invoicing
  • Write runbooks and customer documentation
  • Support enterprise customers and mentor platform engineers

Benefits

  • Welfare benefits
Staff Slurm Cluster & HPC Engineer at Bitdeer | JobStash