Staff Slurm Cluster & HPC Engineer

Bitdeer Technologies Group is a technology company providing Bitcoin mining solutions, mining hardware, data-center infrastructure, and AI cloud services. It serves individual, institutional, and enterprise customers globally.

Singapore, SG
About Bitdeer Technologies Group

Bitdeer provides vertically integrated Bitcoin mining and high-performance computing services. Its operations include mining equipment procurement and manufacturing, datacenter design and construction, equipment management, daily mining operations, cloud mining, and mining-related services. The company also offers AI cloud infrastructure and high-performance computing powered by NVIDIA GPUs for AI and machine-learning workloads, serving customers across global markets.

View jobs by Bitdeer Technologies Group

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

Design, deploy, and operate production Slurm clusters across bare-metal and virtualized environments. Implement topology-aware GPU scheduling, multi-tenant policies, Kubernetes integration, elastic capacity management, containerized job execution, health checks, automation, observability, accounting, and billing integration. Document operations, support enterprise customers, handle incidents, and mentor platform engineers.

Requirements

  • 8+ years in HPC, systems, or cloud infrastructure engineering
  • 4+ years operating production Slurm clusters at 100+ GPU-node scale
  • Deep hands-on experience with Slurm configuration, accounting, authentication, and upgrades
  • Experience with NVIDIA drivers, Fabric Manager, DCGM, MIG, InfiniBand, RoCEv2, GPUDirect RDMA, and NCCL
  • Production Kubernetes and operator or CRD experience
  • Experience with bare-metal provisioning, firmware lifecycle management, and virtualized compute
  • Experience with Terraform and Ansible
  • Knowledge of Lustre, GPFS, WEKA, VAST, or NFS
  • Proficiency in Python and Bash
  • Multi-tenant security and fail-closed authorization experience
  • Clear written and verbal communication in English

Responsibilities

  • Design, deploy, and operate production Slurm clusters
  • Manage Slurm high availability, authentication, accounting, and live upgrades
  • Model GPU fabrics for topology-aware scheduling
  • Validate placement quality with NCCL and multi-node training tests
  • Manage tenant accounts, partitions, QOS, fairshare, preemption, reservations, and TRES limits
  • Lead Slinky slurm-operator implementation
  • Evaluate Slurm and Kubernetes co-scheduling with slurm-bridge
  • Shift GPU nodes between Slurm and Kubernetes capacity
  • Operate Pyxis, Enroot, OCI, and containerd job paths
  • Build GPU cluster health checks and automate draining and job requeue
  • Deliver reproducible clusters with Terraform, Ansible, PXE, Redfish, and IPMI
  • Instrument queue, allocation, utilization, and job-failure metrics
  • Reconcile GPU-hours with metering and invoicing
  • Write runbooks and customer documentation
  • Support enterprise customers and mentor platform engineers

Benefits

  • Welfare benefits
  • Training and mentoring
  • Developmental opportunities
  • Personal accountability and autonomy
  • Fast growth and learning opportunities