Staff Slurm Cluster and HPC Engineer

Bitdeer Technologies Group is a technology company providing Bitcoin mining solutions, mining hardware, data-center infrastructure, and AI cloud services. It serves individual, institutional, and enterprise customers globally.

Singapore, SG
About Bitdeer Technologies Group

Bitdeer provides vertically integrated Bitcoin mining and high-performance computing services. Its operations include mining equipment procurement and manufacturing, datacenter design and construction, equipment management, daily mining operations, cloud mining, and mining-related services. The company also offers AI cloud infrastructure and high-performance computing powered by NVIDIA GPUs for AI and machine-learning workloads, serving customers across global markets.

View jobs by Bitdeer Technologies Group

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will own the architecture, operation, scheduling, and reliability of production Slurm clusters across bare-metal and virtualized GPU infrastructure. You will integrate Slurm with Kubernetes through Slinky, manage GPU fabrics and container runtimes, automate reproducible cluster deployments, monitor health and accounting, support customers, and lead incident response and reliability improvements.

Requirements

  • 8+ years in HPC systems or cloud infrastructure engineering.
  • 4+ years operating production Slurm clusters at 100+ GPU-node scale.
  • Deep experience with Slurm configuration scheduling accounting authentication and live upgrades.
  • Experience with NVIDIA drivers Fabric Manager DCGM MIG InfiniBand RoCEv2 GPUDirect RDMA and NCCL.
  • Production Kubernetes and operator or CRD experience.
  • Experience with Slurm-on-Kubernetes stacks such as Slinky or equivalent.
  • Experience with bare-metal provisioning firmware and BIOS lifecycle management.
  • Experience with KVM or QEMU or equivalent virtualized clusters.
  • Experience with Terraform and Ansible.
  • Knowledge of Lustre GPFS Spectrum Scale WEKA VAST or NFS.
  • Proficiency in Python and Bash.
  • Experience with multi-tenant authorization and isolation.
  • Clear written and verbal English communication.

Responsibilities

  • Design deploy and operate highly available Slurm clusters on bare metal and virtual machines.
  • Model GPU fabric topology and validate placement with NCCL and multi-node training tests.
  • Manage accounts partitions QOS fairshare preemption reservations and tenant TRES limits.
  • Implement Slinky Kubernetes resources and evaluate Slurm bridge co-scheduling.
  • Shift GPU capacity between Slurm and Kubernetes using cloud and power-save mechanisms.
  • Operate Pyxis Enroot OCI and containerd job paths with correct GPU and cgroup controls.
  • Support MPI PMIx modules Spack environments and customer images.
  • Build health checks and automatically drain unhealthy nodes and requeue jobs.
  • Own burn-in and acceptance testing for new racks.
  • Deliver infrastructure with Terraform Ansible golden images PXE Redfish and IPMI.
  • Instrument queue wait allocation efficiency utilization and job failures.
  • Reconcile Slurm GPU accounting with metering and invoicing.
  • Write runbooks and customer documentation and support enterprise onboarding.
  • Act as an escalation point and mentor platform engineers.
Staff Slurm Cluster and HPC Engineer at Bitdeer Technologies Group | JobStash