Senior Slurm Cluster and HPC Engineer

Bitdeer is a technology company providing Bitcoin mining solutions.

Singapore, SG
About Bitdeer

Bitdeer provides full-spectrum Bitcoin mining and high-performance computing solutions, including SEALMINER mining equipment, Minerbase cooling containers, cloud mining, co-mining, mining management applications, mining rights marketplaces, and large-scale data center operations. The company also offers AI cloud infrastructure with GPU computing, model training and deployment capabilities, and turnkey AI data center solutions for enterprise customers and developers. Bitdeer is headquartered in Singapore and operates globally.

View jobs by Bitdeer

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will design, deploy, and operate production Slurm clusters across bare metal and virtual machines. You will manage scheduling policies, GPU-fabric placement, Kubernetes integration, elastic capacity, job runtimes, health checks, automation, observability, billing integration, customer support, and technical documentation.

Requirements

  • 8+ years of HPC, systems, or cloud infrastructure engineering experience.
  • 4+ years operating production Slurm clusters at 100+ GPU-node scale.
  • Hands-on expertise with Slurm configuration, accounting, authentication, scheduling policies, and live upgrades.
  • Knowledge of NVIDIA drivers, Fabric Manager, DCGM, MIG, InfiniBand, RoCEv2, GPUDirect RDMA, and NCCL.
  • Production Kubernetes experience and knowledge of the operator and CRD pattern.
  • Experience with a Slurm-on-Kubernetes stack.
  • Experience delivering bare-metal and virtualized compute infrastructure.
  • Experience with Terraform and Ansible automation.
  • Knowledge of shared and parallel storage for AI workloads.
  • Proficiency in Python and Bash.
  • Knowledge of multi-tenant security and fail-closed authorization.
  • Written and verbal English communication skills and enterprise customer experience.

Responsibilities

  • Design, deploy, and operate production Slurm clusters on bare metal and virtual machines.
  • Manage Slurm high availability, authentication, configuration, and rolling upgrades.
  • Model GPU-fabric topology and validate workload placement.
  • Own multi-tenant scheduling policies, entitlements, and fail-closed access controls.
  • Lead Slinky slurm-operator implementation and evaluate Slurm co-scheduling with Kubernetes.
  • Shift GPU capacity between Slurm batch queues and Kubernetes inference capacity.
  • Operate container job runtimes and support MPI, PMIx, modules, Spack, and customer images.
  • Build health-check systems with automatic node draining and job requeue.
  • Deliver reproducible cluster infrastructure through Terraform, Ansible, golden images, and bare-metal provisioning.
  • Instrument cluster metrics, configure accounting and billing, and reconcile GPU-hour usage.
  • Write runbooks and tenant documentation, support enterprise customers, and mentor platform engineers.