AI Infrastructure Engineer

Bitdeer Technologies Group is a technology company providing Bitcoin mining solutions, mining hardware, data-center infrastructure, and AI cloud services. It serves individual, institutional, and enterprise customers globally.

Singapore, SG
About Bitdeer Technologies Group

Bitdeer provides vertically integrated Bitcoin mining and high-performance computing services. Its operations include mining equipment procurement and manufacturing, datacenter design and construction, equipment management, daily mining operations, cloud mining, and mining-related services. The company also offers AI cloud infrastructure and high-performance computing powered by NVIDIA GPUs for AI and machine-learning workloads, serving customers across global markets.

View jobs by Bitdeer Technologies Group

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

Architect, deploy, and maintain infrastructure for large-scale AI compute environments in a high-density AI data center, including GPU clusters, high-performance networking, storage, provisioning automation, and infrastructure performance optimization.

Requirements

  • Degree in Computer Science, Data Engineering, or a related technical field; a master's degree is preferred.
  • Strong experience with Linux administration, containerization, and GPU infrastructure.
  • Experience with Kubernetes or Slurm.
  • Familiarity with CUDA, NCCL, and Triton Inference Server.
  • Understanding of NVIDIA GB300 and VR NVL72 Scalable Units.
  • Experience in HPC, AI infrastructure, or large-scale distributed compute environments.
  • Experience with InfiniBand, RoCE v2, or high-performance networking architectures.
  • Familiarity with Lustre, BeeGFS, or WekaIO.
  • Experience with Terraform, Ansible, or other Infrastructure as Code frameworks.
  • NVIDIA, Kubernetes, or cloud infrastructure certifications.

Responsibilities

  • Deploy and manage large-scale GPU clusters using Kubernetes or Slurm.
  • Optimize high-speed, low-latency networking for distributed compute.
  • Plan and monitor rack density across AI infrastructure.
  • Implement and maintain high-throughput storage systems for GPU-intensive workloads.
  • Automate infrastructure provisioning and configuration using Terraform, Ansible, or other Infrastructure as Code tools.
  • Troubleshoot and optimize compute, networking, and storage performance in a mission-critical environment.