Senior GPU Systems & Fabric Engineer

Bitdeer Technologies Group is a technology company providing Bitcoin mining solutions, SEALMINER mining equipment, cloud mining, AI cloud infrastructure, high-performance computing, and data-center services. Its customers include miners, developers, AI and machine-learning users, and institutional customers.

Singapore, SG
About Bitdeer Technologies Group

Bitdeer Technologies Group operates across Bitcoin mining, AI cloud, and data-center infrastructure. Its offerings include proprietary SEALMINER equipment, Minerbase cooling systems, cloud-mining plans, co-mining services, mining-farm management tools, and mobile apps. The company also provides GPU-based AI and high-performance computing services for model training and deployment, as well as turnkey AI data-center solutions. Bitdeer is headquartered in Singapore and operates globally, with major operations in the United States, Bhutan, Norway, Canada, Malaysia, and Ethiopia.

View jobs by Bitdeer Technologies Group

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

Architect GPU device plugin and Kubernetes Operator integrations; optimize RDMA, SR-IOV, RoCEv2, and InfiniBand networking; build automated GPU and NIC remediation pipelines; manage MIG and vGPU slicing; tune kernels, drivers, CUDA, and NCCL; support topology-aware placement and data movement; define bare-metal provisioning and hardening standards; investigate performance issues; and mentor team members.

Requirements

  • Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, or a related field
  • 5+ years of systems engineering experience
  • Strong proficiency in Linux kernel internals, C, or Go
  • Hands-on experience with NVIDIA H100/A100 GPU architectures
  • Experience with CUDA runtimes and distributed networking
  • Understanding of containerized environments and Kubernetes device plugin architecture
  • Experience operating, debugging, and scaling bare-metal systems
  • Familiarity with Terraform, Ansible, and CI/CD pipelines
  • Strong problem-solving skills
  • Excellent communication skills

Responsibilities

  • Architect and maintain NVIDIA and AMD GPU device plugin and Kubernetes Operator integrations
  • Configure and optimize RDMA, SR-IOV, RoCEv2, and InfiniBand networking
  • Build automated hardware remediation pipelines using DCGM telemetry
  • Manage MIG and vGPU slicing technologies
  • Tune kernel parameters, device drivers, CUDA, and NCCL
  • Collaborate on topology-aware placement and data movement
  • Define standards for bare-metal provisioning, BIOS and firmware updates, and OS hardening
  • Lead investigations into hardware, fabric, and software performance issues
  • Mentor team members and drive documentation standards

Benefits

  • A culture that values authenticity and diversity of thoughts and backgrounds
  • An inclusive and respectable environment with open workspaces and exciting start-up spirit
  • Fast-growing company with the chance to network with industrial pioneers and enthusiasts
  • Ability to contribute directly and make an impact on the future of the digital asset industry
  • Involvement in new projects and developing processes and systems
  • Personal accountability, autonomy, fast growth, and learning opportunities
  • Attractive welfare benefits and developmental opportunities such as training and mentoring