Senior GPU Systems & Fabric Engineer
Bitdeer is a technology company providing Bitcoin mining solutions.
Funding history
Investors
Projects
About Bitdeer
Bitdeer provides full-spectrum Bitcoin mining and high-performance computing solutions, including SEALMINER mining equipment, Minerbase cooling containers, cloud mining, co-mining, mining management applications, mining rights marketplaces, and large-scale data center operations. The company also offers AI cloud infrastructure with GPU computing, model training and deployment capabilities, and turnkey AI data center solutions for enterprise customers and developers. Bitdeer is headquartered in Singapore and operates globally.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will connect physical GPU and network infrastructure with Kubernetes, expose NVIDIA and AMD hardware capabilities, configure RDMA and high speed networking, build DCGM based hardware remediation, implement MIG and vGPU slicing, tune Linux and AI runtimes, support topology aware placement, maintain bare metal standards, investigate performance issues, and mentor peers.
Requirements
- Bachelor’s or Master’s degree in Computer Science Electrical Engineering or a related field
- 5+ years of systems engineering experience
- Strong proficiency in Linux kernel internals C or Go
- Hands on experience with NVIDIA H100 or A100 GPU architectures CUDA and distributed networking
- Deep understanding of containerized environments and Kubernetes device plugin architecture
- Experience operating debugging and scaling bare metal systems in production or HPC environments
- Familiarity with Terraform Ansible and CI/CD pipelines
- Strong problem solving and communication skills
- Experience in high velocity engineering environments preferred
Responsibilities
- Architect NVIDIA and AMD GPU device plugin and Kubernetes Operator integrations
- Configure and optimize RDMA SR-IOV RoCEv2 and InfiniBand networking
- Build DCGM based hardware remediation pipelines
- Implement MIG and vGPU slicing for multi tenant workloads
- Tune Linux kernel parameters device drivers CUDA and NCCL
- Support topology aware placement and efficient data movement
- Define bare metal provisioning BIOS firmware and OS hardening standards
- Investigate performance issues across hardware fabric and software
- Mentor team members and maintain technical documentation
