GPU Compute and Bare Metal DPU Engineer
Bitdeer is a technology company providing Bitcoin mining solutions.
Maintainer signals as of 9/25/2026
Funding history
Investors
Projects
About Bitdeer
Bitdeer provides full-spectrum Bitcoin mining and high-performance computing solutions, including SEALMINER mining equipment, Minerbase cooling containers, cloud mining, co-mining, mining management applications, mining rights marketplaces, and large-scale data center operations. The company also offers AI cloud infrastructure with GPU computing, model training and deployment capabilities, and turnkey AI data center solutions for enterprise customers and developers. Bitdeer is headquartered in Singapore and operates globally.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
Own the complete lifecycle of bare-metal GPU nodes across multiple regions, from provisioning and delivery through operation, break-fix, and decommissioning. Automate node-delivery pipelines, manage DPU, SmartNIC, and server firmware, improve fleet reliability, lead incident response and root-cause analysis, define GPU infrastructure standards, and participate in a multi-region on-call rotation.
Requirements
- 3+ years in large-scale bare-metal or server-fleet operations, HPC, or cloud infrastructure; 6+ years for Senior level.
- Experience operating GPU servers at scale.
- GPU driver, CUDA, and firmware management experience.
- Strong Linux systems skills.
- Experience with PXE, IPMI, Redfish, OS imaging, and automated provisioning.
- Familiarity with DPU, SmartNIC, and bare-metal networking.
- Ansible, Terraform, Python, or Go experience.
- On-call, incident management, and operational runbook experience.
Responsibilities
- Own bare-metal GPU node provisioning, delivery, operation, break-fix, and decommissioning.
- Build and operate automated node-delivery pipelines.
- Manage DPU, SmartNIC, BMC, BIOS, NIC, and GPU firmware.
- Drive fleet reliability and reduce MTTR.
- Lead incident response and root-cause analysis.
- Improve hardware-health monitoring.
- Build runbooks and tooling to reduce manual work.
- Partner with Storage, Image, and Network teams on provisioning and handoff.
- Define bring-up, rack, capacity, and acceptance standards for GPU SKUs and data-center regions.
- Participate in a multi-region on-call rotation.
Benefits
- Welfare benefits
- Inclusive and respectful work environment
- Training and mentoring
- Developmental opportunities
- Personal accountability, autonomy, and fast growth
