GPU Compute and Bare Metal DPU Engineer
Bitdeer is a technology company providing Bitcoin mining solutions.
Funding history
Investors
Projects
About Bitdeer
Bitdeer provides full-spectrum Bitcoin mining and high-performance computing solutions, including SEALMINER mining equipment, Minerbase cooling containers, cloud mining, co-mining, mining management applications, mining rights marketplaces, and large-scale data center operations. The company also offers AI cloud infrastructure with GPU computing, model training and deployment capabilities, and turnkey AI data center solutions for enterprise customers and developers. Bitdeer is headquartered in Singapore and operates globally.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will own the lifecycle of bare metal GPU nodes from provisioning and delivery through operations break fix firmware management and decommissioning. You will automate node delivery manage GPU and server firmware improve fleet reliability participate in on call operations and define acceptance standards for new hardware and regions.
Requirements
- 3+ years in large scale bare metal or server fleet operations HPC or cloud infrastructure or 6+ years for Senior level
- Experience operating GPU servers at scale including driver CUDA and firmware management
- Strong Linux systems skills
- Experience with PXE IPMI Redfish OS imaging and automated provisioning
- Familiarity with DPU SmartNIC and bare metal networking
- Infrastructure automation skills with Ansible Terraform Python or Go
- Comfort with on call incident management and operational runbooks
- Multi region or large fleet operations experience is a plus
Responsibilities
- Own bare metal GPU node provisioning delivery operation break fix and decommissioning
- Build automated and repeatable node delivery pipelines
- Manage DPU SmartNIC and server firmware version baselines upgrades and validation
- Drive fleet reliability through incident response root cause analysis and health monitoring
- Participate in a 7 by 24 multi region on call rotation
- Build operational runbooks and tooling
- Partner with Storage Image and Network teams on provisioning and handoff
- Define bring up rack capacity and acceptance standards for GPU SKUs and regions
Benefits
- Remote work within San Jose or Austin
