Data Center Operations Engineer
Bitdeer is a technology company providing Bitcoin mining solutions.
Funding history
Investors
Projects
About Bitdeer
Bitdeer provides full-spectrum Bitcoin mining and high-performance computing solutions, including SEALMINER mining equipment, Minerbase cooling containers, cloud mining, co-mining, mining management applications, mining rights marketplaces, and large-scale data center operations. The company also offers AI cloud infrastructure with GPU computing, model training and deployment capabilities, and turnkey AI data center solutions for enterprise customers and developers. Bitdeer is headquartered in Singapore and operates globally.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will operate and maintain AI data center infrastructure to ensure reliable, highly available services. You will install, cable, commission, monitor, troubleshoot, repair, provision, and document cluster systems. You will respond to incidents, complete shift handovers, support deployments, and work a rotating 24x7 schedule.
Requirements
- Bachelor's degree or above in a related technical discipline
- Understanding of data center infrastructure and server hardware architecture
- Familiarity with AI/HPC clusters, GPU servers, x86 servers, storage servers, Ethernet, or InfiniBand networks
- Knowledge of server hardware components, BMC/IPMI, and firmware management
- Familiarity with TCP/IP, Ethernet, VLAN, LACP, InfiniBand, or RoCE
- Understanding of structured cabling systems
- Basic Linux administration, monitoring, troubleshooting, service management, log analysis, network troubleshooting, and shell scripting
- Willingness to work a 24x7 shift rotation
Responsibilities
- Operate and maintain data center infrastructure
- Install, rack, cable, commission, maintain, and troubleshoot AI/HPC cluster infrastructure
- Monitor the health of servers, GPUs, storage, network devices, and related infrastructure
- Replace hardware, upgrade BIOS, BMC, and firmware, and perform hardware diagnostics
- Support server provisioning, operating-system installation, cluster expansion, network validation, and burn-in testing
- Troubleshoot server, GPU, storage, network, switch, and cabling issues
- Perform preventive maintenance and maintain operational records and logs
- Execute incident-response procedures and escalate incidents
- Prepare shift handovers, SOPs, and incident reports
- Support new deployments and operational improvements
- Participate in a three-shift rotation including nights, weekends, and holidays
