Data Center Operations Engineer

Bitdeer is a technology company providing Bitcoin mining solutions.

Singapore, SG
About Bitdeer

Bitdeer provides full-spectrum Bitcoin mining and high-performance computing solutions, including SEALMINER mining equipment, Minerbase cooling containers, cloud mining, co-mining, mining management applications, mining rights marketplaces, and large-scale data center operations. The company also offers AI cloud infrastructure with GPU computing, model training and deployment capabilities, and turnkey AI data center solutions for enterprise customers and developers. Bitdeer is headquartered in Singapore and operates globally.

View jobs by Bitdeer

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will operate and maintain data center infrastructure to ensure reliable service. You will install, commission, monitor, troubleshoot, and repair AI and HPC clusters, servers, storage, switches, and cabling. You will provision servers, install operating systems, perform preventive maintenance, document operations, respond to incidents, and support new deployments while participating in a rotating shift schedule.

Requirements

  • Bachelor's degree or above in a relevant technical discipline
  • Understanding of data center infrastructure and server hardware architecture
  • Familiarity with AI cluster systems, GPU servers, x86 servers, storage servers, Ethernet networks, or InfiniBand networks
  • Knowledge of server hardware components, BMC/IPMI, and firmware management
  • Knowledge of TCP/IP, Ethernet, VLAN, LACP, InfiniBand, or RoCE
  • Understanding of structured cabling systems including DAC, AOC, optical fiber, MPO, and LC connectors
  • Basic Linux administration skills including system monitoring, troubleshooting, service management, log analysis, network troubleshooting, and shell scripting
  • Willingness to work a 24x7 two-shift rotation including night shifts

Responsibilities

  • Operate and maintain data center infrastructure
  • Install, rack, cable, commission, maintain, and troubleshoot AI and HPC cluster infrastructure
  • Monitor the health of servers, GPUs, storage, networking devices, and associated infrastructure
  • Replace hardware, upgrade BIOS, BMC, and firmware, and perform hardware diagnostics
  • Provision servers, install operating systems, expand clusters, validate networks, and conduct burn-in testing
  • Troubleshoot server, GPU, storage, network, switch, and cabling issues
  • Perform routine inspections and preventive maintenance and maintain operational records
  • Execute incident response procedures and escalate and resolve incidents
  • Prepare shift handover reports and maintain operation documents, SOPs, and incident reports
  • Support new deployments and operational improvements with engineering, network, and infrastructure teams
  • Participate in a two-shift rotation including night shifts, weekends, and holidays