SRE L1 Support/Cloud Platform Ops Engineers

Bitdeer Technologies Group is a technology company providing Bitcoin mining solutions, mining hardware, data-center infrastructure, and AI cloud services. It serves individual, institutional, and enterprise customers globally.

Singapore, SG
About Bitdeer Technologies Group

Bitdeer provides vertically integrated Bitcoin mining and high-performance computing services. Its operations include mining equipment procurement and manufacturing, datacenter design and construction, equipment management, daily mining operations, cloud mining, and mining-related services. The company also offers AI cloud infrastructure and high-performance computing powered by NVIDIA GPUs for AI and machine-learning workloads, serving customers across global markets.

View jobs by Bitdeer Technologies Group

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You monitor GPU clusters, networks, storage systems, and environmental sensors; respond to alerts and execute incident runbooks; triage and replace hardware; perform standard remediation; collect diagnostics for escalation; manage incident tickets; complete physical data center tasks; conduct shift handoffs; maintain runbooks; and assist with hardware deployment, firmware updates, and inventory management.

Requirements

  • 2+ years of experience in NOC, data center operations, or IT support
  • Basic Linux system administration
  • Familiarity with Prometheus, Grafana, Nagios, or equivalent monitoring tools
  • Experience with ServiceNow or Jira Service Management
  • Ability to perform rack and stack, cabling, and hardware replacement
  • Strong communication skills
  • Ability to work 8AM-8PM PST shifts with rotation

Responsibilities

  • Monitor GPU cluster health, network status, storage systems, and environmental sensors
  • Respond to alerts and execute runbooks for GPU, network, node, and storage incidents
  • Identify failed GPUs, NICs, PSUs, disks, and cables
  • Execute GPU resets, node drains and reboots, link reseating, and BMC recovery
  • Collect logs, DCGM output, network diagnostics, and hardware health reports for escalation
  • Manage incident tickets through resolution or escalation
  • Perform cable installation, hardware swap-outs, rack and stack, and labeling
  • Execute shift handoffs with the APAC operations team
  • Maintain and update operational runbooks
  • Assist with hardware deployment, firmware updates, and inventory management

Benefits

  • Attractive welfare benefits