Search...

SRE L1 Support/Cloud Platform Ops Engineer

Bitdeer logo
Bitdeer

Bitdeer is a NASDAQ-listed (BTDR) high-performance computing and Bitcoin mining company headquartered in Singapore. It provides end-to-end Bitcoin mining solutions (mining hardware like SEALMINER, Minerbase containers, cloud mining, hosting, and mining farm/data center operations) as well as AI Cloud services offering GPU compute (NVIDIA GB200 NVL72, B200, H200, H100) for AI training and deployment. Its customers range from individual and institutional Bitcoin miners to enterprises and developers needing scalable AI/ML compute infrastructure.

Singapore, SG
About Bitdeer

Bitdeer Technologies Group is a NASDAQ-listed (ticker: BTDR) technology company headquartered in Singapore that describes itself as a "world-leading" high-performance computing platform and Bitcoin mining services provider. The company is vertically integrated across the value chain, spanning IC design and hardware manufacturing (its own SEALMINER ASIC miners and Minerbase mobile cooling containers), infrastructure construction and cloud mining, and artificial intelligence/high-performance computing. Bitdeer handles the full range of mining-related processes for its customers, including equipment procurement, transport logistics, datacenter design and construction, equipment management, and daily operations, and offers institutional services, a hash rate market, and a miner rights trading marketplace via its mobile apps (Bitdeer App and Minerplus App). Since 2013 Bitdeer has built more than 30 data centers globally and currently operates 9 large-scale data centers (including one of North America's largest) with roughly 3GW of diversified energy capacity and tens of exahashes of managed hash rate, with major operations in the United States, Bhutan, Norway, Canada, Malaysia, and Ethiopia. Since 2023 it has also been expanding a global AI infrastructure business (Bitdeer AI Cloud), powered by thousands of NVIDIA GPUs (including GB200 NVL72 and B200, with GB300 NVL72 and B300 planned), offering turnkey AI datacenter solutions and GPU cloud compute for AI training and deployment starting at around $2/hour. Bitdeer serves both individual/retail Bitcoin miners and institutional clients, as well as AI developers and enterprises seeking scalable, energy-efficient compute.

View jobs by Bitdeer

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You provide front-line monitoring and incident response for GPU data centers during an 8AM–8PM PST shift. You execute runbooks, triage hardware and infrastructure issues, collect diagnostics, manage tickets, perform physical data center tasks, coordinate handoffs, update procedures, and supply structured incident data for automation.

Requirements

  • 2+ years in a NOC, data center operations, or IT support role
  • Basic Linux system administration
  • Familiarity with Prometheus, Grafana, Nagios, or equivalent monitoring tools
  • Experience with ServiceNow or Jira Service Management
  • Ability to perform rack and stack, cabling, and hardware replacement
  • Strong communication skills for handoffs, incident documentation, and escalation
  • Ability to work the 8AM–8PM PST shift schedule with 12-hour shifts and rotation
  • Curiosity about automation
  • Comfort with structured data and incident ticketing

Responsibilities

  • Monitor GPU cluster health, network status, storage systems, and environmental sensors
  • Respond to alerts and execute runbooks for common incidents
  • Triage failed GPUs, NICs, PSUs, disks, and cables
  • Perform GPU resets, node drains and reboots, link reseating, and BMC recovery
  • Collect logs, DCGM output, network diagnostics, and hardware health reports
  • Manage incident tickets through resolution or escalation
  • Install cables, swap hardware, rack and stack equipment, and label components
  • Perform structured shift handoffs with the APAC operations team
  • Maintain and update operational runbooks
  • Assist with hardware deployment, firmware updates, and inventory management
  • Tag and document novel incidents for automation
  • Provide structured handoff notes