SRE L1 Support/Cloud Platform Ops Engineer
Bitdeer is a technology company providing Bitcoin mining solutions.
Maintainer signals as of 9/25/2026
Funding history
Investors
Projects
About Bitdeer
Bitdeer provides full-spectrum Bitcoin mining and high-performance computing solutions, including SEALMINER mining equipment, Minerbase cooling containers, cloud mining, co-mining, mining management applications, mining rights marketplaces, and large-scale data center operations. The company also offers AI cloud infrastructure with GPU computing, model training and deployment capabilities, and turnkey AI data center solutions for enterprise customers and developers. Bitdeer is headquartered in Singapore and operates globally.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
Provide front-line monitoring and incident response for GPU data centers during an 8AM–8PM PST shift. Execute runbooks, triage hardware and infrastructure issues, collect diagnostics, manage tickets, perform physical data center tasks, coordinate handoffs, update procedures, and supply structured incident data for automation.
Requirements
- 2+ years in a NOC, data center operations, or IT support role.
- Basic Linux system administration.
- Familiarity with Prometheus, Grafana, Nagios, or equivalent monitoring tools.
- Experience with ServiceNow or Jira Service Management.
- Ability to perform rack and stack, cabling, and hardware replacement.
- Strong communication skills for handoffs, incident documentation, and escalation.
- Ability to work the 8AM–8PM PST shift schedule with 12-hour shifts and rotation.
- Curiosity about automation.
- Comfort with structured data and incident ticketing.
Responsibilities
- Monitor GPU cluster health, network status, storage systems, and environmental sensors.
- Respond to alerts and execute runbooks for common incidents.
- Triage failed GPUs, NICs, PSUs, disks, and cables.
- Perform GPU resets, node drains and reboots, link reseating, and BMC recovery.
- Collect logs, DCGM output, network diagnostics, and hardware health reports.
- Manage incident tickets through resolution or escalation.
- Install cables, swap hardware, rack and stack equipment, and label components.
- Perform structured shift handoffs with the APAC operations team.
- Maintain and update operational runbooks.
- Assist with hardware deployment, firmware updates, and inventory management.
- Tag and document novel incidents for automation.
- Provide structured handoff notes.
