Senior AI Cloud Network Operations Engineer

Bitdeer is a technology company providing Bitcoin mining solutions.

0 current maintainers0 active leadsTeam intelligence

Maintainer signals as of 9/25/2026

Singapore, SG
About Bitdeer

Bitdeer provides full-spectrum Bitcoin mining and high-performance computing solutions, including SEALMINER mining equipment, Minerbase cooling containers, cloud mining, co-mining, mining management applications, mining rights marketplaces, and large-scale data center operations. The company also offers AI cloud infrastructure with GPU computing, model training and deployment capabilities, and turnkey AI data center solutions for enterprise customers and developers. Bitdeer is headquartered in Singapore and operates globally.

View jobs by Bitdeer

Skills

About the Role

Monitor the full network estate around the clock, respond to outages and device failures, manage network tickets, investigate AI network performance issues, produce health reports and post-incident reviews, support customer and R&D change requests, and execute controlled network changes.

Requirements

  • 5+ years of network operations experience
  • Degree in computer science, telecommunications, electronics, or a similar field
  • Knowledge of BGP, OSPF, VXLAN, EVPN, and ECMP
  • Experience with enterprise-grade switches and routers
  • Working knowledge of InfiniBand or RoCEv2
  • Understanding of PFC and ECN congestion control
  • Familiarity with Zabbix, Prometheus, Grafana, Cacti, and MTR
  • Fluency in Chinese and English
  • Experience with large-scale GPU clusters is preferred
  • Familiarity with NCCL and MPI is preferred
  • Proficiency with NVIDIA UFM is preferred
  • CCIE, JNCIE, NCP-AIN, or advanced networking certifications are preferred
  • Optical transmission or global backbone experience is preferred
  • Python or Go network automation experience is preferred

Responsibilities

  • Monitor switches, routers, optical transport, device health, bandwidth, and traffic load
  • Monitor AI-specific network signals
  • Respond to network outages, link failures, and device-down events
  • Investigate jitter, NCCL throughput degradation, and AI network performance issues
  • Manage network request tickets from creation through resolution
  • Produce network health reports and post-incident reviews
  • Maintain the network operations knowledge base
  • Handle network change requests for internal R&D teams and customers
  • Execute network changes under change management procedures

Benefits

  • Attractive welfare benefits
  • Training and mentoring
  • Personal accountability, autonomy, fast growth, and learning opportunities