Senior AI Data Center Network Engineer

Bitdeer is a technology company providing Bitcoin mining solutions.

Singapore, SG
About Bitdeer

Bitdeer provides full-spectrum Bitcoin mining and high-performance computing solutions, including SEALMINER mining equipment, Minerbase cooling containers, cloud mining, co-mining, mining management applications, mining rights marketplaces, and large-scale data center operations. The company also offers AI cloud infrastructure with GPU computing, model training and deployment capabilities, and turnkey AI data center solutions for enterprise customers and developers. Bitdeer is headquartered in Singapore and operates globally.

View jobs by Bitdeer

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will architect, deploy, and operate scalable network solutions for AI data centers and GPU clusters. You will optimize InfiniBand and RoCEv2 fabrics, investigate network incidents, automate infrastructure configuration, monitor fabric health, support container networking, and lead changes, upgrades, capacity expansions, and root cause analysis.

Requirements

  • Bachelor's degree in a related discipline
  • 8-10 years of large-scale data center network operations, architecture, and engineering experience
  • Expertise in TCP/IP, BGP, OSPF, ISIS, VXLAN EVPN, and spine-leaf technologies
  • Experience with InfiniBand, Subnet Manager, RoCEv2, PFC, ECN, and congestion control
  • Experience configuring Cisco Nexus, NVIDIA/Mellanox, or Juniper network equipment
  • Strong Python, Ansible, Terraform, and Linux administration skills
  • Experience with Kubernetes CNI and cloud-native architectures
  • Professional fluency in English and Chinese

Responsibilities

  • Architect scalable, highly available AI data center networks
  • Design and operate spine-leaf, DCN, DCI, backbone, VXLAN EVPN, and SDN networks
  • Design, deploy, and optimize GPU cluster InfiniBand and RoCEv2 fabrics
  • Configure and manage NVIDIA Spectrum and Quantum switches
  • Investigate RDMA packet loss, congestion control, latency, and NCCL timeout issues
  • Develop network automation using Python, Ansible, and Terraform
  • Implement infrastructure as code and configuration management
  • Monitor network telemetry and fabric health using NVIDIA UFM, NetQ, Zabbix, and Prometheus
  • Support Kubernetes, Docker, and hybrid cloud network integration
  • Lead network changes, capacity expansions, firmware upgrades, incident response, mitigations, and root cause analysis