Senior AI Data Center Network Engineer

Bitdeer is a technology company providing Bitcoin mining solutions.

Singapore, SG
About Bitdeer

Bitdeer provides full-spectrum Bitcoin mining and high-performance computing solutions, including SEALMINER mining equipment, Minerbase cooling containers, cloud mining, co-mining, mining management applications, mining rights marketplaces, and large-scale data center operations. The company also offers AI cloud infrastructure with GPU computing, model training and deployment capabilities, and turnkey AI data center solutions for enterprise customers and developers. Bitdeer is headquartered in Singapore and operates globally.

View jobs by Bitdeer

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will architect, deploy, optimize, and operate high-performance networks for large GPU clusters. You will investigate complex network issues, automate configuration management, monitor fabric health, support container and hybrid-cloud networking, and lead network changes, incident response, mitigations, and root-cause analysis.

Requirements

  • Bachelor's degree or above in a related field
  • 8-10 years of large-scale data center network engineering, architecture, or operations experience
  • Expertise in TCP/IP, BGP, OSPF, ISIS, VXLAN EVPN, and Spine-Leaf networking
  • Experience with InfiniBand, Subnet Manager, RoCEv2, PFC, ECN, and congestion control
  • Experience configuring Cisco Nexus, NVIDIA/Mellanox, or Juniper network equipment
  • Network automation and scripting skills in Python, Ansible, and Terraform
  • Linux system-administration skills
  • Experience with Kubernetes CNI and cloud-native architectures
  • Professional fluency in English and Chinese

Responsibilities

  • Architect scalable, highly available AI data center network solutions
  • Design, deploy, and optimize GPU-cluster InfiniBand and RoCEv2 fabrics
  • Configure and manage NVIDIA Spectrum and Quantum switches
  • Investigate network issues involving RDMA, packet loss, congestion, latency, and NCCL timeouts
  • Develop network automation using Python, Ansible, and Terraform
  • Monitor network telemetry and fabric health using UFM, NetQ, Zabbix, and Prometheus
  • Support Kubernetes, Docker, and hybrid-cloud network integration
  • Lead network changes, capacity expansions, firmware upgrades, incident response, and root-cause analysis