Search...

Network Engineer

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will design, deploy, and operate the network infrastructure supporting GPU AI factories across Europe. You will manage physical cabling, high-speed Ethernet and InfiniBand fabrics, routing, overlays, transport tuning, automation pipelines, and production troubleshooting for large GPU clusters and live AI training workloads.

Requirements

  • 4+ years of datacenter networking experience
  • Experience with InfiniBand or 400G/800G Ethernet at scale
  • Deep familiarity with RDMA, RoCE v2, and GPU training cluster communication patterns
  • Solid knowledge of Linux networking internals including DSCP, ECN, PFC, and adaptive routing
  • Experience with Ansible, Terraform, Netbox, and CI/CD pipelines
  • Ability to interpret tcpdump, perftest, and ib_write_bw diagnostics and correlate them with application performance

Responsibilities

  • Design and deploy InfiniBand NDR 400G, HDR, and high-speed Ethernet fabrics for GPU clusters of 1,000+ nodes
  • Configure and operate Arista, Juniper, and Mellanox/NVIDIA equipment
  • Manage BGP, OSPF, and VXLAN overlays
  • Tune RoCE and InfiniBand transport for NCCL and UCX workloads
  • Maintain network automation pipelines across all sites
  • Troubleshoot performance regressions, packet loss, and congestion during live AI training runs
Network Engineer at Sesterce | JobStash