Staff Network Production Engineer, Network Ops
Crusoe is an AI infrastructure and cloud computing company. It provides GPU cloud capacity, managed AI services, inference, fine-tuning, data centers, and energy infrastructure for AI developers and enterprise customers.
Maintainer signals as of 8/23/2026
Funding history
Projects
About Crusoe, Inc
Crusoe designs, builds, and operates energy-first AI infrastructure, including data centers, GPU cloud computing, and modular AI factories. Crusoe Cloud provides GPU clusters, managed Kubernetes and Slurm, storage, networking, observability, managed inference, serverless fine-tuning, and model deployment through Crusoe Intelligence Foundry. Its customers include AI startups, enterprises, and organizations developing training, inference, analytics, and other compute-intensive workloads.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
The Staff Network Production Engineer, Network Ops will implement and operate Crusoe’s global edge, backbone, and data center network for GPU-based HPC clusters. The role provides continuous monitoring and management, rapid incident response, and high availability of network services.
Requirements
- 10+ years of related experience operating at scale in a production environment.
- Strong knowledge of TCP/IP, QoS, BGP, OSPF/IS-IS, EVPN, VXLAN, and MPLS-related technologies including RSVP-TE and LDP.
- Strong understanding of SNMP, IPFIX, sFlow/netflow, and telemetry.
- Familiarity with data center, backbone, and edge network architecture.
- Experience with Python or similar languages for automation and tooling.
- Hands-on experience with Mellanox, Cisco, Arista, Juniper, and other network devices.
- Familiarity with Broadcom and Barefoot switch and router chipsets.
- Knowledge of public cloud connectivity options.
- Understanding of IPv6 and IPv4-IPv6 coexistence technologies.
- Experience with network observability, flow analytics, and dashboarding tools.
- Bachelor’s degree in a relevant field or equivalent experience.
Responsibilities
- Monitor network performance and perform advanced troubleshooting and root-cause analysis.
- Guide post-mortem reviews and improvements.
- Execute network changes across data center, backbone, and edge infrastructure.
- Manage and optimize the global network and public cloud connectivity.
- Collaborate with Network Engineering and cross-functional teams.
- Develop monitoring, alerting, and network availability systems.
- Mentor network engineers and establish incident response and operational readiness practices.
- Manage network vendors and contracts.
- Deliver network health metrics, event statistics, and performance analysis.
- Participate in 24/7 network on-call support.
Benefits
- Pension contributions
- Private health insurance
- Dental insurance
- Income protection
- Life assurance
