Search...

Staff Network Production Engineer, Operations

Crusoe logo
Crusoe

Crusoe is an AI infrastructure company that designs, builds, and operates AI data centers and a cloud platform. It provides managed AI services, GPU compute, model fine-tuning and inference, and infrastructure operations for organizations building and deploying AI workloads.

Maintainer signals as of 8/20/2026

Distributed
About Crusoe

Crusoe, the AI factory company, provides Crusoe Cloud and Crusoe Intelligence Foundry for AI development and production. Its offerings include managed inference, serverless fine-tuning, high-performance NVIDIA and AMD compute, accelerated storage, RDMA networking, managed Kubernetes and Slurm, and operations tooling. The company also designs, builds, and operates modular AI data-center infrastructure using an energy-first approach, serving customers that need scalable training, inference, and AI platform infrastructure.

View jobs by Crusoe

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will own production reliability across global edge, backbone, data center, and GPU cluster networks. You will lead incident response, perform root cause analysis, improve observability, write Python automation, define reliability metrics, maintain operational documentation, and mentor engineers during high-severity network events.

Requirements

  • 8+ years of production network engineering experience focused on operations, incident response, and reliability
  • Experience with streaming telemetry, SNMP, NetFlow or sFlow, Grafana, Prometheus, and ThousandEyes
  • Experience operating RDMA/RoCE lossless fabrics for GPU or HPC workloads
  • Expert knowledge of BGP, EVPN-VXLAN, IS-IS, OSPF, MPLS, QoS, and TCP/IP
  • Proficiency with Arista EOS and Juniper Junos platforms
  • Python proficiency
  • Experience operating large device fleets across multi-region environments with on-call responsibility
  • Bachelor's degree in Computer Science, Electrical Engineering, or a related field, or equivalent practical experience

Responsibilities

  • Own uptime across global edge, backbone, data center, and GPU cluster networks
  • Lead and contribute to high-severity network incident response
  • Drive root cause analyses and track remediation plans to closure
  • Improve network monitoring using telemetry and observability tools
  • Author and maintain runbooks, escalation playbooks, and standard operating procedures
  • Write Python tooling for automated remediation and diagnostics
  • Define and track network reliability metrics and service level objectives
  • Mentor senior engineers

Benefits

  • Competitive compensation and equity packages
  • Restricted Stock Units
  • Paid time off, paid holidays, and leave of absence programs
  • Comprehensive health, dental, and vision insurance
  • Employer contributions to HSA account
  • Paid parental leave
  • Paid life insurance
  • Short-term and long-term disability insurance
  • Professional development and tuition reimbursement
  • Mental health and wellness support
  • Commuter benefits for parking and transit
  • Cell phone stipend
  • 401(k) retirement plan with company match up to 4% of salary
  • Volunteer time off
  • Global travel insurance and emergency assistance
  • Daily meals allowance
  • Additional location-specific perks and programs