Staff Network Production Engineer, Operations

Crusoe is an AI infrastructure and cloud computing company. It provides GPU cloud capacity, managed AI services, inference, fine-tuning, data centers, and energy infrastructure for AI developers and enterprise customers.

Maintainer signals as of 8/23/2026

Distributed
About Crusoe, Inc

Crusoe designs, builds, and operates energy-first AI infrastructure, including data centers, GPU cloud computing, and modular AI factories. Crusoe Cloud provides GPU clusters, managed Kubernetes and Slurm, storage, networking, observability, managed inference, serverless fine-tuning, and model deployment through Crusoe Intelligence Foundry. Its customers include AI startups, enterprises, and organizations developing training, inference, analytics, and other compute-intensive workloads.

View jobs by Crusoe, Inc

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

Own production reliability across global edge, backbone, data center, and GPU cluster networks. Lead incident response, perform root cause analysis, improve observability, write Python automation, define reliability metrics, maintain operational documentation, and mentor engineers during high-severity network events.

Requirements

  • 8+ years of production network engineering experience focused on operations, incident response, and reliability
  • Experience with streaming telemetry, SNMP, NetFlow or sFlow, Grafana, Prometheus, and ThousandEyes
  • Experience operating RDMA/RoCE lossless fabrics for GPU or HPC workloads
  • Expert knowledge of BGP, EVPN-VXLAN, IS-IS, OSPF, MPLS, QoS, and TCP/IP
  • Proficiency with Arista EOS and Juniper Junos platforms
  • Python proficiency
  • Experience operating large device fleets across multi-region environments with on-call responsibility
  • Bachelor's degree in Computer Science, Electrical Engineering, or a related field, or equivalent practical experience

Responsibilities

  • Own uptime across global edge, backbone, data center, and GPU cluster networks
  • Lead and contribute to high-severity network incident response
  • Drive root cause analyses and track remediation plans to closure
  • Improve network monitoring using telemetry and observability tools
  • Author and maintain runbooks, escalation playbooks, and standard operating procedures
  • Write Python tooling for automated remediation and diagnostics
  • Define and track network reliability metrics and service level objectives
  • Mentor senior engineers

Benefits

  • Competitive compensation and equity packages
  • Restricted Stock Units
  • Paid time off, paid holidays, and leave of absence programs
  • Comprehensive health, dental, and vision insurance
  • Employer contributions to HSA account
  • Paid parental leave
  • Paid life insurance
  • Short-term and long-term disability insurance
  • Professional development and tuition reimbursement
  • Mental health and wellness support
  • Commuter benefits for parking and transit
  • Cell phone stipend
  • 401(k) retirement plan with company match up to 4% of salary
  • Volunteer time off
  • Global travel insurance and emergency assistance
  • Daily meals allowance
  • Additional location-specific perks and programs