Staff Software Engineer (Cloud Infrastructure)

Crusoe is an AI infrastructure and cloud computing company. It provides GPU cloud capacity, managed AI services, inference, fine-tuning, data centers, and energy infrastructure for AI developers and enterprise customers.

Maintainer signals as of 8/23/2026

Distributed
About Crusoe, Inc

Crusoe designs, builds, and operates energy-first AI infrastructure, including data centers, GPU cloud computing, and modular AI factories. Crusoe Cloud provides GPU clusters, managed Kubernetes and Slurm, storage, networking, observability, managed inference, serverless fine-tuning, and model deployment through Crusoe Intelligence Foundry. Its customers include AI startups, enterprises, and organizations developing training, inference, analytics, and other compute-intensive workloads.

View jobs by Crusoe, Inc

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

Diagnose, maintain, and repair high-performance GPU compute clusters; automate hardware fault diagnosis; develop software for NVIDIA and AMD GPU platforms; perform component repairs and validation testing; manage firmware upgrades; document maintenance; create troubleshooting procedures; and investigate systemic failures.

Requirements

  • Ability to code in Golang
  • Experience diagnosing and repairing high-density rack-mounted compute hardware in production environments
  • Deep understanding of GPU architectures and hands-on experience with GPU-based systems
  • Experience supporting NVIDIA A100, H200, GB200, B200, and AMD 350X and 355X platforms
  • Familiarity with InfiniBand, NVLink, and RoCE
  • Strong Linux experience with Ubuntu, Rocky Linux, and CentOS
  • Proficiency with NVIDIA DCGM and NVIDIA field diagnostic utilities
  • Experience with enterprise server hardware, power delivery, and cooling systems
  • Experience working with hardware vendors and escalations
  • Background in large-scale GPU fleet operations or hyperscale data center environments

Responsibilities

  • Automate diagnosis and troubleshooting of hardware faults within GPU racks
  • Develop software to troubleshoot and support GPU platforms
  • Execute component-level diagnosis and remediation for failed or degraded hardware
  • Perform field-replaceable unit repairs for GPUs, power supplies, cooling systems, interconnects, and networking hardware
  • Conduct post-repair validation, burn-in, Torch, and NVIDIA NCCL testing
  • Implement preventative maintenance procedures
  • Perform firmware and BIOS upgrades across the GPU fleet
  • Document maintenance activities, failures, and resolutions
  • Develop and update standard operating procedures for troubleshooting, repair, and validation
  • Identify root causes of systemic failures and implement preventative solutions
  • Participate in a rotating infrastructure on-call schedule

Benefits

  • Hybrid work schedule
  • Industry competitive pay
  • Restricted Stock Units
  • Health insurance options including HDHP and PPO
  • Vision insurance
  • Dental insurance
  • Employer contributions to HSA accounts
  • Paid parental leave
  • Paid life insurance
  • Short-term and long-term disability insurance
  • Teladoc
  • 401(k) with 100% match up to 4% of salary
  • Generous paid time off and holiday schedule
  • Cell phone reimbursement
  • Tuition reimbursement
  • Subscription to the Calm app
  • MetLife Legal
  • Company-paid commuter benefit of $300 per pay period