Staff Cloud Support Engineer

Crusoe is an AI infrastructure and cloud computing company. It provides GPU cloud capacity, managed AI services, inference, fine-tuning, data centers, and energy infrastructure for AI developers and enterprise customers.

Maintainer signals as of 8/23/2026

Distributed
About Crusoe, Inc

Crusoe designs, builds, and operates energy-first AI infrastructure, including data centers, GPU cloud computing, and modular AI factories. Crusoe Cloud provides GPU clusters, managed Kubernetes and Slurm, storage, networking, observability, managed inference, serverless fine-tuning, and model deployment through Crusoe Intelligence Foundry. Its customers include AI startups, enterprises, and organizations developing training, inference, analytics, and other compute-intensive workloads.

View jobs by Crusoe, Inc

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

Serve as the highest-level escalation point for complex incidents, lead cross-functional root-cause investigations, design systemic reliability improvements, influence Kubernetes and workload orchestration architecture, troubleshoot AI and machine-learning infrastructure, advise customers during high-risk incidents, deliver executive-ready analyses, mentor engineers, and define support standards.

Requirements

  • 8+ years of experience in SRE, DevOps, HPC, or Cloud Infrastructure roles
  • Advanced Linux systems expertise
  • Deep Kubernetes operational experience at CKA level or higher
  • Strong knowledge of InfiniBand, RDMA, RoCE, and SDN
  • Experience supporting AI/ML workloads at scale on GPU clusters
  • Track record of resolving multi-layer distributed system failures
  • Strong customer communication and executive-facing presence

Responsibilities

  • Serve as the highest-level escalation point for complex P1/P0 incidents
  • Lead cross-functional root-cause investigations
  • Design systemic fixes with SRE and software teams
  • Improve node validation, burn-in, performance baselining, and release readiness
  • Influence Kubernetes architecture and workload orchestration
  • Reduce MTTR and incident recurrence
  • Troubleshoot NCCL, InfiniBand, GPU driver, and firmware issues
  • Support AI training and inference workloads
  • Deliver executive-ready root-cause analyses
  • Mentor engineers and define technical standards

Benefits

  • Restricted Stock Units
  • Paid time off
  • Paid holidays
  • Comprehensive health insurance
  • Dental insurance
  • Vision insurance
  • Employer HSA contributions
  • Paid parental leave
  • Paid life insurance
  • Short-term disability insurance
  • Long-term disability insurance
  • Professional development
  • Tuition reimbursement
  • Mental health and wellness support
  • Commuter benefits
  • Cell phone stipend
  • 401(k) retirement plan with company match up to 4% of salary
  • Volunteer time off