Search...

Staff Cloud Support Engineer

Crusoe logo
Crusoe

Crusoe is an AI infrastructure company that designs, builds, and operates AI data centers and a cloud platform. It provides managed AI services, GPU compute, model fine-tuning and inference, and infrastructure operations for organizations building and deploying AI workloads.

Maintainer signals as of 8/23/2026

Distributed
About Crusoe

Crusoe, the AI factory company, provides Crusoe Cloud and Crusoe Intelligence Foundry for AI development and production. Its offerings include managed inference, serverless fine-tuning, high-performance NVIDIA and AMD compute, accelerated storage, RDMA networking, managed Kubernetes and Slurm, and operations tooling. The company also designs, builds, and operates modular AI data-center infrastructure using an energy-first approach, serving customers that need scalable training, inference, and AI platform infrastructure.

View jobs by Crusoe

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You serve as the highest-level escalation point for complex incidents, lead cross-functional root-cause investigations, and design systemic reliability improvements. You influence Kubernetes and workload orchestration architecture, troubleshoot AI and machine-learning infrastructure, advise customers during high-risk incidents, deliver executive-ready root-cause analyses, mentor engineers, and define support standards.

Requirements

  • 8+ years of experience in SRE, DevOps, HPC, or Cloud Infrastructure roles
  • Advanced Linux systems expertise
  • Deep Kubernetes operational experience at CKA level or higher
  • Strong knowledge of Infiniband, RDMA, RoCE, and SDN
  • Experience supporting AI/ML workloads at scale on GPU clusters
  • Track record of resolving multi-layer distributed system failures
  • Strong customer communication and executive-facing presence

Responsibilities

  • Serve as the highest-level escalation point for complex P1/P0 incidents
  • Lead cross-functional root-cause investigations
  • Design systemic fixes with SRE and software teams
  • Improve node validation, burn-in, performance baselining, and release readiness
  • Influence Kubernetes architecture and workload orchestration
  • Reduce MTTR and incident recurrence
  • Troubleshoot NCCL, Infiniband, GPU driver, and firmware issues
  • Support AI training and inference workloads
  • Deliver executive-ready root-cause analyses
  • Mentor engineers and define technical standards

Benefits

  • Restricted Stock Units
  • Paid time off
  • Paid holidays
  • Comprehensive health insurance
  • Dental insurance
  • Vision insurance
  • Employer HSA contributions
  • Paid parental leave
  • Paid life insurance
  • Short-term disability insurance
  • Long-term disability insurance
  • Professional development
  • Tuition reimbursement
  • Mental health and wellness support
  • Commuter benefits
  • Cell phone stipend
  • 401(k) retirement plan with company match up to 4% of salary
  • Volunteer time off