Staff Cloud Support Engineer
Crusoe is an AI infrastructure and cloud computing company. It provides GPU cloud capacity, managed AI services, inference, fine-tuning, data centers, and energy infrastructure for AI developers and enterprise customers.
Maintainer signals as of 8/23/2026
Funding history
Projects
About Crusoe, Inc
Crusoe designs, builds, and operates energy-first AI infrastructure, including data centers, GPU cloud computing, and modular AI factories. Crusoe Cloud provides GPU clusters, managed Kubernetes and Slurm, storage, networking, observability, managed inference, serverless fine-tuning, and model deployment through Crusoe Intelligence Foundry. Its customers include AI startups, enterprises, and organizations developing training, inference, analytics, and other compute-intensive workloads.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
Serve as the highest-level escalation point for complex incidents, lead cross-functional root-cause investigations, design systemic reliability improvements, influence Kubernetes and workload orchestration architecture, troubleshoot AI and machine-learning infrastructure, advise customers during high-risk incidents, deliver executive-ready analyses, mentor engineers, and define support standards.
Requirements
- 8+ years of experience in SRE, DevOps, HPC, or Cloud Infrastructure roles
- Advanced Linux systems expertise
- Deep Kubernetes operational experience at CKA level or higher
- Strong knowledge of InfiniBand, RDMA, RoCE, and SDN
- Experience supporting AI/ML workloads at scale on GPU clusters
- Track record of resolving multi-layer distributed system failures
- Strong customer communication and executive-facing presence
Responsibilities
- Serve as the highest-level escalation point for complex P1/P0 incidents
- Lead cross-functional root-cause investigations
- Design systemic fixes with SRE and software teams
- Improve node validation, burn-in, performance baselining, and release readiness
- Influence Kubernetes architecture and workload orchestration
- Reduce MTTR and incident recurrence
- Troubleshoot NCCL, InfiniBand, GPU driver, and firmware issues
- Support AI training and inference workloads
- Deliver executive-ready root-cause analyses
- Mentor engineers and define technical standards
Benefits
- Restricted Stock Units
- Paid time off
- Paid holidays
- Comprehensive health insurance
- Dental insurance
- Vision insurance
- Employer HSA contributions
- Paid parental leave
- Paid life insurance
- Short-term disability insurance
- Long-term disability insurance
- Professional development
- Tuition reimbursement
- Mental health and wellness support
- Commuter benefits
- Cell phone stipend
- 401(k) retirement plan with company match up to 4% of salary
- Volunteer time off
