Senior Cloud Support Engineer - Weekend Shift

Crusoe is an AI infrastructure and cloud computing company. It provides GPU cloud capacity, managed AI services, inference, fine-tuning, data centers, and energy infrastructure for AI developers and enterprise customers.

Maintainer signals as of 8/23/2026

Distributed
About Crusoe, Inc

Crusoe designs, builds, and operates energy-first AI infrastructure, including data centers, GPU cloud computing, and modular AI factories. Crusoe Cloud provides GPU clusters, managed Kubernetes and Slurm, storage, networking, observability, managed inference, serverless fine-tuning, and model deployment through Crusoe Intelligence Foundry. Its customers include AI startups, enterprises, and organizations developing training, inference, analytics, and other compute-intensive workloads.

View jobs by Crusoe, Inc

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

Provide technical customer support through Zendesk, meet service-level agreements, and maintain high customer satisfaction. Participate in a 24/7 on-call rotation, lead incident triage and communication, troubleshoot virtual machines and hardware failures, manage alerts and maintenance testing, collaborate on root-cause investigations, and create training materials, documentation, and standard operating procedures.

Requirements

  • Bachelor's degree in IT, Computer Science, Engineering, or a related field, or 4+ years of equivalent technical experience
  • Strong command-line interface skills in Linux environments
  • Proficiency with Git
  • 5+ years of customer support experience, ideally in cloud, storage, or networking environments
  • Experience with Kubernetes, Slurm, Terraform, and Grafana
  • Familiarity with AWS, Azure, or GCP
  • Excellent communication and customer service skills
  • Understanding of InfiniBand, RDMA, RoCE, and Software Defined Networking

Responsibilities

  • Provide technical support through Zendesk
  • Participate in a 24/7 on-call rotation
  • Lead incident triage and response communication
  • Diagnose and resolve virtual machine, hardware, and scaling issues
  • Manage alert triage and maintenance preparation
  • Conduct node delivery testing
  • Collaborate with SRE, Networking, and Storage teams on root-cause analysis
  • Follow ticketing and on-call handoff processes
  • Develop onboarding materials, knowledge base documentation, and standard operating procedures

Benefits

  • Restricted Stock Units
  • Paid time off
  • Paid holidays
  • Leave of absence programs
  • Comprehensive health insurance
  • Dental insurance
  • Vision insurance
  • Employer HSA contributions
  • Paid parental leave
  • Paid life insurance
  • Short-term disability insurance
  • Long-term disability insurance
  • Professional development
  • Tuition reimbursement
  • Mental health and wellness support
  • Commuter benefits
  • Cell phone stipend
  • 401(k) retirement plan with company match up to 4% of salary
  • Volunteer time off
  • Global travel insurance
  • Emergency assistance
  • Daily meals allowance