Senior Cloud Support Engineer - Weekend Shift

Crusoe is an AI infrastructure and cloud computing company. It provides GPU cloud capacity, managed AI services, inference, fine-tuning, data centers, and energy infrastructure for AI developers and enterprise customers.

Maintainer signals as of 8/23/2026

Distributed
About Crusoe, Inc

Crusoe designs, builds, and operates energy-first AI infrastructure, including data centers, GPU cloud computing, and modular AI factories. Crusoe Cloud provides GPU clusters, managed Kubernetes and Slurm, storage, networking, observability, managed inference, serverless fine-tuning, and model deployment through Crusoe Intelligence Foundry. Its customers include AI startups, enterprises, and organizations developing training, inference, analytics, and other compute-intensive workloads.

View jobs by Crusoe, Inc

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

Provide technical customer support through Zendesk, participate in on-call rotations, manage incidents, troubleshoot virtual machines and hardware failures, triage alerts, prepare maintenance windows, collaborate on root-cause analysis, and create training and knowledge-base materials.

Requirements

  • Bachelor's degree in IT, Computer Science, Engineering, or a related field, or 4+ years of equivalent technical experience
  • Strong Linux command-line interface skills
  • Proficiency with Git
  • 5+ years of experience in customer support, ideally in cloud, storage, or networking environments
  • Experience with Kubernetes, Slurm, Terraform, and Grafana
  • Familiarity with AWS, Azure, or GCP
  • Excellent communication and customer service skills
  • Understanding of Infiniband, RDMA, RoCE, and Software Defined Networking
  • Experience with automation tools and scripting languages

Responsibilities

  • Provide technical customer support through Zendesk
  • Meet service-level agreements and maintain high customer satisfaction
  • Participate in a 24/7 on-call rotation
  • Manage incident triage, communication, and response
  • Diagnose and resolve virtual machine, hardware, and scaling-test issues
  • Manage alert triage
  • Prepare for maintenance windows
  • Conduct node delivery testing
  • Collaborate with SRE, Networking, and Storage teams on root-cause analysis
  • Follow global ticketing and on-call handoff processes
  • Develop onboarding materials, knowledge-base documentation, and standard operating procedures

Benefits

  • Competitive compensation and equity packages
  • Restricted Stock Units
  • Paid time off
  • Paid holidays and leave of absence programs
  • Comprehensive health, dental and vision insurance
  • Employer HSA contributions
  • Paid parental leave
  • Paid life insurance
  • Short-term and long-term disability insurance
  • Professional development and tuition reimbursement
  • Mental health and wellness support
  • Commuter benefits
  • Cell phone stipend
  • 401(k) retirement plan with company match up to 4% of salary
  • Volunteer time off
  • Global travel insurance and emergency assistance
  • Daily meals allowance
  • Location-specific perks and programs