Senior Staff Data Center Operations Engineer, GPU Hardware Architecture

Crusoe is an AI infrastructure and cloud computing company. It provides GPU cloud capacity, managed AI services, inference, fine-tuning, data centers, and energy infrastructure for AI developers and enterprise customers.

Maintainer signals as of 8/23/2026

Distributed
About Crusoe, Inc

Crusoe designs, builds, and operates energy-first AI infrastructure, including data centers, GPU cloud computing, and modular AI factories. Crusoe Cloud provides GPU clusters, managed Kubernetes and Slurm, storage, networking, observability, managed inference, serverless fine-tuning, and model deployment through Crusoe Intelligence Foundry. Its customers include AI startups, enterprises, and organizations developing training, inference, analytics, and other compute-intensive workloads.

View jobs by Crusoe, Inc

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

Serve as the technical authority for GPU platforms, translate GPU power and thermal roadmaps into facility requirements, analyze fleet telemetry for predictive maintenance, architect spare strategies, create repair SOPs and diagnostic tools, lead complex hardware root-cause analyses, and guide vendor technical relationships.

Requirements

  • 10+ years in hardware engineering, systems architecture, or data center infrastructure
  • Experience educating and influencing engineering and operations teams
  • Experience managing or architecting GPU clusters at scale
  • Expert knowledge of NVIDIA Hopper, Blackwell, and Rubin architectures
  • Expert knowledge of AMD Instinct architectures
  • Mastery of NVLink, NVSwitch, and InfiniBand
  • Ability to translate GPU data sheets into mechanical engineering requirements
  • Proficiency in Python, Go, or Bash
  • Experience with DCGM and ROCm
  • Experience using large datasets or machine learning frameworks
  • Experience using failure telemetry for sparing and field-service workflows
  • Deep understanding of Direct-to-Chip cooling, fluid dynamics, pressure-drop curves, and dripless couplings
  • B.S. or M.S. in electrical engineering, computer engineering, or a related technical field

Responsibilities

  • Provide technical guidance on upcoming GPU platforms and facility designs
  • Translate GPU power and thermal roadmaps into power, cooling, and rack-spacing requirements
  • Analyze fleet telemetry for predictive maintenance
  • Identify pre-failure patterns in HBM and NVLink components
  • Architect site-level spare strategies and define critical spare inventories
  • Create SOPs for GPU repairs and manifold maintenance
  • Develop diagnostic tools for NVLink, PCIe, and thermal issues
  • Act as the Tier-3 escalation point for complex hardware failures
  • Lead root-cause analyses across hardware and facility systems
  • Maintain a 24-month NVIDIA and AMD architecture roadmap
  • Educate stakeholders on HBM, interconnect, and liquid-cooling impacts
  • Support OEM and VAR technical relationships

Benefits

  • Restricted stock units
  • Paid time off
  • Paid holidays
  • Health insurance
  • Dental insurance
  • Vision insurance
  • Employer HSA contributions
  • Paid parental leave
  • Paid life insurance
  • Short-term disability insurance
  • Long-term disability insurance
  • Professional development
  • Tuition reimbursement
  • Mental health support
  • Wellness support
  • Commuter benefits
  • Cell phone stipend
  • 401(k) retirement plan
  • 401(k) company match up to 4% of salary
  • Volunteer time off