Senior Hardware Systems Engineer, Performance

Crusoe is an AI infrastructure and cloud computing company. It provides GPU compute, AI model training and inference, developer tools, managed orchestration, data centers, and energy infrastructure for AI builders and enterprise customers.

Maintainer signals as of 8/23/2026

Distributed
About Crusoe, Inc

Crusoe designs, builds, and operates energy-first AI infrastructure, including high-performance data centers, modular AI factories, and Crusoe Cloud. Its cloud platform provides NVIDIA and AMD GPU compute, scalable storage, high-performance networking, managed Kubernetes and Slurm, observability, model training, fine-tuning, and inference services. Crusoe serves AI developers, startups, enterprises, and teams running production-scale AI workloads.

View jobs by Crusoe, Inc

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

Evaluate and bring next-generation compute platforms into production while defining performance characterization strategies. Analyze AI workloads, tune cluster configurations, profile system behavior, debug hardware and software issues, and guide platform architecture, technology selection, and infrastructure strategy.

Requirements

  • 5-6+ years of experience in hardware systems engineering, platform engineering, performance engineering, machine learning systems engineering, infrastructure engineering, or related areas
  • Experience with large-scale GPU or accelerated computing infrastructure
  • Experience with distributed training or inference workloads and parallelism strategies
  • Experience with workload benchmarking, performance profiling, and system optimization
  • Understanding of CPU, GPU, memory, storage, networking, PCIe, InfiniBand, and NVLink
  • Experience with system bring-up, validation, performance characterization, and root-cause analysis
  • Experience developing automation, testing, diagnostics, or data-analysis frameworks
  • Bachelor’s or Master’s degree in Electrical Engineering, Computer Engineering, Computer Science, or equivalent experience

Responsibilities

  • Drive the lifecycle of next-generation compute platforms
  • Define and execute performance characterization and validation strategies
  • Characterize training and inference workloads
  • Translate workload insights into cluster tuning and configuration recommendations
  • Build workload performance profiles and reference configurations
  • Analyze system and workload performance and identify bottlenecks
  • Debug compute, memory, storage, networking, accelerator, and firmware issues
  • Partner on prototyping, qualification, NPI, and production readiness
  • Influence platform architecture, technology selection, and hardware roadmaps

Benefits

  • Industry competitive pay
  • Restricted Stock Units
  • Health insurance package options including HDHP and PPO
  • Vision insurance
  • Dental insurance
  • Employer contributions to HSA accounts
  • Paid parental leave
  • Paid life insurance
  • Short-term disability insurance
  • Long-term disability insurance
  • Teladoc
  • 401(k) with 100% match up to 4% of salary
  • Paid time off
  • Paid holidays
  • Cell phone reimbursement
  • Tuition reimbursement
  • Subscription to the Calm app
  • MetLife Legal
  • Company-paid commuter benefit of $300 per month