Search...

Staff Hardware Systems Engineer Performance

Crusoe logo
Crusoe

Crusoe is an AI infrastructure company that designs, builds, and operates AI data centers and a cloud platform. It provides managed AI services, GPU compute, model fine-tuning and inference, and infrastructure operations for organizations building and deploying AI workloads.

Maintainer signals as of 8/14/2026

Distributed
About Crusoe

Crusoe, the AI factory company, provides Crusoe Cloud and Crusoe Intelligence Foundry for AI development and production. Its offerings include managed inference, serverless fine-tuning, high-performance NVIDIA and AMD compute, accelerated storage, RDMA networking, managed Kubernetes and Slurm, and operations tooling. The company also designs, builds, and operates modular AI data-center infrastructure using an energy-first approach, serving customers that need scalable training, inference, and AI platform infrastructure.

View jobs by Crusoe

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will drive the lifecycle of next-generation compute platforms, define performance characterization strategies, study training and inference workloads, tune cluster configurations, analyze system bottlenecks, debug hardware and software issues, and influence platform architecture and infrastructure strategy.

Requirements

  • 8+ years of experience in hardware systems, platform, performance, ML systems, infrastructure engineering, or related areas
  • Experience with large-scale GPU or accelerated computing infrastructure for AI/ML or HPC workloads
  • Experience with distributed training or inference workloads at scale
  • Experience with workload benchmarking, performance profiling, and system optimization
  • Understanding of CPU, GPU, memory, storage, networking, PCIe, InfiniBand, and NVLink
  • Experience with system bring-up, validation, performance characterization, and root-cause analysis
  • Experience developing automation, testing, diagnostics, or data-analysis frameworks using Python, Shell, or similar languages
  • Bachelor’s or Master’s degree in Electrical Engineering, Computer Engineering, Computer Science, or equivalent experience

Responsibilities

  • Drive the lifecycle of next-generation compute platforms
  • Define and execute performance characterization and validation strategies
  • Conduct workload characterization studies across training and inference models
  • Translate workload and platform insights into cluster tuning recommendations
  • Build workload performance profiles and reference configurations
  • Analyze system and workload performance and identify bottlenecks
  • Lead system-level debugging across compute, memory, storage, networking, accelerators, and firmware
  • Partner on prototyping, qualification, NPI, and production readiness
  • Collaborate across hardware, firmware, networking, software, infrastructure, reliability, and operations
  • Influence platform architecture, technology selection, and hardware roadmaps

Benefits

  • Restricted Stock Units
  • Health insurance package options including HDHP and PPO
  • Vision insurance
  • Dental insurance
  • Employer HSA contributions
  • Paid parental leave
  • Paid life insurance
  • Short-term disability insurance
  • Long-term disability insurance
  • Teladoc
  • 401(k) with a 100% match up to 4% of salary
  • Paid time off
  • Paid holidays
  • Cell phone reimbursement
  • Tuition reimbursement
  • Calm app subscription
  • MetLife Legal
  • Company-paid commuter benefit of $300 per month
Staff Hardware Systems Engineer Performance at Crusoe | JobStash