Search...

Senior Production Engineer Compute

Crusoe logo
Crusoe

Crusoe is an AI infrastructure company that designs, builds, and operates AI data centers and a cloud platform. It provides managed AI services, GPU compute, model fine-tuning and inference, and infrastructure operations for organizations building and deploying AI workloads.

Maintainer signals as of 8/14/2026

Distributed
About Crusoe

Crusoe, the AI factory company, provides Crusoe Cloud and Crusoe Intelligence Foundry for AI development and production. Its offerings include managed inference, serverless fine-tuning, high-performance NVIDIA and AMD compute, accelerated storage, RDMA networking, managed Kubernetes and Slurm, and operations tooling. The company also designs, builds, and operates modular AI data-center infrastructure using an energy-first approach, serving customers that need scalable training, inference, and AI platform infrastructure.

View jobs by Crusoe

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will support virtualization, hypervisor, and kernel-level performance across compute infrastructure. You will deploy and optimize bare-metal and virtualized platforms, build automation and observability tools, troubleshoot kernel and hardware issues, tune system subsystems, improve guest VM reliability, and validate emerging compute hardware for AI and HPC workloads.

Requirements

  • 5+ years of professional experience in Compute SRE, Linux systems engineering, or compute infrastructure
  • Proficiency in Linux kernel internals
  • Experience with KVM, Xen, QEMU, VMware, or similar virtualization technologies
  • Familiarity with SmartNICs, DPUs, and kernel bypass techniques
  • Expert-level skills in Go, C, or Rust
  • Experience with kdump, kexec, and kernel panic analysis
  • Proficiency in infrastructure-as-code and CI/CD practices
  • Understanding of compute scheduling, resource management, and high-throughput networking
  • Experience with custom Linux distributions or kernels is a plus
  • Exposure to AI model infrastructure and GPU clusters is a plus

Responsibilities

  • Develop automation and observability tools for compute infrastructure
  • Support and scale the virtualization stack
  • Identify and resolve performance bottlenecks and driver issues
  • Optimize hardware offloads
  • Optimize CPU, GPU, and DPU/NIC performance for AI and HPC workloads
  • Perform root cause analysis for kernel crashes and integration problems
  • Integrate hypervisor enhancements
  • Tune process scheduling, NUMA configuration, memory management, and interrupt handling
  • Implement and validate support for SmartNICs, BlueField devices, and TPUs

Benefits

  • Industry competitive pay
  • Restricted Stock Units
  • Health insurance
  • Vision insurance
  • Dental insurance
  • Employer HSA contributions
  • Paid parental leave
  • Paid life insurance
  • Short-term disability insurance
  • Long-term disability insurance
  • Teladoc
  • 401(k) with 100% match up to 4% of salary
  • Paid time off
  • Paid holidays
  • Cell phone reimbursement
  • Tuition reimbursement
  • Calm app subscription
  • MetLife Legal
  • Company-paid commuter benefit of $300 per month
Senior Production Engineer Compute at Crusoe | JobStash