Senior Production Engineer Compute
Crusoe is an AI infrastructure and cloud computing company. It provides GPU cloud capacity, managed AI services, inference, fine-tuning, data centers, and energy infrastructure for AI developers and enterprise customers.
Maintainer signals as of 8/23/2026
Funding history
Projects
About Crusoe, Inc
Crusoe designs, builds, and operates energy-first AI infrastructure, including data centers, GPU cloud computing, and modular AI factories. Crusoe Cloud provides GPU clusters, managed Kubernetes and Slurm, storage, networking, observability, managed inference, serverless fine-tuning, and model deployment through Crusoe Intelligence Foundry. Its customers include AI startups, enterprises, and organizations developing training, inference, analytics, and other compute-intensive workloads.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
Support virtualization, hypervisor, and kernel-level performance across compute infrastructure by deploying and optimizing bare-metal and virtualized platforms, building automation and observability tools, troubleshooting kernel and hardware issues, tuning system subsystems, improving guest VM reliability, and validating emerging compute hardware for AI and HPC workloads.
Requirements
- 5+ years of professional experience in Compute SRE, Linux systems engineering, or compute infrastructure
- Proficiency in Linux kernel internals
- Experience with KVM, Xen, QEMU, VMware, or similar virtualization technologies
- Familiarity with SmartNICs, DPUs, and kernel bypass techniques
- Expert-level skills in Go, C, or Rust
- Experience with kdump, kexec, and kernel panic analysis
- Proficiency in infrastructure-as-code and CI/CD practices
- Understanding of compute scheduling, resource management, and high-throughput networking
- Experience with custom Linux distributions or kernels is a plus
- Exposure to AI model infrastructure and GPU clusters is a plus
Responsibilities
- Develop automation and observability tools for compute infrastructure
- Support and scale the virtualization stack
- Identify and resolve performance bottlenecks and driver issues
- Optimize hardware offloads
- Optimize CPU, GPU, and DPU/NIC performance for AI and HPC workloads
- Perform root cause analysis for kernel crashes and integration problems
- Integrate hypervisor enhancements
- Tune process scheduling, NUMA configuration, memory management, and interrupt handling
- Implement and validate support for SmartNICs, BlueField devices, and TPUs
Benefits
- Industry competitive pay
- Restricted Stock Units
- Health insurance
- Vision insurance
- Dental insurance
- Employer HSA contributions
- Paid parental leave
- Paid life insurance
- Short-term disability insurance
- Long-term disability insurance
- Teladoc
- 401(k) with 100% match up to 4% of salary
- Paid time off
- Paid holidays
- Cell phone reimbursement
- Tuition reimbursement
- Calm app subscription
- MetLife Legal
- Company-paid commuter benefit of $300 per month
