Staff Production Engineer, Compute
Crusoe is an AI infrastructure company that designs, builds, and operates AI data centers and a cloud platform. It provides managed AI services, GPU compute, model fine-tuning and inference, and infrastructure operations for organizations building and deploying AI workloads.
Maintainer signals as of 8/14/2026
Funding history
Projects
About Crusoe
Crusoe, the AI factory company, provides Crusoe Cloud and Crusoe Intelligence Foundry for AI development and production. Its offerings include managed inference, serverless fine-tuning, high-performance NVIDIA and AMD compute, accelerated storage, RDMA networking, managed Kubernetes and Slurm, and operations tooling. The company also designs, builds, and operates modular AI data-center infrastructure using an energy-first approach, serving customers that need scalable training, inference, and AI platform infrastructure.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will support virtualization, hypervisors, and kernel-level performance across compute infrastructure. You will build automation and observability tools, scale KVM and QEMU platforms, debug kernel and hardware issues, optimize AI and HPC workloads, tune kernel subsystems, and validate emerging compute hardware.
Requirements
- 8+ years of experience in Compute SRE, Linux system engineering, or compute infrastructure roles
- Strong proficiency in Linux kernel internals
- Experience with KVM, Xen, QEMU, or VMware
- Familiarity with SmartNICs, DPUs, and kernel bypass techniques
- Expert-level skill in Go, C, or Rust
- Experience with kdump, kexec, and kernel panic analysis
- Proficiency with Infrastructure as Code tooling and CI/CD practices
- Strong understanding of compute scheduling, resource management, and high-throughput networking
Responsibilities
- Develop automation and observability tools for compute infrastructure
- Support and scale the virtualization stack
- Identify and resolve performance bottlenecks and driver issues
- Optimize hardware offloads and AI and HPC workloads
- Perform root cause analysis for kernel crashes and integration problems
- Integrate hypervisor enhancements for guest VM reliability and workload isolation
- Tune process scheduling, NUMA configuration, memory management, and interrupt handling
- Implement and validate support for SmartNICs, BlueField devices, and TPUs
Benefits
- Competitive compensation
- Restricted Stock Units
- Paid time off and paid holidays
- Comprehensive health, dental, and vision insurance
- Employer contributions to HSA account
- Paid parental leave
- Paid life insurance
- Short-term and long-term disability insurance
- Professional development and tuition reimbursement
- Mental health and wellness support
- Commuter benefits for parking and transit
- Cell phone stipend
- 401(k) retirement plan with company match up to 4% of salary
- Volunteer time off
