Principal Systems Software Engineer
Crusoe is an AI infrastructure company that designs, builds, and operates AI data centers and a cloud platform. It provides managed AI services, GPU compute, model fine-tuning and inference, and infrastructure operations for organizations building and deploying AI workloads.
Projects
About Crusoe
Crusoe, the AI factory company, provides Crusoe Cloud and Crusoe Intelligence Foundry for AI development and production. Its offerings include managed inference, serverless fine-tuning, high-performance NVIDIA and AMD compute, accelerated storage, RDMA networking, managed Kubernetes and Slurm, and operations tooling. The company also designs, builds, and operates modular AI data-center infrastructure using an energy-first approach, serving customers that need scalable training, inference, and AI platform infrastructure.
Skills
About the Role
You will architect next-generation AI infrastructure spanning bare-metal, IaaS, and container platforms. You will lead I/O-path and cloud-fabric design, prototype and productionize systems for memory, networking, and compute, define technical strategy through RFCs and white papers, debug kernel-level issues, and represent the organization in open-source and industry forums.
Requirements
- 12+ years of experience designing and shipping core infrastructure at a major hyperscaler or specialized HPC cloud
- Authoritative knowledge of the Linux kernel, virtualization internals, and high-performance networking
- Ability to design software that maximizes NVIDIA or AMD GPU and high-speed NIC performance
- Experience leading cross-functional teams through high-ambiguity projects and delivering production-ready mission-critical systems
- Significant field contributions, such as patents, major open-source contributions, or distributed-systems research
- Bachelor’s or Master’s degree in Computer Science, Computer Engineering, a related analytical field, or equivalent professional experience
Responsibilities
- Architect bare-metal systems that deliver GPU throughput through InfiniBand and RDMA fabrics
- Design thin virtualization layers using KVM or custom micro-VMs
- Build a container substrate using Kubernetes or Slurm for AI workloads across heterogeneous GPU nodes
- Lead the architectural design of the internal cloud fabric for SR-IOV, RDMA, and virtualized GPU scheduling
- Lead R&D workstreams to prototype and productionize memory, networking, and compute systems
- Draft white papers and RFCs for the compute and networking stack
- Resolve complex I/O-path race conditions and optimize kernel-level memory pinning for GPU clusters
- Represent Crusoe in open-source communities and industry forums
Benefits
- Restricted Stock Units
- Paid time off and paid holidays
- Health, dental, and vision insurance
- Employer contributions to an HSA account
- Paid parental leave
- Paid life insurance
- Short-term and long-term disability insurance
- Professional development and tuition reimbursement
- Mental health and wellness support
- Commuter benefits
- Cell phone stipend
- 401(k) plan with company match up to 4% of salary
- Volunteer time off
