Senior Software Engineer (DCIE)
Crusoe is an AI infrastructure company that designs, builds, and operates AI data centers and a cloud platform. It provides managed AI services, GPU compute, model fine-tuning and inference, and infrastructure operations for organizations building and deploying AI workloads.
Maintainer signals as of 8/12/2026
Funding history
Projects
About Crusoe
Crusoe, the AI factory company, provides Crusoe Cloud and Crusoe Intelligence Foundry for AI development and production. Its offerings include managed inference, serverless fine-tuning, high-performance NVIDIA and AMD compute, accelerated storage, RDMA networking, managed Kubernetes and Slurm, and operations tooling. The company also designs, builds, and operates modular AI data-center infrastructure using an energy-first approach, serving customers that need scalable training, inference, and AI platform infrastructure.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You develop software that manages GPU servers and data centers. You build diagnostics, observability, automation, repair, validation, and operational tooling for high-performance GPU clusters, deploy and monitor those tools, and support facilities management and liquid cooling systems.
Requirements
- 4-6 years of software engineering experience
- Distributed systems expertise
- Reliability expertise
- Cloud platform expertise
- Experience with Kubernetes, infrastructure as code, and GCP
- Proficiency in Go, Python, Java, or Rust
- Analytical skills
- Problem-solving skills
- Communication skills
- Collaboration skills
- Experience with Temporal and Kubernetes
- Experience working with hardware vendors
- GPU fleet operations or hyperscale data center experience
Responsibilities
- Develop deep-level diagnostics and troubleshoot hardware faults within GPU racks and high-density compute systems
- Develop troubleshooting and automation tooling for NVIDIA and AMD GPU platforms
- Develop automation and AI agents for component-level diagnosis and hardware remediation
- Develop tooling and AI agents for managing critical data center environments
- Develop post-repair validation and testing tools using burn-in, PyTorch, and NVIDIA NCCL
- Deploy, monitor, and provide operational support for developed tooling
- Develop automation and operational tooling for facility power management and direct liquid cooling systems
Benefits
- Restricted Stock Units
- Health insurance including HDHP and PPO options
- Vision insurance
- Dental insurance
- Employer contributions to HSA accounts
- Paid parental leave
- Paid life insurance
- Short-term disability insurance
- Long-term disability insurance
- Teladoc
- 401(k) with a 100% match up to 4% of salary
- Paid time off
- Paid holidays
- Cell phone reimbursement
- Tuition reimbursement
- Calm app subscription
- MetLife Legal
- Company-paid commuter benefit
