Senior Software Engineer (DCIE)
Crusoe is an AI infrastructure and cloud computing company. It provides GPU cloud capacity, managed AI services, inference, fine-tuning, data centers, and energy infrastructure for AI developers and enterprise customers.
Maintainer signals as of 8/23/2026
Funding history
Projects
About Crusoe, Inc
Crusoe designs, builds, and operates energy-first AI infrastructure, including data centers, GPU cloud computing, and modular AI factories. Crusoe Cloud provides GPU clusters, managed Kubernetes and Slurm, storage, networking, observability, managed inference, serverless fine-tuning, and model deployment through Crusoe Intelligence Foundry. Its customers include AI startups, enterprises, and organizations developing training, inference, analytics, and other compute-intensive workloads.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You develop software that manages GPU servers and data centers. You build diagnostics, observability, automation, repair, validation, and operational tooling for high-performance GPU clusters, deploy and monitor those tools, and support facilities management and liquid cooling systems.
Requirements
- 4-6 years of software engineering experience
- Distributed systems expertise
- Reliability expertise
- Cloud platform expertise
- Experience with Kubernetes, infrastructure as code, and GCP
- Proficiency in Go, Python, Java, or Rust
- Analytical skills
- Problem-solving skills
- Communication skills
- Collaboration skills
- Experience with Temporal and Kubernetes
- Experience working with hardware vendors
- GPU fleet operations or hyperscale data center experience
Responsibilities
- Develop deep-level diagnostics and troubleshoot hardware faults within GPU racks and high-density compute systems
- Develop troubleshooting and automation tooling for NVIDIA and AMD GPU platforms
- Develop automation and AI agents for component-level diagnosis and hardware remediation
- Develop tooling and AI agents for managing critical data center environments
- Develop post-repair validation and testing tools using burn-in, PyTorch, and NVIDIA NCCL
- Deploy, monitor, and provide operational support for developed tooling
- Develop automation and operational tooling for facility power management and direct liquid cooling systems
Benefits
- Restricted Stock Units
- Health insurance including HDHP and PPO options
- Vision insurance
- Dental insurance
- Employer contributions to HSA accounts
- Paid parental leave
- Paid life insurance
- Short-term disability insurance
- Long-term disability insurance
- Teladoc
- 401(k) with a 100% match up to 4% of salary
- Paid time off
- Paid holidays
- Cell phone reimbursement
- Tuition reimbursement
- Calm app subscription
- MetLife Legal
- Company-paid commuter benefit
