Senior Production Engineer, Storage

Crusoe is an AI infrastructure and cloud computing company. It provides GPU cloud capacity, managed AI services, inference, fine-tuning, data centers, and energy infrastructure for AI developers and enterprise customers.

Maintainer signals as of 8/23/2026

Distributed
About Crusoe, Inc

Crusoe designs, builds, and operates energy-first AI infrastructure, including data centers, GPU cloud computing, and modular AI factories. Crusoe Cloud provides GPU clusters, managed Kubernetes and Slurm, storage, networking, observability, managed inference, serverless fine-tuning, and model deployment through Crusoe Intelligence Foundry. Its customers include AI startups, enterprises, and organizations developing training, inference, analytics, and other compute-intensive workloads.

View jobs by Crusoe, Inc

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will build automation and self-healing tools for distributed cloud storage infrastructure covering block, file, and object storage. You will improve replication, encryption, backup, restore, failover, availability, and performance; investigate incidents; diagnose low-level I/O issues; and optimize storage systems and architecture.

Requirements

  • Bachelor's degree in Computer Science, Electrical Engineering, or a related technical field, or equivalent practical experience
  • 5+ years of professional experience in Storage SRE, systems engineering, or storage engineering
  • Hands-on experience with enterprise storage platforms such as Pure Storage or EMC
  • Understanding of object, block, and file storage paradigms
  • Proficiency in Go, Python, Java, or C
  • Experience with Infrastructure as Code and deployment tooling
  • Knowledge of Linux I/O subsystems, memory management, and storage scheduling
  • Familiarity with NFS, SMB, iSCSI, or NVMe-oF
  • Experience with containerized workloads and Kubernetes or Docker
  • Incident response, troubleshooting, and documentation experience
  • Experience operating managed storage services at scale
  • Experience with distributed storage systems is a bonus
  • Experience with hybrid storage models is a bonus

Responsibilities

  • Build automation and self-healing tools
  • Monitor and maintain distributed cloud storage infrastructure
  • Improve data replication, encryption, backup, restore, and failover
  • Support user-facing storage services
  • Tune storage performance
  • Investigate and resolve storage incidents
  • Diagnose low-level I/O issues
  • Optimize I/O paths, cache policies, and file systems
  • Contribute to fault-tolerant storage backend architecture

Benefits

  • Industry competitive pay
  • Restricted Stock Units
  • Health insurance options including HDHP and PPO
  • Vision insurance
  • Dental insurance
  • Employer HSA contributions
  • Paid parental leave
  • Paid life insurance
  • Short-term and long-term disability insurance
  • Teladoc
  • 401(k) with 100% match up to 4% of salary
  • Paid time off
  • Paid holidays
  • Cell phone reimbursement
  • Tuition reimbursement
  • Calm app subscription
  • MetLife Legal
  • Company-paid commuter benefit of $300 per month