Search...

Senior Production Engineer Storage

Crusoe logo
Crusoe

Crusoe is an AI infrastructure company that designs, builds, and operates AI data centers and a cloud platform. It provides managed AI services, GPU compute, model fine-tuning and inference, and infrastructure operations for organizations building and deploying AI workloads.

Series C13 current maintainers10 active leads5 new active leads9 lead step-downs1 early lead departureTeam intelligence

Maintainer signals as of 8/12/2026

Distributed
About Crusoe

Crusoe, the AI factory company, provides Crusoe Cloud and Crusoe Intelligence Foundry for AI development and production. Its offerings include managed inference, serverless fine-tuning, high-performance NVIDIA and AMD compute, accelerated storage, RDMA networking, managed Kubernetes and Slurm, and operations tooling. The company also designs, builds, and operates modular AI data-center infrastructure using an energy-first approach, serving customers that need scalable training, inference, and AI platform infrastructure.

View jobs by Crusoe

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will build automation and self-healing tools for distributed cloud storage infrastructure covering block, file, and object storage. You will improve replication, encryption, backup, restore, failover, availability, and performance; investigate incidents; diagnose low-level I/O issues; and optimize storage systems and architecture.

Requirements

  • Bachelor's degree in Computer Science, Electrical Engineering, or a related technical field, or equivalent practical experience
  • 5+ years of professional experience in Storage SRE, systems engineering, or storage engineering
  • Hands-on experience with enterprise storage platforms such as Pure Storage or EMC
  • Understanding of object, block, and file storage paradigms
  • Proficiency in Go, Python, Java, or C
  • Experience with Infrastructure as Code and deployment tooling
  • Knowledge of Linux I/O subsystems, memory management, and storage scheduling
  • Familiarity with NFS, SMB, iSCSI, or NVMe-oF
  • Experience with containerized workloads and Kubernetes or Docker
  • Incident response, troubleshooting, and documentation experience
  • Experience operating managed storage services at scale
  • Experience with distributed storage systems is a bonus
  • Experience with hybrid storage models is a bonus

Responsibilities

  • Build automation and self-healing tools
  • Monitor and maintain distributed cloud storage infrastructure
  • Improve data replication, encryption, backup, restore, and failover
  • Support user-facing storage services
  • Tune storage performance
  • Investigate and resolve storage incidents
  • Diagnose low-level I/O issues
  • Optimize I/O paths, cache policies, and file systems
  • Contribute to fault-tolerant storage backend architecture

Benefits

  • Industry competitive pay
  • Restricted Stock Units
  • Health insurance options including HDHP and PPO
  • Vision insurance
  • Dental insurance
  • Employer HSA contributions
  • Paid parental leave
  • Paid life insurance
  • Short-term and long-term disability insurance
  • Teladoc
  • 401(k) with 100% match up to 4% of salary
  • Paid time off
  • Paid holidays
  • Cell phone reimbursement
  • Tuition reimbursement
  • Calm app subscription
  • MetLife Legal
  • Company-paid commuter benefit of $300 per month