Staff Production Engineer, Core PE
Crusoe is an AI infrastructure and cloud computing company. It provides GPU cloud capacity, managed AI services, inference, fine-tuning, data centers, and energy infrastructure for AI developers and enterprise customers.
Maintainer signals as of 8/23/2026
Funding history
Projects
About Crusoe, Inc
Crusoe designs, builds, and operates energy-first AI infrastructure, including data centers, GPU cloud computing, and modular AI factories. Crusoe Cloud provides GPU clusters, managed Kubernetes and Slurm, storage, networking, observability, managed inference, serverless fine-tuning, and model deployment through Crusoe Intelligence Foundry. Its customers include AI startups, enterprises, and organizations developing training, inference, analytics, and other compute-intensive workloads.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
Lead reliability strategy for a GPU cloud, define availability metrics, manage production incidents, improve observability, identify reliability risks, build self-healing automation, strengthen disaster recovery, establish operational processes, and mentor production and infrastructure engineers.
Requirements
- Bachelor's degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience
- 8+ years of experience in Production Engineering, SRE, or large-scale infrastructure operations
- Experience supporting GPU workloads, HPC environments, or latency- and throughput-sensitive distributed systems
- Experience building or managing compute, storage, or networking platforms
- Deep knowledge of Linux/Unix systems
- Understanding of Kubernetes, distributed systems, virtualization, and AWS/GCP cloud platforms
- Experience with incident management and reliability frameworks
- Hands-on experience with Prometheus and Grafana
- Experience with Terraform or Ansible
- Proficiency in Go, Python, C, or C++
- Experience troubleshooting complex production issues
- Commitment to reliability engineering, automation, and operational excellence
Responsibilities
- Define and evolve cloud platform availability metrics
- Drive production incident response and lead post-incident reviews
- Conduct root cause analysis for service disruptions
- Architect and improve infrastructure observability
- Identify reliability risks and performance bottlenecks
- Develop automation and tooling for self-healing infrastructure
- Strengthen service resilience and disaster recovery
- Define operational processes and reliability best practices
- Mentor junior and mid-level engineers
Benefits
- Industry competitive pay
- Restricted Stock Units
- Health insurance options including HDHP and PPO
- Vision insurance
- Dental insurance
- Employer contributions to HSA accounts
- Paid parental leave
- Paid life insurance
- Short-term and long-term disability insurance
- Teladoc
- 401(k) with 100% match up to 4% of salary
- Generous paid time off and holiday schedule
- Cell phone reimbursement
- Tuition reimbursement
- Subscription to the Calm app
- MetLife Legal
- Company-paid commuter benefit of $300 per month
