Principal Production Engineer
Crusoe is an AI infrastructure and cloud computing company. It provides GPU cloud capacity, managed AI services, inference, fine-tuning, data centers, and energy infrastructure for AI developers and enterprise customers.
Maintainer signals as of 8/23/2026
Funding history
Projects
About Crusoe, Inc
Crusoe designs, builds, and operates energy-first AI infrastructure, including data centers, GPU cloud computing, and modular AI factories. Crusoe Cloud provides GPU clusters, managed Kubernetes and Slurm, storage, networking, observability, managed inference, serverless fine-tuning, and model deployment through Crusoe Intelligence Foundry. Its customers include AI startups, enterprises, and organizations developing training, inference, analytics, and other compute-intensive workloads.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
Own the reliability, scalability, and operational excellence of cloud infrastructure across compute, storage, networking, and supporting platform tooling. Define service-level objectives, lead incident response, drive systemic improvements, build observability and on-call tooling, influence architecture decisions, set production engineering standards, and mentor senior and staff engineers.
Requirements
- 15+ years of experience in infrastructure, networking, or production engineering at internet-scale companies
- Strong systems fundamentals in Linux, distributed systems, storage, and compute scheduling
- Hands-on data center infrastructure experience, including power and thermal constraints
- Ability to write code for automation, instrumentation, and tooling
- Excellent analytical and problem-solving skills
- Strong incident command skills and experience running blameless retrospectives
Responsibilities
- Own the reliability and scalability of cloud infrastructure across compute, storage, and networking
- Define service-level objectives, lead incident response, and drive systemic reliability improvements
- Build and mature observability, telemetry, instrumentation, and on-call tooling
- Drive reliability improvements across the cloud stack and influence architecture decisions
- Advise senior leadership on observability trends and long-term technology investments
- Set standards for on-call culture, incident frameworks, and reliability practices
- Mentor senior and staff engineers and elevate technical depth
Benefits
- Restricted Stock Units
- Health insurance options including HDHP and PPO plans, vision, and dental coverage
- Employer contributions to HSA accounts
- Paid parental leave
- Paid life insurance
- Short-term and long-term disability insurance
- Teladoc
- 401(k) with a 100% match up to 4% of salary
- Generous paid time off and holiday schedule
- Cell phone reimbursement
- Tuition reimbursement
- Subscription to the Calm app
- MetLife Legal
- Company-paid commuter benefit of $300 per month
