Principal Production Engineer
Crusoe is an AI infrastructure company that designs, builds, and operates AI data centers and a cloud platform. It provides managed AI services, GPU compute, model fine-tuning and inference, and infrastructure operations for organizations building and deploying AI workloads.
Projects
About Crusoe
Crusoe, the AI factory company, provides Crusoe Cloud and Crusoe Intelligence Foundry for AI development and production. Its offerings include managed inference, serverless fine-tuning, high-performance NVIDIA and AMD compute, accelerated storage, RDMA networking, managed Kubernetes and Slurm, and operations tooling. The company also designs, builds, and operates modular AI data-center infrastructure using an energy-first approach, serving customers that need scalable training, inference, and AI platform infrastructure.
Skills
About the Role
You will own the reliability, scalability, and operational excellence of cloud infrastructure across compute, storage, networking, and supporting platform tooling. You will define service-level objectives, lead incident response, and drive systemic improvements that reduce operational toil. You will build observability and on-call tooling, influence architecture decisions with software, hardware, and network engineering partners, set production engineering standards, and mentor senior and staff engineers.
Requirements
- 15+ years of experience in infrastructure, networking, or production engineering at internet-scale companies
- Strong systems fundamentals in Linux, distributed systems, storage, and compute scheduling
- Hands-on data center infrastructure experience, including power and thermal constraints
- Ability to write code for automation, instrumentation, and tooling
- Excellent analytical and problem-solving skills
- Strong incident command skills and experience running blameless retrospectives
Responsibilities
- Own the reliability and scalability of cloud infrastructure across compute, storage, and networking
- Define service-level objectives, lead incident response, and drive systemic reliability improvements
- Build and mature observability, telemetry, instrumentation, and on-call tooling
- Drive reliability improvements across the cloud stack and influence architecture decisions early
- Advise senior leadership on observability trends and long-term technology investments
- Set standards for on-call culture, incident frameworks, and reliability practices
- Mentor senior and staff engineers and elevate technical depth
Benefits
- Restricted Stock Units
- Health insurance options including HDHP and PPO plans, vision, and dental coverage for employees and dependents
- Employer contributions to HSA accounts
- Paid parental leave
- Paid life insurance
- Short-term and long-term disability insurance
- Teladoc
- 401(k) with a 100% match up to 4% of salary
- Generous paid time off and holiday schedule
- Cell phone reimbursement
- Tuition reimbursement
- Subscription to the Calm app
- MetLife Legal
- Company-paid commuter benefit of $300 per month
