Senior Production Engineer Managed Cloud
Crusoe is an AI infrastructure company that designs, builds, and operates AI data centers and a cloud platform. It provides managed AI services, GPU compute, model fine-tuning and inference, and infrastructure operations for organizations building and deploying AI workloads.
Maintainer signals as of 8/12/2026
Funding history
Projects
About Crusoe
Crusoe, the AI factory company, provides Crusoe Cloud and Crusoe Intelligence Foundry for AI development and production. Its offerings include managed inference, serverless fine-tuning, high-performance NVIDIA and AMD compute, accelerated storage, RDMA networking, managed Kubernetes and Slurm, and operations tooling. The company also designs, builds, and operates modular AI data-center infrastructure using an energy-first approach, serving customers that need scalable training, inference, and AI platform infrastructure.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will design and operate reliable managed cloud services at scale. You will define SLIs and SLOs, improve observability and performance, investigate distributed-system reliability issues, optimize AI infrastructure, and contribute to fault-tolerant systems architecture.
Requirements
- Strong software engineering background
- Experience building production-grade systems beyond scripting or Bash
- Experience designing and implementing distributed systems
- Experience defining and measuring SLIs and SLOs
- Experience building monitoring and observability systems
- Experience driving performance and reliability improvements
- Experience designing fault-tolerant systems and automated testing strategies
- Proficiency in Python, Go, Java, or C++
- Familiarity with Kubernetes or container orchestration platforms
Responsibilities
- Design and operate reliable managed cloud services
- Define and improve SLIs and SLOs
- Collaborate with AI, platform, and infrastructure teams
- Optimize training and inference clusters
- Build telemetry and performance-tuning strategies
- Investigate and resolve reliability issues
- Contribute to distributed systems architecture
Benefits
- Industry competitive pay
- Restricted Stock Units
- Health insurance options including HDHP and PPO
- Vision insurance
- Dental insurance
- Employer HSA contributions
- Paid parental leave
- Paid life insurance
- Short-term and long-term disability insurance
- Teladoc
- 401(k) with 100% match up to 4% of salary
- Paid time off
- Paid holidays
- Cell phone reimbursement
- Tuition reimbursement
- Calm app subscription
- MetLife Legal
- Company-paid commuter benefit of $300 per month
