Senior Production Engineer, Managed Cloud
Crusoe is an AI infrastructure and cloud computing company. It provides GPU cloud capacity, managed AI services, inference, fine-tuning, data centers, and energy infrastructure for AI developers and enterprise customers.
Maintainer signals as of 8/23/2026
Funding history
Projects
About Crusoe, Inc
Crusoe designs, builds, and operates energy-first AI infrastructure, including data centers, GPU cloud computing, and modular AI factories. Crusoe Cloud provides GPU clusters, managed Kubernetes and Slurm, storage, networking, observability, managed inference, serverless fine-tuning, and model deployment through Crusoe Intelligence Foundry. Its customers include AI startups, enterprises, and organizations developing training, inference, analytics, and other compute-intensive workloads.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
Design and operate reliable managed cloud services at scale, define SLIs and SLOs, improve observability and performance, investigate distributed-system reliability issues, optimize AI infrastructure, and contribute to fault-tolerant systems architecture.
Requirements
- Strong software engineering background and experience building production-grade systems beyond scripting or Bash
- Experience designing and implementing distributed systems
- SRE experience defining and measuring SLIs and SLOs
- Experience building monitoring and observability systems
- Experience driving performance and reliability improvements
- Experience designing fault-tolerant systems and automated testing strategies
- Proficiency in Python, Go, Java, or C++
- Familiarity with Kubernetes or container orchestration platforms
- Strong collaboration and communication skills
- Ability to thrive in a fast-paced, mission-driven environment
Responsibilities
- Design and operate reliable managed AI services focused on serving and scaling LLM workloads
- Define, measure, and improve SLIs and SLOs
- Collaborate with AI, platform, and infrastructure teams to optimize training and inference clusters
- Build telemetry and performance-tuning strategies for latency-sensitive services
- Investigate and resolve reliability issues using telemetry, logs, and profiling
- Contribute to next-generation distributed systems architecture
Benefits
- Industry competitive pay
- Restricted Stock Units
- Health insurance options including HDHP and PPO
- Vision insurance
- Dental insurance
- Employer HSA contributions
- Paid parental leave
- Paid life insurance
- Short-term and long-term disability insurance
- Teladoc
- 401(k) with 100% match up to 4% of salary
- Paid time off
- Paid holidays
- Cell phone reimbursement
- Tuition reimbursement
- Calm app subscription
- MetLife Legal
- Company-paid commuter benefit of $300 per month
