Senior Production Engineer Core PE
Crusoe is an AI infrastructure company that designs, builds, and operates AI data centers and a cloud platform. It provides managed AI services, GPU compute, model fine-tuning and inference, and infrastructure operations for organizations building and deploying AI workloads.
Maintainer signals as of 8/14/2026
Funding history
Projects
About Crusoe
Crusoe, the AI factory company, provides Crusoe Cloud and Crusoe Intelligence Foundry for AI development and production. Its offerings include managed inference, serverless fine-tuning, high-performance NVIDIA and AMD compute, accelerated storage, RDMA networking, managed Kubernetes and Slurm, and operations tooling. The company also designs, builds, and operates modular AI data-center infrastructure using an energy-first approach, serving customers that need scalable training, inference, and AI platform infrastructure.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will strengthen the reliability, scalability, and performance of a GPU cloud platform. You will define and improve SLIs and SLOs, respond to incidents, build observability and automation, identify reliability risks, improve recovery and self-healing capabilities, and partner with compute, networking, storage, and platform teams.
Requirements
- 5+ years of experience in Production Engineering, SRE, or large-scale infrastructure operations
- Experience supporting GPU workloads, HPC environments, or latency- and throughput-sensitive distributed systems
- Strong knowledge of Linux and Unix systems
- Experience building or managing compute, storage, or networking platforms
- Understanding of Kubernetes, distributed systems, virtualization, AWS, or GCP
- Familiarity with incident management and reliability frameworks such as SRE or ITIL
- Experience with Prometheus and Grafana, or willingness to deepen expertise
- Familiarity with Terraform or Ansible
- Scripting or programming experience in Go, Python, C, or C++
- Bachelor's degree in Computer Science, Electrical Engineering, or a related technical field, or equivalent experience
Responsibilities
- Define, measure, and improve availability metrics, SLIs, and SLOs
- Participate in production incident response
- Diagnose and resolve service disruptions
- Contribute to post-incident reviews and root cause analysis
- Build and improve infrastructure observability with Prometheus, Grafana, Alertmanager, and OpenTelemetry
- Identify reliability risks and performance bottlenecks
- Develop automation and tooling to reduce operational toil
- Improve recovery times and enable self-healing infrastructure
- Strengthen service resilience and disaster recovery capabilities
- Improve operational processes and reliability practices
Benefits
- Industry competitive pay
- Restricted Stock Units
- Health insurance
- Vision insurance
- Dental insurance
- Employer HSA contributions
- Paid parental leave
- Paid life insurance
- Short-term disability insurance
- Long-term disability insurance
- Teladoc
- 401(k) with 100% match up to 4% of salary
- Paid time off
- Paid holidays
- Cell phone reimbursement
- Tuition reimbursement
- Calm app subscription
- MetLife Legal
- Company-paid commuter benefit of $300 per month
