Senior Production Engineer, Core PE
Crusoe is an AI infrastructure and cloud computing company. It provides GPU cloud capacity, managed AI services, inference, fine-tuning, data centers, and energy infrastructure for AI developers and enterprise customers.
Maintainer signals as of 8/23/2026
Funding history
Projects
About Crusoe, Inc
Crusoe designs, builds, and operates energy-first AI infrastructure, including data centers, GPU cloud computing, and modular AI factories. Crusoe Cloud provides GPU clusters, managed Kubernetes and Slurm, storage, networking, observability, managed inference, serverless fine-tuning, and model deployment through Crusoe Intelligence Foundry. Its customers include AI startups, enterprises, and organizations developing training, inference, analytics, and other compute-intensive workloads.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
Improve the reliability, scalability, and performance of Crusoe’s GPU cloud by defining availability metrics, responding to incidents, building observability and automation, identifying reliability risks, strengthening disaster recovery, and improving operational practices across large-scale AI infrastructure.
Requirements
- 5+ years of experience in Production Engineering, SRE, or large-scale infrastructure operations
- Experience supporting GPU workloads, HPC environments, or latency- and throughput-sensitive distributed systems
- Strong knowledge of Linux/Unix systems
- Experience building or managing compute, storage, or networking platforms
- Understanding of Kubernetes, distributed systems, virtualization, and cloud platforms such as AWS and GCP
- Familiarity with incident management practices and reliability frameworks
- Experience with Prometheus, Grafana, or observability tools
- Familiarity with Terraform or Ansible
- Scripting or programming experience with Go, Python, C, or C++
- Strong communication and cross-team collaboration skills
Responsibilities
- Define, measure, and improve SLIs and SLOs for the cloud platform
- Respond to production incidents and contribute to post-incident reviews and root cause analysis
- Build and improve observability using Prometheus, Grafana, Alertmanager, and OpenTelemetry
- Identify reliability risks, performance bottlenecks, and potential production issues
- Develop automation and tooling to reduce operational toil and enable self-healing infrastructure
- Strengthen service resilience and disaster recovery capabilities
- Improve operational processes and reliability practices
Benefits
- Pension contributions
- Private health insurance
- Dental insurance
- Income protection
- Life assurance
