Senior Production Engineer, Core PE

Crusoe is an AI infrastructure and cloud computing company. It provides GPU cloud capacity, managed AI services, inference, fine-tuning, data centers, and energy infrastructure for AI developers and enterprise customers.

Maintainer signals as of 8/23/2026

Distributed
About Crusoe, Inc

Crusoe designs, builds, and operates energy-first AI infrastructure, including data centers, GPU cloud computing, and modular AI factories. Crusoe Cloud provides GPU clusters, managed Kubernetes and Slurm, storage, networking, observability, managed inference, serverless fine-tuning, and model deployment through Crusoe Intelligence Foundry. Its customers include AI startups, enterprises, and organizations developing training, inference, analytics, and other compute-intensive workloads.

View jobs by Crusoe, Inc

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will strengthen the reliability, scalability, and performance of a GPU cloud platform. You will define and improve SLIs and SLOs, respond to incidents, build observability and automation, identify reliability risks, improve recovery and self-healing capabilities, and partner with compute, networking, storage, and platform teams.

Requirements

  • 5+ years of experience in Production Engineering, SRE, or large-scale infrastructure operations
  • Experience supporting GPU workloads, HPC environments, or latency- and throughput-sensitive distributed systems
  • Strong knowledge of Linux and Unix systems
  • Experience building or managing compute, storage, or networking platforms
  • Understanding of Kubernetes, distributed systems, virtualization, AWS, or GCP
  • Familiarity with incident management and reliability frameworks such as SRE or ITIL
  • Experience with Prometheus and Grafana, or willingness to deepen expertise
  • Familiarity with Terraform or Ansible
  • Scripting or programming experience in Go, Python, C, or C++
  • Bachelor's degree in Computer Science, Electrical Engineering, or a related technical field, or equivalent experience

Responsibilities

  • Define, measure, and improve availability metrics, SLIs, and SLOs
  • Participate in production incident response
  • Diagnose and resolve service disruptions
  • Contribute to post-incident reviews and root cause analysis
  • Build and improve infrastructure observability with Prometheus, Grafana, Alertmanager, and OpenTelemetry
  • Identify reliability risks and performance bottlenecks
  • Develop automation and tooling to reduce operational toil
  • Improve recovery times and enable self-healing infrastructure
  • Strengthen service resilience and disaster recovery capabilities
  • Improve operational processes and reliability practices

Benefits

  • Industry competitive pay
  • Restricted Stock Units
  • Health insurance
  • Vision insurance
  • Dental insurance
  • Employer HSA contributions
  • Paid parental leave
  • Paid life insurance
  • Short-term disability insurance
  • Long-term disability insurance
  • Teladoc
  • 401(k) with 100% match up to 4% of salary
  • Paid time off
  • Paid holidays
  • Cell phone reimbursement
  • Tuition reimbursement
  • Calm app subscription
  • MetLife Legal
  • Company-paid commuter benefit of $300 per month