Staff Production Engineer, Core PE

Crusoe is an AI infrastructure and cloud computing company. It provides GPU cloud capacity, managed AI services, inference, fine-tuning, data centers, and energy infrastructure for AI developers and enterprise customers.

Maintainer signals as of 8/23/2026

Distributed
About Crusoe, Inc

Crusoe designs, builds, and operates energy-first AI infrastructure, including data centers, GPU cloud computing, and modular AI factories. Crusoe Cloud provides GPU clusters, managed Kubernetes and Slurm, storage, networking, observability, managed inference, serverless fine-tuning, and model deployment through Crusoe Intelligence Foundry. Its customers include AI startups, enterprises, and organizations developing training, inference, analytics, and other compute-intensive workloads.

View jobs by Crusoe, Inc

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

Lead reliability strategy for a GPU cloud, define availability metrics, manage production incidents, improve observability, identify reliability risks, build self-healing automation, strengthen disaster recovery, establish operational processes, and mentor production and infrastructure engineers.

Requirements

  • Bachelor's degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience
  • 8+ years of experience in Production Engineering, SRE, or large-scale infrastructure operations
  • Experience supporting GPU workloads, HPC environments, or latency- and throughput-sensitive distributed systems
  • Experience building or managing compute, storage, or networking platforms
  • Deep knowledge of Linux/Unix systems
  • Understanding of Kubernetes, distributed systems, virtualization, and AWS/GCP cloud platforms
  • Experience with incident management and reliability frameworks
  • Hands-on experience with Prometheus and Grafana
  • Experience with Terraform or Ansible
  • Proficiency in Go, Python, C, or C++
  • Experience troubleshooting complex production issues
  • Commitment to reliability engineering, automation, and operational excellence

Responsibilities

  • Define and evolve cloud platform availability metrics
  • Drive production incident response and lead post-incident reviews
  • Conduct root cause analysis for service disruptions
  • Architect and improve infrastructure observability
  • Identify reliability risks and performance bottlenecks
  • Develop automation and tooling for self-healing infrastructure
  • Strengthen service resilience and disaster recovery
  • Define operational processes and reliability best practices
  • Mentor junior and mid-level engineers

Benefits

  • Industry competitive pay
  • Restricted Stock Units
  • Health insurance options including HDHP and PPO
  • Vision insurance
  • Dental insurance
  • Employer contributions to HSA accounts
  • Paid parental leave
  • Paid life insurance
  • Short-term and long-term disability insurance
  • Teladoc
  • 401(k) with 100% match up to 4% of salary
  • Generous paid time off and holiday schedule
  • Cell phone reimbursement
  • Tuition reimbursement
  • Subscription to the Calm app
  • MetLife Legal
  • Company-paid commuter benefit of $300 per month