Search...

Staff Production Engineer, Core PE

Crusoe logo
Crusoe

Crusoe is an AI infrastructure company that designs, builds, and operates AI data centers and a cloud platform. It provides managed AI services, GPU compute, model fine-tuning and inference, and infrastructure operations for organizations building and deploying AI workloads.

Maintainer signals as of 8/14/2026

Distributed
About Crusoe

Crusoe, the AI factory company, provides Crusoe Cloud and Crusoe Intelligence Foundry for AI development and production. Its offerings include managed inference, serverless fine-tuning, high-performance NVIDIA and AMD compute, accelerated storage, RDMA networking, managed Kubernetes and Slurm, and operations tooling. The company also designs, builds, and operates modular AI data-center infrastructure using an energy-first approach, serving customers that need scalable training, inference, and AI platform infrastructure.

View jobs by Crusoe

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will lead reliability strategy for a GPU cloud, define availability metrics, manage production incidents, improve observability, identify reliability risks, build self-healing automation, strengthen disaster recovery, establish operational processes, and mentor production and infrastructure engineers.

Requirements

  • Bachelor's degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience
  • 8+ years of experience in Production Engineering, SRE, or large-scale infrastructure operations
  • Experience supporting GPU workloads, HPC environments, or latency- and throughput-sensitive distributed systems
  • Experience building or managing compute, storage, or networking platforms
  • Deep knowledge of Linux/Unix systems
  • Understanding of Kubernetes, distributed systems, virtualization, and AWS/GCP cloud platforms
  • Experience with incident management and reliability frameworks
  • Hands-on experience with Prometheus and Grafana
  • Experience with Terraform or Ansible
  • Proficiency in Go, Python, C, or C++
  • Experience troubleshooting complex production issues
  • Commitment to reliability engineering, automation, and operational excellence

Responsibilities

  • Define and evolve cloud platform availability metrics
  • Drive production incident response and lead post-incident reviews
  • Conduct root cause analysis for service disruptions
  • Architect and improve infrastructure observability
  • Identify reliability risks and performance bottlenecks
  • Develop automation and tooling for self-healing infrastructure
  • Strengthen service resilience and disaster recovery
  • Define operational processes and reliability best practices
  • Mentor junior and mid-level engineers

Benefits

  • Industry competitive pay
  • Restricted Stock Units
  • Health insurance options including HDHP and PPO
  • Vision insurance
  • Dental insurance
  • Employer contributions to HSA accounts
  • Paid parental leave
  • Paid life insurance
  • Short-term and long-term disability insurance
  • Teladoc
  • 401(k) with 100% match up to 4% of salary
  • Generous paid time off and holiday schedule
  • Cell phone reimbursement
  • Tuition reimbursement
  • Subscription to the Calm app
  • MetLife Legal
  • Company-paid commuter benefit of $300 per month