Search...

Senior Production Engineer, Core PE

Crusoe logo
Crusoe

Crusoe is an AI infrastructure company that designs, builds, and operates AI data centers and a cloud platform. It provides managed AI services, GPU compute, model fine-tuning and inference, and infrastructure operations for organizations building and deploying AI workloads.

Maintainer signals as of 8/14/2026

Distributed
About Crusoe

Crusoe, the AI factory company, provides Crusoe Cloud and Crusoe Intelligence Foundry for AI development and production. Its offerings include managed inference, serverless fine-tuning, high-performance NVIDIA and AMD compute, accelerated storage, RDMA networking, managed Kubernetes and Slurm, and operations tooling. The company also designs, builds, and operates modular AI data-center infrastructure using an energy-first approach, serving customers that need scalable training, inference, and AI platform infrastructure.

View jobs by Crusoe

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will improve the reliability, scalability, and performance of Crusoe’s GPU cloud. You will define availability metrics, respond to incidents, build observability and automation, identify reliability risks, strengthen disaster recovery, and improve operational practices across large-scale AI infrastructure.

Requirements

  • 5+ years of experience in Production Engineering, SRE, or large-scale infrastructure operations
  • Experience supporting GPU workloads, HPC environments, or latency- and throughput-sensitive distributed systems
  • Strong knowledge of Linux/Unix systems
  • Experience building or managing compute, storage, or networking platforms
  • Understanding of Kubernetes, distributed systems, virtualization, and cloud platforms such as AWS and GCP
  • Familiarity with incident management practices and reliability frameworks
  • Experience with Prometheus, Grafana, or observability tools
  • Familiarity with Terraform or Ansible
  • Scripting or programming experience with Go, Python, C, or C++
  • Strong communication and cross-team collaboration skills

Responsibilities

  • Define, measure, and improve SLIs and SLOs for the cloud platform
  • Respond to production incidents and contribute to post-incident reviews and root cause analysis
  • Build and improve observability using Prometheus, Grafana, Alertmanager, and OpenTelemetry
  • Identify reliability risks, performance bottlenecks, and potential production issues
  • Develop automation and tooling to reduce operational toil and enable self-healing infrastructure
  • Strengthen service resilience and disaster recovery capabilities
  • Improve operational processes and reliability practices

Benefits

  • Pension contributions
  • Private health insurance
  • Dental insurance
  • Income protection
  • Life assurance
Senior Production Engineer, Core PE at Crusoe | JobStash