Manager, Engineering

Crusoe is an AI infrastructure and cloud computing company. It provides GPU cloud capacity, managed AI services, inference, fine-tuning, data centers, and energy infrastructure for AI developers and enterprise customers.

Maintainer signals as of 8/23/2026

Distributed
About Crusoe, Inc

Crusoe designs, builds, and operates energy-first AI infrastructure, including data centers, GPU cloud computing, and modular AI factories. Crusoe Cloud provides GPU clusters, managed Kubernetes and Slurm, storage, networking, observability, managed inference, serverless fine-tuning, and model deployment through Crusoe Intelligence Foundry. Its customers include AI startups, enterprises, and organizations developing training, inference, analytics, and other compute-intensive workloads.

View jobs by Crusoe, Inc

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will establish and lead the Production Engineering presence in Tel Aviv while remaining hands-on with code and incident response. You will recruit and mentor systems engineers, operate a global follow-the-sun on-call rotation, improve monitoring and alert quality, automate operational workflows, and enforce production-readiness and change-control standards. You will reserve engineering capacity for strategic automation, tooling, and firmware optimization, converting recurring physical interventions into software-defined remediation.

Requirements

  • 8+ years of experience in infrastructure, SRE, or production engineering
  • 2+ years of experience leading first-line engineering teams in a high-growth neocloud, hyperscaler, or large-scale distributed environment
  • Hands-on software engineering proficiency in Go, Python, C++, or a comparable systems language
  • Expert knowledge of Linux internals, container orchestration at scale, and root-cause analysis across physical-to-virtual boundaries
  • Experience operating tiered on-call models, defining SLIs, SLOs, and error budgets, and reducing paging fatigue

Responsibilities

  • Recruit, mentor, and establish a Production Engineering team in Tel Aviv
  • Partner with US and Dublin teams to operate a follow-the-sun global on-call rotation
  • Champion blameless post-mortems focused on systemic failures
  • Reduce alerts and improve fleet signal-to-noise ratios
  • Automate routine workflows using runbook automation
  • Build predictive monitoring for SEV1 and SEV2 events
  • Enforce production readiness reviews and change control across compute, storage, networking, and platform teams
  • Protect engineering capacity for automation, tooling, and firmware optimization
  • Convert recurring physical interventions into software-defined auto-remediation

Benefits

  • Pension contributions