Manager, Engineering
Crusoe is an AI infrastructure company that designs, builds, and operates AI data centers and a cloud platform. It provides managed AI services, GPU compute, model fine-tuning and inference, and infrastructure operations for organizations building and deploying AI workloads.
Projects
About Crusoe
Crusoe, the AI factory company, provides Crusoe Cloud and Crusoe Intelligence Foundry for AI development and production. Its offerings include managed inference, serverless fine-tuning, high-performance NVIDIA and AMD compute, accelerated storage, RDMA networking, managed Kubernetes and Slurm, and operations tooling. The company also designs, builds, and operates modular AI data-center infrastructure using an energy-first approach, serving customers that need scalable training, inference, and AI platform infrastructure.
Skills
About the Role
You will establish and lead the Production Engineering presence in Tel Aviv while remaining hands-on with code and incident response. You will recruit and mentor systems engineers, operate a global follow-the-sun on-call rotation, improve monitoring and alert quality, automate operational workflows, and enforce production-readiness and change-control standards. You will reserve engineering capacity for strategic automation, tooling, and firmware optimization, converting recurring physical interventions into software-defined remediation.
Requirements
- 8+ years of experience in infrastructure, SRE, or production engineering
- 2+ years of experience leading first-line engineering teams in a high-growth neocloud, hyperscaler, or large-scale distributed environment
- Hands-on software engineering proficiency in Go, Python, C++, or a comparable systems language
- Expert knowledge of Linux internals, container orchestration at scale, and root-cause analysis across physical-to-virtual boundaries
- Experience operating tiered on-call models, defining SLIs, SLOs, and error budgets, and reducing paging fatigue
Responsibilities
- Recruit, mentor, and establish a Production Engineering team in Tel Aviv
- Partner with US and Dublin teams to operate a follow-the-sun global on-call rotation
- Champion blameless post-mortems focused on systemic failures
- Reduce alerts and improve fleet signal-to-noise ratios
- Automate routine workflows using runbook automation
- Build predictive monitoring for SEV1 and SEV2 events
- Enforce production readiness reviews and change control across compute, storage, networking, and platform teams
- Protect engineering capacity for automation, tooling, and firmware optimization
- Convert recurring physical interventions into software-defined auto-remediation
Benefits
- Pension contributions
