Engineering Manager, Managed Platform Services

Crusoe is an AI infrastructure and cloud computing company. It provides GPU cloud capacity, managed AI services, inference, fine-tuning, data centers, and energy infrastructure for AI developers and enterprise customers.

Maintainer signals as of 8/23/2026

Distributed
About Crusoe, Inc

Crusoe designs, builds, and operates energy-first AI infrastructure, including data centers, GPU cloud computing, and modular AI factories. Crusoe Cloud provides GPU clusters, managed Kubernetes and Slurm, storage, networking, observability, managed inference, serverless fine-tuning, and model deployment through Crusoe Intelligence Foundry. Its customers include AI startups, enterprises, and organizations developing training, inference, analytics, and other compute-intensive workloads.

View jobs by Crusoe, Inc

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

Lead the Command Center Insights & Actions team, owning alerting infrastructure, control plane APIs, automated actions, telemetry insights, anomaly detection, GPU profiling, and remediation systems while managing complex projects and coaching engineers.

Requirements

  • Hands-on background in ML, heuristics, or rule-based systems
  • Experience with anomaly detection, threshold design, and automated remediation logic
  • People management experience
  • Technical leadership in ambiguous problem spaces
  • Technical communication skills
  • Experience owning and delivering complex projects end-to-end
  • Experience building and operating global services at scale
  • Organizational and prioritization skills
  • Background in data platforms and data science
  • Background in observability platforms or products
  • Familiarity with GPU profiling tools or infrastructure diagnostics

Responsibilities

  • Own and execute the Insights and Actions roadmap
  • Drive alerting infrastructure, control plane APIs, and automated action systems
  • Develop telemetry-derived insights, including straggler node detection and GPU profiling
  • Collaborate with product and engineering leadership on requirements
  • Partner with product, design, and engineering teams
  • Lead complex initiatives involving multiple engineers
  • Champion process improvements and operational excellence
  • Coach and mentor engineers from new graduate to Staff level
  • Set performance expectations and define career paths

Benefits

  • Restricted Stock Units
  • Paid time off and paid holidays
  • Comprehensive health, dental, and vision insurance
  • Employer HSA contributions
  • Paid parental leave
  • Life and disability insurance
  • Professional development and tuition reimbursement
  • Mental health and wellness support
  • Commuter benefits
  • Cell phone stipend
  • 401(k) plan with company match up to 4% of salary
  • Volunteer time off
  • Global travel insurance and emergency assistance
  • Daily meals allowance
  • Location-specific perks