Search...

Engineering Manager Managed Platform Services

Crusoe logo
Crusoe

Crusoe is an AI infrastructure company that designs, builds, and operates AI data centers and a cloud platform. It provides managed AI services, GPU compute, model fine-tuning and inference, and infrastructure operations for organizations building and deploying AI workloads.

Series C13 current maintainers10 active leads5 new active leads9 lead step-downs1 early lead departureTeam intelligence

Maintainer signals as of 8/12/2026

Distributed
About Crusoe

Crusoe, the AI factory company, provides Crusoe Cloud and Crusoe Intelligence Foundry for AI development and production. Its offerings include managed inference, serverless fine-tuning, high-performance NVIDIA and AMD compute, accelerated storage, RDMA networking, managed Kubernetes and Slurm, and operations tooling. The company also designs, builds, and operates modular AI data-center infrastructure using an energy-first approach, serving customers that need scalable training, inference, and AI platform infrastructure.

View jobs by Crusoe

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will lead the Command Center Insights and Actions team and own its technical roadmap. You will guide alerting infrastructure, control plane APIs, automated actions, telemetry insights, anomaly detection, and remediation systems, clarify product requirements, manage complex projects, improve operations, and coach engineers across experience levels.

Requirements

  • Hands-on background in ML, heuristics, or rule-based systems
  • Experience with anomaly detection, threshold design, and automated remediation logic
  • People management experience
  • Technical leadership in ambiguous problem spaces
  • Technical communication skills
  • Experience owning and delivering complex projects end-to-end
  • Experience building and operating global services at scale
  • Organizational and prioritization skills
  • Background in data platforms and data science
  • Background in observability platforms or products
  • Familiarity with GPU profiling tools or infrastructure diagnostics

Responsibilities

  • Own and execute the Insights and Actions roadmap
  • Drive alerting infrastructure, control plane APIs, and automated action systems
  • Develop telemetry-derived insights such as straggler node detection and GPU profiling
  • Collaborate with product and engineering leadership on requirements
  • Partner with product, design, and engineering teams
  • Lead complex initiatives involving multiple engineers
  • Champion process improvements and operational excellence
  • Coach and mentor engineers from new graduate to Staff level
  • Set performance expectations and define career paths

Benefits

  • Restricted Stock Units
  • Paid time off
  • Paid holidays
  • Leave of absence programs
  • Comprehensive health insurance
  • Dental insurance
  • Vision insurance
  • Employer contributions to HSA account
  • Paid parental leave
  • Paid life insurance
  • Short-term disability insurance
  • Long-term disability insurance
  • Professional development
  • Tuition reimbursement
  • Mental health and wellness support
  • Commuter benefits
  • Cell phone stipend
  • 401(k) retirement plan with company match up to 4% of salary
  • Volunteer time off
  • Global travel insurance and emergency assistance
  • Daily meals allowance
  • Additional location-specific perks