Search...

Principal Engineer, Conductor Platform (CAPE)

Crusoe logo
Crusoe

Crusoe is an AI infrastructure company that designs, builds, and operates AI data centers and a cloud platform. It provides managed AI services, GPU compute, model fine-tuning and inference, and infrastructure operations for organizations building and deploying AI workloads.

Series C13 current maintainers10 active leads5 new active leads9 lead step-downs1 early lead departureTeam intelligence

Maintainer signals as of 8/12/2026

Distributed
About Crusoe

Crusoe, the AI factory company, provides Crusoe Cloud and Crusoe Intelligence Foundry for AI development and production. Its offerings include managed inference, serverless fine-tuning, high-performance NVIDIA and AMD compute, accelerated storage, RDMA networking, managed Kubernetes and Slurm, and operations tooling. The company also designs, builds, and operates modular AI data-center infrastructure using an energy-first approach, serving customers that need scalable training, inference, and AI platform infrastructure.

View jobs by Crusoe

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will lead the architecture and development of a control plane for Crusoe's AI infrastructure platform. You will define the platform architecture, topology graph, reconciliation loop, policy engine, hardware lifecycle automation, observability plane, site autonomy, automated remediation, and multi-team orchestration while driving technical standards, roadmaps, design reviews, and execution across teams.

Requirements

  • 10+ years building infrastructure-layer systems at scale
  • Deep experience with distributed systems design
  • Experience with state reconciliation, event-driven orchestration, workflow engines, and policy enforcement
  • Experience architecting platforms composed of many services
  • Hands-on fluency with bare-metal provisioning, firmware and BIOS management, GPU telemetry, and InfiniBand or RoCE fabrics
  • Experience designing large-scale observability or telemetry platforms
  • Demonstrated technical leadership across multiple teams
  • Strong software engineering fundamentals in Go, Rust, C++, or a similar systems language
  • Experience with graph data models for infrastructure topology or security attestation and SBOM tooling

Responsibilities

  • Own the end-to-end platform architecture
  • Define topology graph and infrastructure abstractions
  • Design reconciliation between intended and runtime state
  • Design policy-gated hardware lifecycle automation
  • Orchestrate provisioning, imaging, firmware upgrades, repair, and re-admission
  • Build a unified observability plane
  • Enable site autonomy across network partitions
  • Automate detection, remediation, migration, quarantine, and re-admission
  • Replace cross-team handoffs with policy-checked work orders
  • Set architecture and operating standards
  • Lead design reviews and build-versus-adopt decisions
  • Own the phased roadmap and milestones
  • Mentor senior engineers
  • Partner with hardware, network, datacenter, and product leadership

Benefits

  • Competitive compensation and equity packages
  • Restricted Stock Units
  • Paid time off
  • Paid holidays
  • Leave of absence programs
  • Comprehensive health insurance
  • Dental insurance
  • Vision insurance
  • Employer HSA contributions
  • Paid parental leave
  • Paid life insurance
  • Short-term disability insurance
  • Long-term disability insurance
  • Professional development
  • Tuition reimbursement
  • Mental health and wellness support
  • Commuter benefits
  • Cell phone stipend
  • 401(k) retirement plan with company match up to 4% of salary
  • Volunteer time off
  • Global travel insurance and emergency assistance
  • Daily meals allowance
  • Location-specific perks and programs