Staff Technical Program Manager

Crusoe is an AI infrastructure and cloud computing company. It provides GPU compute, AI model training and inference, developer tools, managed orchestration, data centers, and energy infrastructure for AI builders and enterprise customers.

Maintainer signals as of 8/23/2026

Distributed
About Crusoe, Inc

Crusoe designs, builds, and operates energy-first AI infrastructure, including high-performance data centers, modular AI factories, and Crusoe Cloud. Its cloud platform provides NVIDIA and AMD GPU compute, scalable storage, high-performance networking, managed Kubernetes and Slurm, observability, model training, fine-tuning, and inference services. Crusoe serves AI developers, startups, enterprises, and teams running production-scale AI workloads.

View jobs by Crusoe, Inc

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

Connect model engineering, IaaS, product, and data center operations to deliver a reliable inference platform. Own multi-quarter program delivery, coordinate model onboarding and optimization, manage production readiness, build execution frameworks, identify risks, and align technical stakeholders.

Requirements

  • 7+ years of Technical Program Manager experience
  • LLM inference and model serving knowledge
  • Familiarity with batching strategies and quantization approaches
  • Understanding of latency, throughput, and cost trade-offs
  • Multi-tenant systems experience
  • Familiarity with isolation, quota management, and SLA enforcement
  • Familiarity with fine-tuning and alignment workflows
  • Ability to build execution models in low-structure environments
  • Exceptional written and verbal communication
  • Daily use of AI tools for program execution
  • Ability to influence engineering, product, and infrastructure leadership without direct authority

Responsibilities

  • Own multi-quarter release planning and dependency governance
  • Deliver executive communications across the Managed Inference platform
  • Drive model version rollouts and inference optimization campaigns
  • Prepare SLA readiness for new GPU hardware
  • Manage multi-tenant capacity planning
  • Coordinate Model Engineering, IaaS, Cloud Foundations, Data Center Operations, and external model providers
  • Identify risks across model serving, reliability, capacity, and vendor timelines
  • Build TPM execution frameworks
  • Maintain execution dashboards
  • Deliver data-driven executive updates
  • Plan model onboarding on new GPU generations
  • Validate firmware, drivers, CUDA, and ROCm stacks
  • Define inference commissioning criteria
  • Drive alignment across technical stakeholders

Benefits

  • Pension contributions
  • Additional perks supporting work-life balance