Staff Technical Program Manager
Crusoe is an AI infrastructure company that designs, builds, and operates AI data centers and a cloud platform. It provides managed AI services, GPU compute, model fine-tuning and inference, and infrastructure operations for organizations building and deploying AI workloads.
Maintainer signals as of 8/14/2026
Funding history
Projects
About Crusoe
Crusoe, the AI factory company, provides Crusoe Cloud and Crusoe Intelligence Foundry for AI development and production. Its offerings include managed inference, serverless fine-tuning, high-performance NVIDIA and AMD compute, accelerated storage, RDMA networking, managed Kubernetes and Slurm, and operations tooling. The company also designs, builds, and operates modular AI data-center infrastructure using an energy-first approach, serving customers that need scalable training, inference, and AI platform infrastructure.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will connect model engineering, IaaS, product, and data center operations to deliver a reliable inference platform. You will own multi-quarter program delivery, coordinate model onboarding and optimization, manage production readiness, build execution frameworks, identify risks, and align technical stakeholders.
Requirements
- 7+ years of Technical Program Manager experience
- LLM inference and model serving knowledge
- Familiarity with batching strategies and quantization approaches
- Understanding of latency, throughput, and cost trade-offs
- Multi-tenant systems experience
- Familiarity with isolation, quota management, and SLA enforcement
- Familiarity with fine-tuning and alignment workflows
- Ability to build execution models in low-structure environments
- Exceptional written and verbal communication
- Daily use of AI tools for program execution
- Ability to influence engineering, product, and infrastructure leadership without direct authority
Responsibilities
- Own multi-quarter release planning and dependency governance
- Deliver executive communications across the Managed Inference platform
- Drive model version rollouts and inference optimization campaigns
- Prepare SLA readiness for new GPU hardware
- Manage multi-tenant capacity planning
- Coordinate Model Engineering, IaaS, Cloud Foundations, Data Center Operations, and external model providers
- Identify risks across model serving, reliability, capacity, and vendor timelines
- Build TPM execution frameworks
- Maintain execution dashboards
- Deliver data-driven executive updates
- Plan model onboarding on new GPU generations
- Validate firmware, drivers, CUDA, and ROCm stacks
- Define inference commissioning criteria
- Drive alignment across technical stakeholders
Benefits
- Pension contributions
- Additional perks supporting work-life balance
