Senior Staff Software Engineer, AI Model Lifecycle
Crusoe is an AI infrastructure and cloud computing company. It provides GPU cloud capacity, managed AI services, inference, fine-tuning, data centers, and energy infrastructure for AI developers and enterprise customers.
Maintainer signals as of 8/23/2026
Funding history
Projects
About Crusoe, Inc
Crusoe designs, builds, and operates energy-first AI infrastructure, including data centers, GPU cloud computing, and modular AI factories. Crusoe Cloud provides GPU clusters, managed Kubernetes and Slurm, storage, networking, observability, managed inference, serverless fine-tuning, and model deployment through Crusoe Intelligence Foundry. Its customers include AI startups, enterprises, and organizations developing training, inference, analytics, and other compute-intensive workloads.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
Build a managed platform for the AI application development lifecycle, including fine-tuning systems and training pipelines for foundation models, multi-node orchestration and failure recovery, reinforcement learning and distillation workflows, and management of datasets, models, experiments, evaluations, and lineage.
Requirements
- Advanced degree in Computer Science, Engineering, or a related field
- 8+ years of industry experience leading impactful AI projects
- Experience in Generative AI, large language models, and multimodal models
- Hands-on experience training, fine-tuning, and aligning LLMs with reinforcement learning and reinforcement fine-tuning
- Proactive and collaborative approach with the ability to work autonomously
- Passion for building cutting-edge AI products and solving challenging technical problems
- Proficiency in Golang or Python
- Proficiency with PyTorch
Responsibilities
- Manage fine-tuning systems for foundation models
- Orchestrate multi-node training and checkpointing
- Implement end-to-end LLM training pipelines
- Develop reinforcement learning and reinforcement fine-tuning workflows
- Build distillation and preference or policy optimization pipelines
- Manage dataset, model, and experiment versioning
- Maintain lineage and evaluation systems
- Enable reproducible fine-tuning at scale
Benefits
- Restricted Stock Units
- Paid time off
- Paid holidays
- Comprehensive health insurance
- Dental insurance
- Vision insurance
- Employer HSA contributions
- Paid parental leave
- Paid life insurance
- Short-term disability insurance
- Long-term disability insurance
- Professional development
- Tuition reimbursement
- Mental health and wellness support
- Commuter benefits
- Cell phone stipend
- 401(k) retirement plan with company match up to 4% of salary
- Volunteer time off
