Search...

Senior Software Engineer, AI Model Lifecycle

Crusoe logo
Crusoe

Crusoe is an AI infrastructure company that designs, builds, and operates AI data centers and a cloud platform. It provides managed AI services, GPU compute, model fine-tuning and inference, and infrastructure operations for organizations building and deploying AI workloads.

Series C13 current maintainers10 active leads5 new active leads9 lead step-downs1 early lead departureTeam intelligence

Maintainer signals as of 8/12/2026

Distributed
About Crusoe

Crusoe, the AI factory company, provides Crusoe Cloud and Crusoe Intelligence Foundry for AI development and production. Its offerings include managed inference, serverless fine-tuning, high-performance NVIDIA and AMD compute, accelerated storage, RDMA networking, managed Kubernetes and Slurm, and operations tooling. The company also designs, builds, and operates modular AI data-center infrastructure using an energy-first approach, serving customers that need scalable training, inference, and AI platform infrastructure.

View jobs by Crusoe

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You build a managed platform for the application development lifecycle with a focus on machine learning models and large language models. You manage fine-tuning systems, training pipelines, reinforcement learning workflows, and dataset, model, and experiment management at scale.

Requirements

  • Advanced degree in Computer Science, Engineering, or a related field
  • 4-5+ years of industry experience leading impactful AI projects
  • Generative AI experience
  • Large language model experience
  • Multimodal model experience
  • Experience training, fine-tuning, and aligning large language models
  • Reinforcement learning experience
  • Reinforcement fine-tuning experience
  • Autonomous work
  • Collaboration skills
  • Proficiency in Golang or Python
  • PyTorch experience
  • Open-source AI project contributions
  • GPU performance optimization experience
  • Inference framework experience

Responsibilities

  • Manage fine-tuning systems for large foundation models, including multi-node orchestration, checkpointing, failure recovery, and cost-efficient scaling
  • Implement and maintain end-to-end training pipelines for large language models
  • Apply reinforcement fine-tuning and reinforcement learning to fine-tuning and training workflows
  • Develop distillation and reinforcement learning pipelines for preference optimization, policy optimization, and reward modeling
  • Manage dataset, model, and experiment versioning, lineage, evaluation, and reproducible fine-tuning at scale

Benefits

  • Restricted Stock Units
  • Paid time off
  • Paid holidays
  • Health insurance
  • Dental insurance
  • Vision insurance
  • Employer contributions to HSA account
  • Paid parental leave
  • Paid life insurance
  • Short-term disability insurance
  • Long-term disability insurance
  • Tuition reimbursement
  • Mental health and wellness support
  • Commuter benefits
  • Cell phone stipend
  • 401(k) retirement plan with company match up to 4% of salary
  • Volunteer time off