Senior Software Engineer, AI Model Lifecycle

Crusoe is an AI infrastructure and cloud computing company. It provides GPU cloud capacity, managed AI services, inference, fine-tuning, data centers, and energy infrastructure for AI developers and enterprise customers.

Maintainer signals as of 8/23/2026

Distributed
About Crusoe, Inc

Crusoe designs, builds, and operates energy-first AI infrastructure, including data centers, GPU cloud computing, and modular AI factories. Crusoe Cloud provides GPU clusters, managed Kubernetes and Slurm, storage, networking, observability, managed inference, serverless fine-tuning, and model deployment through Crusoe Intelligence Foundry. Its customers include AI startups, enterprises, and organizations developing training, inference, analytics, and other compute-intensive workloads.

View jobs by Crusoe, Inc

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

Build a managed platform for the application development lifecycle focused on machine learning models and large language models. Manage fine-tuning systems, training pipelines, reinforcement learning workflows, and dataset, model, and experiment management at scale.

Requirements

  • Advanced degree in Computer Science, Engineering, or a related field.
  • 4-5+ years of industry experience leading impactful AI projects.
  • Generative AI experience, including large language and multimodal models.
  • Experience training, fine-tuning, and aligning large language models using reinforcement learning and reinforcement fine-tuning techniques.
  • Ability to work autonomously and collaboratively.
  • Proficiency in Golang or Python and PyTorch.
  • Open-source AI project contributions.
  • GPU performance optimization and inference framework experience.

Responsibilities

  • Manage fine-tuning systems for large foundation models, including multi-node orchestration, checkpointing, failure recovery, and cost-efficient scaling.
  • Implement and maintain end-to-end training pipelines for large language models.
  • Apply reinforcement fine-tuning and reinforcement learning to fine-tuning and training workflows.
  • Develop distillation and reinforcement learning pipelines for preference optimization, policy optimization, and reward modeling.
  • Manage dataset, model, and experiment versioning, lineage, evaluation, and reproducible fine-tuning at scale.

Benefits

  • Restricted Stock Units
  • Paid time off
  • Paid holidays
  • Health insurance
  • Dental insurance
  • Vision insurance
  • Employer contributions to HSA account
  • Paid parental leave
  • Paid life insurance
  • Short-term disability insurance
  • Long-term disability insurance
  • Tuition reimbursement
  • Mental health and wellness support
  • Commuter benefits
  • Cell phone stipend
  • 401(k) retirement plan with company match up to 4% of salary
  • Volunteer time off