Search...

Senior Staff Applied AI Inference Engineer

Crusoe logo
Crusoe

Crusoe is an AI infrastructure company that designs, builds, and operates AI data centers and a cloud platform. It provides managed AI services, GPU compute, model fine-tuning and inference, and infrastructure operations for organizations building and deploying AI workloads.

Maintainer signals as of 8/14/2026

Distributed
About Crusoe

Crusoe, the AI factory company, provides Crusoe Cloud and Crusoe Intelligence Foundry for AI development and production. Its offerings include managed inference, serverless fine-tuning, high-performance NVIDIA and AMD compute, accelerated storage, RDMA networking, managed Kubernetes and Slurm, and operations tooling. The company also designs, builds, and operates modular AI data-center infrastructure using an energy-first approach, serving customers that need scalable training, inference, and AI platform infrastructure.

View jobs by Crusoe

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will make large language models run faster, cheaper, and more reliably in production. You will own the inference stack end to end, optimize serving architectures and frameworks, profile performance at the kernel level, tailor deployments to customer workloads, build production software, and take experiments through to monitored production services.

Requirements

  • Bachelor's, Master's, or Ph.D. in Computer Science, Engineering, Mathematics, or a related field
  • Production coding experience in Python, C++, or another general-purpose language
  • Experience optimizing large language models for high-throughput and low-latency inference
  • Familiarity with vLLM or SGLang
  • Experience profiling and analyzing performance at the kernel level
  • Understanding of GPU architecture and behavior
  • Hands-on experience with large language models
  • Knowledge of AI/ML pipelines and model development and deployment
  • Strong communication skills

Responsibilities

  • Bring current inference techniques into production
  • Design and optimize serving architectures
  • Profile and analyze serving stacks and CUDA kernels
  • Adapt and scale optimization methods across ML models
  • Tune deployments for latency, throughput, and cost
  • Tailor deployments to customer models and constraints
  • Build and support production software and product features
  • Develop proofs of concept and ship tested results
  • Own delivery from experimentation through production
  • Draft features and product requirement documents

Benefits

  • Equity packages
  • Restricted Stock Units
  • Paid time off
  • Paid holidays
  • Leave of absence programs
  • Comprehensive health insurance
  • Dental insurance
  • Vision insurance
  • Employer HSA contributions
  • Paid parental leave
  • Paid life insurance
  • Short-term disability insurance
  • Long-term disability insurance
  • Professional development
  • Tuition reimbursement
  • Mental health and wellness support
  • Commuter benefits
  • Cell phone stipend
  • 401(k) retirement plan with company match up to 4% of salary
  • Volunteer time off
  • Global travel insurance
  • Emergency assistance
  • Daily meals allowance
  • Location-specific perks and programs