Senior Staff Applied AI Inference Engineer
Crusoe is an AI infrastructure and cloud computing company. It provides GPU cloud capacity, managed AI services, inference, fine-tuning, data centers, and energy infrastructure for AI developers and enterprise customers.
Maintainer signals as of 8/23/2026
Funding history
Projects
About Crusoe, Inc
Crusoe designs, builds, and operates energy-first AI infrastructure, including data centers, GPU cloud computing, and modular AI factories. Crusoe Cloud provides GPU clusters, managed Kubernetes and Slurm, storage, networking, observability, managed inference, serverless fine-tuning, and model deployment through Crusoe Intelligence Foundry. Its customers include AI startups, enterprises, and organizations developing training, inference, analytics, and other compute-intensive workloads.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will make large language models run faster, cheaper, and more reliably in production by owning the inference stack end to end, optimizing serving architectures and frameworks, profiling kernel-level performance, tailoring deployments to customer workloads, building production software, and taking experiments through to monitored production services.
Requirements
- Bachelor's, Master's, or Ph.D. in Computer Science, Engineering, Mathematics, or a related field
- Production coding experience in Python, C++, or another general-purpose language
- Experience optimizing large language models for high-throughput and low-latency inference
- Familiarity with vLLM or SGLang
- Experience profiling and analyzing performance at the kernel level
- Understanding of GPU architecture and behavior
- Hands-on experience with large language models
- Knowledge of AI/ML pipelines and model development and deployment
- Strong communication skills
Responsibilities
- Bring current inference techniques into production
- Design and optimize serving architectures
- Profile and analyze serving stacks and CUDA kernels
- Adapt and scale optimization methods across ML models
- Tune deployments for latency, throughput, and cost
- Tailor deployments to customer models and constraints
- Build and support production software and product features
- Develop proofs of concept and ship tested results
- Own delivery from experimentation through production
- Draft features and product requirement documents
Benefits
- Equity packages
- Restricted Stock Units
- Paid time off
- Paid holidays
- Leave of absence programs
- Comprehensive health insurance
- Dental insurance
- Vision insurance
- Employer HSA contributions
- Paid parental leave
- Paid life insurance
- Short-term disability insurance
- Long-term disability insurance
- Professional development
- Tuition reimbursement
- Mental health and wellness support
- Commuter benefits
- Cell phone stipend
- 401(k) retirement plan with company match up to 4% of salary
- Volunteer time off
- Global travel insurance
- Emergency assistance
- Daily meals allowance
- Location-specific perks and programs
