Senior GPU Engineer

Vultr is an active independent cloud-infrastructure company providing global compute, GPU, bare-metal, storage, Kubernetes, and serverless AI inference services.

0 current maintainers0 active leadsTeam intelligence

Maintainer signals as of 9/25/2026

West Palm Beach, United States
About Vultr

Vultr provides globally available cloud infrastructure for developers, enterprises, and AI innovators, including Cloud Compute, Cloud GPU, Bare Metal, Cloud Storage, managed Kubernetes, and serverless inference.

View jobs by Vultr

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will lead GPU platform qualification and cluster validation, resolve system-performance bottlenecks, improve provisioning automation and reliability, set validation standards, mentor engineers, and lead cross-functional infrastructure initiatives.

Requirements

  • 5+ years of GPU infrastructure, HPC, or distributed-systems experience
  • Linux systems and server hardware expertise
  • Experience with large-scale GPU clusters
  • Strong Python programming skills
  • Experience with Ansible or similar automation frameworks
  • Experience defining validation standards and performance baselines
  • Debugging skills across hardware, operating-system, and network layers
  • Familiarity with high-speed networking concepts
  • Communication and cross-team collaboration skills

Responsibilities

  • Own end-to-end validation of GPU clusters and new hardware platforms
  • Lead hardware qualification and bring-up for new GPU platforms
  • Analyze and resolve bottlenecks across GPU, CPU, PCIe, and network layers
  • Establish validation frameworks, test suites, and performance baselines
  • Develop automation frameworks for cluster provisioning and validation
  • Define validation methodologies for GPU infrastructure
  • Troubleshoot distributed-system issues, including communication libraries such as NCCL
  • Improve reliability through proactive testing and tuning
  • Mentor engineers
  • Drive cross-functional initiatives to improve GPU cluster reliability and efficiency

Benefits

  • Company-paid medical, dental, and vision insurance premiums
  • 401(k) matching up to 4% with immediate vesting
  • 11 holidays, paid time off accrual, and PTO rollover
  • Increased PTO at three- and ten-year anniversaries
  • One-month paid sabbatical every five years
  • Annual anniversary bonus
  • Remote office setup stipend
  • Internet reimbursement up to $75 per month
  • Gym membership reimbursement up to $50 per month
  • Company-paid Wellable subscription