HPC Infrastructure Engineer - GPU Clusters

ElevenLabs is an AI research and product company building generative audio AI platforms for voice, music, image, video, and conversational agents.

London, United Kingdom
About ElevenLabs

Founded in 2022 by Piotr Dąbkowski and Mateusz Staniszewski, ElevenLabs develops generative AI for audio and human-technology interaction. Its current offerings include ElevenAgents for voice and chat agents, ElevenCreative for generating and editing speech, music, image, and video, and ElevenAPI for developer access to AI audio models, alongside text-to-speech, voice cloning, dubbing, voice changing, and sound-effects tools.

View jobs by ElevenLabs

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will operate and improve GPU infrastructure across provisioning scheduling monitoring upgrades and capacity planning. You will automate node health checks remediation and capacity burn-in, manage the Linux and NVIDIA software stack, tune job scheduling and storage, investigate performance issues, evaluate GPU providers, and handle hardware and security operations.

Requirements

  • Production experience running large-scale Linux server or GPU environments
  • Knowledge of NVIDIA drivers CUDA NCCL and DCGM or deep systems experience
  • Comfort with bare-metal environments server hardware and high-speed networking
  • Python or Bash automation skills
  • Experience with Ansible or Terraform
  • Ability to analyze metrics logs and PromQL data
  • Experience supporting ML training workloads is beneficial
  • Experience evaluating GPU cloud providers is beneficial
  • Experience with parallel filesystems or large-scale object storage is beneficial
  • Experience with BMC IPMI Redfish automation or PXE provisioning is beneficial
  • Awareness of power and cooling for dense GPU deployments is beneficial

Responsibilities

  • Operate and improve the GPU fleet
  • Automate node health checks draining remediation and burn-in pipelines
  • Manage OS images NVIDIA drivers CUDA container runtimes NCCL and high-speed networking
  • Run and tune Slurm or similar job scheduling
  • Build and maintain storage for datasets and checkpoints
  • Diagnose and resolve GPU cluster performance problems
  • Evaluate rented GPU capacity and provider SLAs
  • Rack cable and diagnose hardware when needed
  • Maintain access control network isolation and secrets management

Benefits

  • Annual discretionary professional development stipend
  • Annual discretionary social travel stipend
  • Annual company offsite
  • Monthly coworking stipend