Software Engineer GPU Performance
HeyGenVisit HeyGen website
HeyGen is an AI video platform for creating avatar-led videos, generating videos from prompts, and translating video with lip sync.
Los Angeles, United States
Funding history
About HeyGen
HeyGen Technology Inc. provides AI video creation software and APIs, including AI avatars, video generation, translation/localization, and real-time avatar capabilities.
Skills
About the Role
You will improve the performance and efficiency of AI model execution and inference systems. You will profile GPU workloads, identify bottlenecks, optimize latency and throughput, develop or integrate GPU kernels, and build regression benchmarks for video workloads. You will measure the effects of changes on cost and output quality and bring reliable optimizations into production.
Requirements
- Experience optimizing GPU-based AI workloads or high-performance computing systems.
- Proficiency in Python.
- Experience with PyTorch or a similar machine learning framework.
- Knowledge of GPU hardware, including memory bandwidth, cache behavior, tensor cores, and CPU–GPU data movement.
- Experience using profiling tools to identify application-level bottlenecks and validate improvements.
- Ability to turn performance experiments into reliable production changes and communicate tradeoffs clearly.
- Experience with CUDA, Triton, or C++ GPU programming.
- Experience optimizing video, image, audio, diffusion, or Transformer models.
- Familiarity with multi-GPU inference, GPU interconnects, quantization, or large-scale model serving.
- Experience building performance benchmarks or regression testing infrastructure.
- Experience in a fast-paced technology environment.
Responsibilities
- Profile GPU utilization, kernel execution, memory bandwidth, and CPU–GPU data movement using NVIDIA Nsight Systems, Nsight Compute, and PyTorch Profiler.
- Identify bottlenecks in model execution, preprocessing, and inference serving, and measure optimization impact.
- Improve performance through batching, scheduling, memory management, and GPU utilization.
- Develop or integrate high-performance GPU kernels.
- Build benchmarks and automated checks for performance regressions across video workloads.
- Bring performance optimizations into production with AI researchers and infrastructure engineers.
- Measure changes in latency, throughput, cost, and output quality.
