Search...

HPC Specialist

DRW logo
DRW

Global principal trading firm; parent of Cumberland (crypto market maker).

Chicago, Illinois, USA
About DRW

DRW (drw.com) is a Chicago-based global principal trading firm across asset classes; its crypto arm is Cumberland, a leading digital-asset market maker.

View jobs by DRW

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will deploy, maintain, and optimize GPU infrastructure for large-scale LLM inference. You will architect multi-node model serving, manage GPU-enabled Kubernetes clusters, configure networking and storage, troubleshoot performance bottlenecks, automate infrastructure, and improve reliability through monitoring, alerting, capacity planning, and incident response.

Requirements

  • Bachelor's or Master's degree in Computer Science, Systems Engineering, or a related field
  • 5+ years in DevOps, SRE, or infrastructure engineering roles
  • Experience with GPU infrastructure, vLLM, SGLang, and GPU driver management
  • Experience optimizing deep learning workloads on GPU clusters
  • Deep Linux systems knowledge
  • Experience with Ansible, Terraform, or similar infrastructure as code tools
  • Understanding of distributed systems, TCP/IP, HTTP/2, and load balancing
  • Proficiency in Python and Bash scripting
  • Experience with Prometheus, Grafana, or similar monitoring tools

Responsibilities

  • Deploy, maintain, and optimize GPU infrastructure for LLM inference
  • Architect distributed serving solutions for multi-node, multi-GPU deployments
  • Manage GPU-enabled Kubernetes clusters
  • Configure load balancers, firewalls, and inter-node communication
  • Implement and optimize storage for model weights and inference caches
  • Troubleshoot performance bottlenecks across the technology stack
  • Evaluate GPU technologies, model serving frameworks, and infrastructure optimizations
  • Profile model performance and implement inference acceleration techniques
  • Improve reliability through monitoring, alerting, capacity planning, and incident response
HPC Specialist at DRW | JobStash