HPC Specialist
DRW is a Chicago-based diversified trading firm that uses technology, quantitative research, and risk management to provide liquidity across traditional and digital-asset markets.
About DRW
Founded by Don Wilson in 1992, DRW trades for its own account across global markets and operates strategies in cryptoassets, venture capital, real estate, carbon markets, and public-equity investments. Its crypto business, Cumberland, has provided institutional crypto-asset liquidity since 2014.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will deploy, maintain, and optimize GPU infrastructure for large-scale LLM inference. You will architect multi-node model serving, manage GPU-enabled Kubernetes clusters, configure networking and storage, troubleshoot performance bottlenecks, automate infrastructure, and improve reliability through monitoring, alerting, capacity planning, and incident response.
Requirements
- Bachelor's or Master's degree in Computer Science, Systems Engineering, or related field.
- 5+ years in DevOps, SRE, or infrastructure engineering roles.
- Strong experience with GPU infrastructure, model serving frameworks (vLLM, SGLang), and GPU driver management.
- Hands-on experience optimizing deep learning workloads (inference or training) on GPU clusters.
- Deep Linux systems knowledge including network configuration, storage optimization, and Kubernetes orchestration.
- Experience with infrastructure as code tools (Ansible, Terraform, or similar).
- Strong understanding of distributed systems, networking protocols (TCP/IP, HTTP/2), and load balancing.
- Proficiency in Python and Bash scripting for automation.
- Experience with monitoring and observability tools (Prometheus, Grafana, or similar).
Responsibilities
- Deploy, maintain, and optimize GPU infrastructure for large-scale LLM inference workloads, including provisioning, configuration, and deployment of GPU server fleets.
- Architect and implement distributed serving solutions for multi-node, multi-GPU model deployments.
- Manage GPU-enabled Kubernetes clusters for LLM and ML workloads.
- Configure network infrastructure including load balancers, firewalls, and inter-node communication for GPU clusters.
- Implement and optimize storage solutions for model weights and inference caches.
- Troubleshoot performance bottlenecks across the stack: hardware, drivers, networking, and application layer.
- Research and evaluate emerging GPU technologies, model serving frameworks, and infrastructure optimizations.
- Collaborate with ML engineers to profile model performance and implement inference acceleration techniques.
- Drive reliability improvements through monitoring, alerting, capacity planning, and incident response.
