AI Inference Platform Engineer

1 day agoSalary: 200K - 250KChicagoAiJobs by DRW

DRW is a Chicago-based diversified trading firm that uses technology, quantitative research, and risk management to provide liquidity across traditional and digital-asset markets.

Chicago, United States
About DRW

Founded by Don Wilson in 1992, DRW trades for its own account across global markets and operates strategies in cryptoassets, venture capital, real estate, carbon markets, and public-equity investments. Its crypto business, Cumberland, has provided institutional crypto-asset liquidity since 2014.

View jobs by DRW

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will build, operate, and optimize a production AI inference platform. You will onboard and configure new models, profile performance across GPUs and distributed systems, improve reliability and cost efficiency, manage model serving lifecycles, and operate multi-tenant GPU scheduling and isolation.

Requirements

  • Hands-on experience serving LLMs on NVIDIA GPUs.
  • Deep expertise in TensorRT-LLM, vLLM, SGLang, or another modern inference runtime.
  • Knowledge of inference optimization techniques including continuous batching, scheduling, chunked prefill, speculative decoding, quantization, CUDA Graphs, and paged attention.
  • Understanding of KV cache architecture.
  • Experience measuring model quality equivalence using evaluation harnesses, benchmarks, and regression detection.
  • Experience designing and tuning distributed inference systems.
  • Experience with multi-tenant GPU scheduling, workload isolation, and QoS.
  • Proficiency with Nsight, DCGM, OpenTelemetry, Prometheus, and Grafana.
  • Strong Linux and systems performance fundamentals.
  • Production experience with model serving infrastructure, CI/CD, automated testing, observability, and production readiness.

Responsibilities

  • Optimize LLM inference performance across NVIDIA GPU architectures and inference runtimes.
  • Build performance profiling and observability across GPU kernels and multi-node inference systems.
  • Design and optimize KV cache and distributed inference architectures.
  • Own day-zero model onboarding and select serving configurations.
  • Maintain validated performance profiles and quality regression testing.
  • Measure and monitor quality equivalence across serving configurations.
  • Manage model versioning, compatibility, staging, canarying, promotion, rollback, and retirement.
  • Automate model deployment, distribution, production readiness, observability, and reliable operation.
  • Optimize model placement, scaling, and resource allocation across the inference fleet.
  • Design and operate multi-tenant GPU scheduling and workload isolation.

Benefits

  • Annual discretionary bonus
  • Group medical, pharmacy, dental, and vision insurance
  • 401k with discretionary employer match
  • Short-term and long-term disability insurance
  • Life and AD&D insurance
  • Health savings accounts
  • Flexible spending accounts