Senior ML Systems Engineer Inference

Runpod is an AI developer cloud providing GPU compute, serverless inference, Pods, and clusters for building, training, fine-tuning, deploying, and scaling AI workloads.

Recently funded0 current maintainers0 active leadsTeam intelligence

Maintainer signals as of 9/25/2026

San Francisco, United States
About Runpod

Runpod Inc. operates a globally distributed GPU cloud platform for AI developers, offering on-demand GPU infrastructure and serverless services across the AI development lifecycle.

View jobs by Runpod

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will lead end-to-end LLM serving performance work. You will define rigorous performance measurements, diagnose bottlenecks across the serving stack, improve efficiency for large models on GPU deployments, and turn findings into reliable production runtimes and configurations. You will collaborate on inference offerings, evaluate emerging ecosystem tools, and implement runtime fixes.

Requirements

  • 5+ years of professional system engineering experience
  • Production or serious benchmark-scale experience with vLLM, SGLang, or a comparable serving engine
  • Strong Python software engineering skills
  • Understanding of LLM inference performance, batching, memory, parallelism, latency, and throughput
  • Experience with quantization, speculative decoding, or distributed serving
  • Benchmarking, performance analysis, and GPU profiling skills
  • Clear written communication of results and decisions

Responsibilities

  • Define rigorous, repeatable inference-performance measurements
  • Profile and diagnose serving-stack performance problems
  • Improve serving efficiency for large models on single-node and multi-node GPU deployments
  • Create production-ready runtimes, configurations, and defaults
  • Shape inference offerings with product and infrastructure stakeholders
  • Evaluate, adopt, build, and contribute to inference ecosystem projects
  • Trace serving-engine bottlenecks and implement fixes

Benefits

  • Equity through stock options
  • Medical, dental, and vision plans
  • Flexible PTO
  • Remote-first work
  • $1,200 home office and equipment stipend