vLLM Server Parameters for Inference Throughput

Connor Blier
Founding GTM

A single speedometer-style gauge inside a circular dial, with a needle pointing to the upper-right, symbolizing tuning settings for maximum throughput.

vLLM Server Parameters for Inference Throughput

Tune vLLM server parameters for inference throughput by treating the GPU as one fixed memory budget shared by weights, activation memory and KV cache. Set gpu_memory_utilization first, size the KV cache from what remains, push max-num-seqs as high as that cache allows, then validate with warmup and request ordering.

The one-budget mental model

The single most useful way to think about vLLM server parameters for inference throughput is as a memory accounting problem, not a list of flags. A GPU exposes one fixed pool of VRAM, and three claimants compete for it: the model weights, the activation (and non-Torch) memory needed to run a forward pass, and the KV cache that holds attention state for in-flight requests. Throughput is overwhelmingly a function of how many tokens and how many concurrent sequences you can keep resident in that third claimant - so every parameter decision is really a decision about how much of the budget the KV cache gets to keep.

The vLLM community makes the same point directly: because "system performance is affected by the available GPU memory for inference, and the available memory for inference is in turn constrained by the server-side parameters," choosing the right parameters is the core tuning problem (discuss.vllm.ai). The order in which you set them matters, because each one changes how much budget the next one has to work with.

Before you tune anything, measure the baseline on the platform you will actually bill against. Vendor spec sheets are not your round-trip latency - in our benchmarks of vLLM, SGLang and TensorRT for a Llama 3.1 API, we measured vLLM's lowest TTFT at 123ms and a peak of 460 tokens per second (SGLang, batch size 64) on a single H100, all as total round-trip time to the Cerebrium platform rather than isolated kernel timings. Establish your own number the same way before you start moving flags.

Step 1 - set the GPU memory budget (gpu_memory_utilization)

gpu_memory_utilization is the outer envelope: it tells vLLM what fraction of total VRAM it may claim. Everything else is carved out of this. Push it as high as you safely can - more of the envelope means more room for KV cache - but leave headroom, because the measurements that follow (non-Torch buffers in particular) live outside Torch's own allocator.

The quantities you are implicitly dividing up here are concrete and measurable: "Model weight memory | GPU memory occupied by model weights" and "Non_torch memory | Memory not allocated by Torch (NCCL, low-level driver buffers, etc.)" (discuss.vllm.ai). Subtract both from your utilization envelope and what is left is the space available to the KV cache. Quantization shrinks the weight claimant and hands that space straight to the cache, which is why FP8 deployments reach throughput a higher-precision build cannot.

Step 2 - size the KV cache

With weights and non-Torch overhead subtracted, the remainder is your KV cache. vLLM reports this back to you, and it is the number to watch: the "Available KV cache memory | Available KV cache size" and the "GPU KV cache size | Number of tokens that can be accommodated under the current available KV cache" (discuss.vllm.ai). That token count is the real throughput ceiling: it is the total number of tokens, summed across every concurrent request, that the server can hold at once.

If the reported token capacity is lower than your workload needs, the fix is upstream - raise gpu_memory_utilization, quantize the weights, or move to a larger-memory GPU. These trade-offs are what our H100 vs H200 throughput guide and the 2026 GPU buyer's guide exist to help with, and for mixture-of-experts models the weight footprint math is covered in our MoE VRAM footprint breakdown.

Step 3 - push concurrency (max-num-seqs)

max-num-seqs caps how many sequences vLLM batches together per step. Higher concurrency is where throughput comes from - but it is bounded by the KV cache you sized in Step 2, because every concurrent sequence needs its own slice of cache for its context. Set it too high and requests either queue or trigger preemption; set it too low and the GPU sits underfed.

The principled way to pick the value is to relate concurrency to the activation memory each additional sequence costs. vLLM surfaces "Peak Activation memory | Peak activation memory usage," which is "a critical measurement for computing the conversion coefficient between activation memory and max-num-seqs" (discuss.vllm.ai). In practice: raise max-num-seqs until either the KV token budget or peak activation memory is the binding constraint, then back off to leave headroom.

Step 4 - account for activation memory

Activation memory is the quiet third claimant. It scales with batch size and sequence length, and because it peaks mid-forward-pass it is the parameter most likely to produce an intermittent OOM that a short test never reveals. Measure peak activation memory at your intended max-num-seqs and longest expected input, and treat that peak - not the average - as the amount you must reserve inside the gpu_memory_utilization envelope. If the peak eats into the KV cache you planned, lower max-num-seqs rather than discovering the ceiling in production.

Step 5 - order your requests

Once the memory budget is maximized, request scheduling is the next lever. How you order incoming requests changes how effectively the KV cache is reused and evicted: as the Quail work put it, "with a structured query in hand, you can order requests to better cache (and evict) KV" - a technique they credit for throughput more than 10x a vLLM baseline on H100 (modal.com). Grouping requests that share a prefix lets the cache serve the shared tokens once instead of per request. For multi-turn agents this compounds; see our deep dives on prefix caching for multi-turn agents and token-load-aware routing.

Step 6 - warm up before taking traffic

None of the above is real until the server is warm. vLLM's startup does substantial work before it can serve: the initialization path "includes importing Python modules, loading PyTorch, assembling model weights, copying them onto the GPU, and running the framework's warmup path - torch.compile, CUDA graph capture, KV cache initialization" (cerebrium.ai). Benchmark only after KV cache initialization and CUDA graph capture have completed, or your throughput numbers will blend cold-start cost with steady-state. On serverless this matters doubly - the cold-start work repeats on every scale-up event.

Worked tuning order (checklist)

  1. Baseline. Measure throughput and TTFT end to end, on your target platform, warm.

  2. gpu_memory_utilization. Raise toward the safe maximum; subtract measured model-weight and non-Torch memory.

  3. KV cache. Read back available KV cache size and token capacity - this is your throughput ceiling.

  4. max-num-seqs. Raise until KV token budget or peak activation memory binds; back off for headroom.

  5. Activation memory. Reserve the measured peak, not the average, at your longest input.

  6. Request ordering. Group shared-prefix requests to improve cache reuse.

  7. Warmup. Complete torch.compile, CUDA graph capture and KV cache init before measuring or serving.

  8. Re-measure and iterate - every change in one step shifts the budget for the others.

If you are deciding between serving engines before you tune at all, start from our vLLM, SGLang and TensorRT benchmark, and if you need an OpenAI-compatible endpoint for open-source LLMs the same budget logic applies underneath it.

Frequently asked questions

Which vLLM parameter should I set first for throughput?
Set gpu_memory_utilization first. It defines the total VRAM envelope that weights, activation memory and KV cache are all carved from, so every later decision depends on it. Raise it toward the safe maximum, then subtract measured model-weight and non-Torch memory to see what the KV cache can have.
How does KV cache size limit throughput?
vLLM reports the available KV cache size and the number of tokens it can accommodate. That token count - summed across all concurrent requests - is the real throughput ceiling. If it is too low, raise gpu_memory_utilization, quantize the weights, or move to a larger-memory GPU.
How do I choose max-num-seqs?
Relate it to peak activation memory per sequence and to KV token capacity. Raise max-num-seqs until one of those becomes the binding constraint, then back off for headroom. Too high causes queuing or preemption; too low leaves the GPU underfed.
Does request ordering really affect throughput?
Yes. Ordering requests so that shared prefixes land together lets the KV cache be reused and evicted efficiently. The Quail project credits structured request ordering for throughput more than 10x a vLLM baseline on H100.
Why warm up before benchmarking?
vLLM's startup runs torch.compile, CUDA graph capture and KV cache initialization before it can serve. Measuring before that completes blends cold-start cost into your steady-state throughput and gives misleading numbers.

Get started with Cerebrium

Deploy AI models on serverless GPUs in minutes, with no infrastructure to manage. Start for free and pay only for the compute you use.

Sign up free

Sources

  1. discuss.vllm.ai
    “Since system performance is affected by the available GPU memory for inference, and the available memory for inference is in turn constrained by the server-side parameters of the VLLM framework, choosing appropriate server-side parameters during performance testing is a relatively important issue.”

    Establishes GPU memory as the primary constraint that server-side parameters govern.

  2. discuss.vllm.ai
    “Available KV cache memory | Available KV cache size GPU KV cache size | Number of tokens that can be accommodated under the current available KV cache”

    The KV cache metrics vLLM reports back, used as the throughput ceiling.

  3. discuss.vllm.ai
    “Peak Activation memory | Peak activation memory usage”

    Peak activation memory as the measurement tying max-num-seqs to memory.

  4. discuss.vllm.ai
    “Model weight memory | GPU memory occupied by model weights Non_torch memory | Memory not allocated by Torch (NCCL, low-level driver buffers, etc.)”

    Weight and non-Torch memory subtracted from the budget before KV cache sizing.

  5. modal.com
    “Spoilers: the big win is that with a structured query in hand, you can order requests to better cache (and evict) KV.”

    Request ordering improves KV cache reuse and throughput.

  6. cerebrium.ai
    “That initialization path includes importing Python modules, loading PyTorch, assembling model weights, copying them onto the GPU, and running the framework's warmup path - torch.compile, CUDA graph capture, KV cache initialization, and whatever else the serving stack needs before it can take traffic.”

    Warmup path that must complete before benchmarking or serving.

  7. cerebrium.ai
    “From our tests, it seems that vLLM is the best framework for use cases looking for the lowest TTFT of 123ms which is extremely quick if you take into consideration network latency.”

    Cerebrium-measured vLLM TTFT baseline figure.

  8. cerebrium.ai
    “As for throughput use cases, it seems SGLang was a clear winner achieving a throughput of 460 tokens per second on a batch size of 64.”

    Cerebrium-measured peak throughput on a single H100.


Related resources

See all
A shield containing a padlock, symbolising dedicated, secured GPU capacity reserved for production workloads.
Reserved GPU Capacity for Production LLM Inference
A single microchip rendered as a square outline with pins on all four sides, containing two short vertical bars inside representing a compact, low-bit numeric value.
4-Bit LLM Inference on H100: Throughput, VRAM & Cost
A large four-pointed sparkle star at the centre of the frame, surrounded by two smaller sparkle stars of the same shape near its lower corners, all rendered as solid black shapes suggesting bursts of instant activation.
Modal vs RunPod: LLM Inference Cost & Cold Starts