Token-Load-Aware Routing for LLM Serving
Connor Blier
Founding GTM
Token-Load-Aware Routing for LLM Serving
Token-load-aware routing for LLM serving picks a replica by combining KV cache prefix overlap with the token load already queued on each engine. Overlap cuts prefill and TTFT, but token load imbalance is what exhausts KV capacity during bursts and drives p99 latency, so both must be scored together.
TL;DR: KV overlap is a proxy for load, not the objective
Most teams route to whichever replica already holds a request's prefix, because reuse is easy to reason about. That instinct is right but incomplete. KV cache overlap is a proxy for cheap prefill; it is not the thing you are actually trying to protect. The thing you are protecting is tail latency, and tail latency dies when token load piles up unevenly across your fleet. Anyscale frames the tradeoff plainly: the challenge is "how to take advantage of KV cache reuse while balancing heterogeneity across engine replicas."
Token-load-aware routing reconciles the two. It scores each candidate replica on prefix overlap and on the token load it is already carrying, then routes to the one that wins on the combined objective rather than on overlap alone.
What token load means on a replica
A replica's token load is the sum of two very different kinds of work. Prefill processes the input prompt in one compute-heavy pass and produces the first token. Decode generates each subsequent token one step at a time, and its throughput scales with how many requests share the batch - as vLLM puts it, "more requests a decode engine serves at once, higher throughput we get."
Those two phases compete for the same GPU. A burst of long-prompt requests floods prefill; a pile of long-generation requests keeps decode slots occupied for minutes. A router that only counts requests-in-flight sees both as "one request" and mis-estimates the real load. Token load is the honest unit because it tracks the work, not the request count.
The tension: overlap vs. imbalance
Routing purely by KV overlap concentrates traffic. If replica A holds the popular system prompt, every matching request wants A - and "routing a request to a replica that already holds its prefix can significantly reduce prefill computation and TTFT." That is real savings. But it is also a hotspot.
When a burst arrives, that concentration bites. Together.ai's account of a coding-agent workload describes it exactly: "prefill capacity and KV cache headroom get exhausted exactly during the burst, which is when it matters most, producing exactly the kind of multi-minute request queuing that agentic coding workflows can't tolerate." The overlap decision that looked optimal request-by-request produced a queue that blew up p99.
This matters most for the traffic shape agents actually generate: "a peak-load, relatively-low-TPS pattern where bursts in concurrency matter far more than raw tokens-per-second." Average throughput can look healthy while the tail is on fire.
How token-load-aware routers decide
The orchestration layer's decisions are not cosmetic - "how effectively the orchestration layer distributes heterogeneous request streams across a fleet of LLM engine replicas directly impacts serving TTFT, TPOT, and throughput." A token-load-aware router turns that into a two-term score.
Overlap term. Estimate how much of the request's prefix a replica already holds. Modern systems compute this at fine granularity: "vLLM emits events as cache blocks are created and removed, giving the router an up-to-date view of KV cache state across replicas and allowing it to compute overlap at the KV cache block level for each request."
Load term. Estimate the token work already committed on that replica - queued prefill tokens plus active decode slots.
The router then routes to the replica that maximizes overlap benefit net of the queuing cost the extra load would add. When a high-overlap replica is saturated, it accepts a colder cache elsewhere to keep the tail flat. Overlap is a discount, not a mandate.
Concurrency and KV cache capacity constraints
The reason token load and cache reuse are the same conversation is that they share one budget. "What caps concurrency is KV cache capacity, since every active request holds a slice of it." Every request you admit for its cheap-prefill overlap also consumes cache that a concurrent decode needs. Push overlap too hard and you starve concurrency; the replica can no longer keep enough requests batched to sustain throughput, and both TTFT and TPOT degrade together.
A good router therefore treats KV capacity as the constraint and overlap as one lever inside it - never as a goal that overrides headroom. If prefix reuse is central to your workload, the mechanics are worth understanding on their own; see our deeper write-ups on prefix caching for multi-turn LLM agents and speculative decoding throughput.
Why the underlying latency numbers matter
Routing only pays off if each replica is fast to begin with. In our benchmarks of framework choice on a single H100 - 256 input tokens, batch size 1, full roundtrip to the platform - we measured a lowest TTFT of 123ms with vLLM, and 460 tokens/sec throughput with SGLang at batch size 64. Framework choice sets the floor your router then defends. The full method and per-framework curves are in our vLLM vs SGLang vs TensorRT benchmark.
Capacity elasticity matters too: a replica that cold-starts slowly can't absorb a burst. In our testing of a TensorRT-LLM Llama 3 8B engine, cold start takes "roughly ~10-15s to load the model into GPU memory," and that same single-A10 setup reached ~1700 output tokens/sec (FP8). If your autoscaler adds replicas to shed token load, that cold-start window is the gap your router has to cover with the fleet it already has. For the cost side of running this at volume, see LLM inference cost at scale and the 2026 GPU buyer's guide.
Checklist for Lead ML Engineers
Measure token load, not request count. Track queued prefill tokens and active decode slots per replica.
Score overlap and load together. Treat prefix reuse as a discount on prefill cost, subtracted from queuing cost - not as the routing objective.
Watch p99, not the average. Bursty, low-TPS agent traffic hides behind healthy mean throughput.
Respect KV capacity as the hard limit. Concurrency is capped by cache; over-concentrating overlap starves it.
Account for cold starts. Know how long a new replica takes before it can take load off a hotspot.
Workloads that are heavy on concurrent sessions - AI agents and voice especially - feel this tension first. Start from the tail latency you must hold, then let overlap earn its place inside that budget.
Frequently asked questions
- Is KV cache overlap a bad routing signal?
- No. Overlap is a strong signal because reusing a held prefix significantly reduces prefill computation and TTFT. The mistake is treating it as the objective. It's a proxy for cheap prefill; the real objective is protecting tail latency, which requires also scoring the token load already on each replica.
- Why does token load imbalance hurt p99 more than average latency?
- Agent traffic tends to be a peak-load, low-TPS pattern where bursts in concurrency matter more than raw tokens-per-second. During a burst, an overloaded replica exhausts prefill capacity and KV headroom and queues requests for minutes, spiking p99 while the average stays deceptively healthy.
- What actually caps concurrency on an LLM replica?
- KV cache capacity. Every active request holds a slice of the cache, so once it's full the replica can't batch more requests, and both throughput and latency degrade. That's why a router must treat KV capacity as the constraint, not chase overlap past the point of headroom.
- How does a router know the KV cache state across replicas?
- Engines like vLLM emit events as cache blocks are created and removed, giving the router an up-to-date, block-level view of cache state so it can compute prefix overlap per request in real time.
Get started with Cerebrium
Deploy AI models on serverless GPUs in minutes, with no infrastructure to manage. Start for free and pay only for the compute you use.
Sources
- anyscale.com
“Together, these properties create a unique challenge for the orchestration layer: how to take advantage of KV cache reuse while balancing heterogeneity across engine replicas.”
Frames the core overlap-vs-load tension the article reconciles.
- anyscale.com
“Routing a request to a replica that already holds its prefix can significantly reduce prefill computation and TTFT.”
Supports the benefit of overlap-based routing.
- anyscale.com
“How effectively the orchestration layer distributes heterogeneous request streams across a fleet of LLM engine replicas directly impacts serving TTFT, TPOT, and throughput.”
Establishes that routing decisions materially affect serving metrics.
- anyscale.com
“vLLM emits events as cache blocks are created and removed, giving the router an up-to-date view of KV cache state across replicas and allowing it to compute overlap at the KV cache block level for each request.”
Explains how the router obtains block-level overlap data.
- together.ai
“prefill capacity and KV cache headroom get exhausted exactly during the burst, which is when it matters most, producing exactly the kind of multi-minute request queuing that agentic coding workflows can't tolerate.”
Illustrates the queuing failure caused by token load imbalance under burst.
- together.ai
“a peak-load, relatively-low-TPS pattern where bursts in concurrency matter far more than raw tokens-per-second.”
Characterizes the bursty agent traffic shape where p99 matters most.
- vllm.ai
“What caps concurrency is KV cache capacity, since every active request holds a slice of it.”
Grounds the KV-capacity-as-constraint section.
- vllm.ai
“This metric depends on concurrency: more requests a decode engine serves at once, higher throughput we get.”
Explains decode throughput scaling with concurrency.
- cerebrium.ai
“From our tests, it seems that vLLM is the best framework for use cases looking for the lowest TTFT of 123ms which is extremely quick if you take into consideration network latency.”
First-hand Cerebrium TTFT measurement.
- cerebrium.ai
“The following benchmark was done with 256 input tokens on 1xH100 of batch size 1. This was the total roundtrip time of a request made to the Cerebrium platform.”
Methodology for the first-hand benchmark numbers.
- cerebrium.ai
“As for throughput use cases, it seems SGLang was a clear winner achieving a throughput of 460 tokens per second on a batch size of 64.”
First-hand Cerebrium throughput measurement.
- cerebrium.ai
“This code will run on every cold start and takes roughly ~10-15s to load the model into GPU memory.”
First-hand cold-start figure relevant to burst absorption.
- cerebrium.ai
“In this tutorial we will achieve ~1700 output tokens per second (FP8)on a single Nvidia A10 instance however you can go up to ~4500 output tokens per second on a single Nvidia A100 40GB instance or even ~19,000 tokens on a H100.”
First-hand single-A10 throughput figure (A10 only).