Disaggregated Prefill/Decode LLM Serving: A Guide
Connor Blier
Founding GTM
Disaggregated Prefill/Decode LLM Serving: A Guide
Disaggregated prefill/decode LLM serving splits inference into two phases that run on separate GPU pods: compute-bound prefill on one set, memory-bound decode on another. Splitting them stops the phases contending for the same hardware, but it adds KV cache handoff and KV-aware routing you must engineer for.
LLM inference is not one workload. It is two, and they fight over the same GPU. Disaggregated prefill/decode serving is the architecture that pulls them apart onto separate hardware so each phase runs on the resource it actually needs. This guide explains what that split buys you, the constraints it introduces, and when it isn't worth the complexity.
What is disaggregated prefill/decode serving
As Modular puts it, "LLM Inference has two phases, and they stress hardware differently." Prefill reads your whole prompt and builds the KV cache; decode then generates output tokens one at a time. In a disaggregated setup those two phases no longer share a GPU. Instead, "some pods might specialize in prefill, others in decode," and a request flows from a prefill pod to a decode pod with the KV cache handed off in between.
The payoff is that each pool can be sized, batched and scaled for its own bottleneck rather than for an uneasy average of both.
Why prefill and decode contend for the same GPU
Prefill is compute-heavy. "Prefill processes the entire prompt in parallel. It's compute-bound. This means GPU cores are saturated doing dense [matrix multiplications]." Decode is the opposite: it is dominated by reading the KV cache, and its concurrency is bounded by memory. Per vLLM, "KV cache capacity" determines "request concurrency since every active request holds a slice of it."
Run both on one GPU and they interfere. A long prompt saturating the cores stalls the token-by-token decode of other users, and a large batch of decoders eats the memory a fresh prefill needs. Splitting the phases removes that cross-interference. If you want to squeeze more out of a single co-located GPU first, speculative decoding on serverless GPUs is a lighter-weight lever to try before you disaggregate.
PD topologies and pod specialization
The defining move of disaggregation is pod specialization: prefill-only pools and decode-only pools. These are not interchangeable web servers. As Modular notes, the backends "are GPU pods holding large, local KV caches in high-bandwidth, RAM or SSD memory," and "a single inference call sometimes needs two backends in sequence."
That statefulness is the whole point and the whole problem. A prefill pod and a decode pod play different roles, hold different state, and cannot be swapped freely the way a stateless HTTP fleet can. Your capacity planning becomes two independent ratios instead of one, and getting the prefill:decode ratio wrong leaves one pool idle while the other queues. For a broader view of what this does to your bill, see LLM inference cost at scale.
KV cache handoff constraints
The KV cache built during prefill has to reach the decode pod, and this is where disaggregation gets sharp edges. vLLM warns: "prefill and decode compute their block sizes independently, but KV cache transfer requires both to match." A mismatch means the transfer simply does not work.
Cache residency matters too. That state "is expensive to rebuild, not uniformly available across the cluster, and often determines whether a request returns quickly or spends seconds recomputing previous work." So block size synchronization between engines and deciding where cache lives are first-class design decisions, not afterthoughts. If most of your traffic is multi-turn, prefix caching for multi-turn LLM agents covers how reused prefixes change the math.
KV-aware routing: why stateless routing fails
Once backends are stateful and specialized, classic load balancing collapses. The old assumptions about "interchangeable backends" and "independent requests" no longer hold, so round-robin or least-requests routing sends work to a pod that has to recompute a prefix another pod already holds.
KV-aware (cache-aware) routing fixes this by choosing pods on cache residency. Modular is blunt: "Cache state is the primary driver of prefill latency variance at scale. A router that selects pods based on cache residency eliminates prefill compute proportional to the shared prefix length for every hit." In other words, the router is not a nicety in a disaggregated system, it is what makes the topology pay off.
Orchestrator options
A few self-hosted orchestrators support these patterns as of 2026:
NVIDIA Dynamo / llm-d is built for the full pattern, listing "disaggregated prefill/decode, KV-aware routing" among its features.
vLLM offers "prefix caching per instance" rather than cluster-wide disaggregation, which covers cache reuse within a single engine.
LocalAI (v3) adds "prefix-cache-aware across replicas" routing.
Before reaching for a disaggregated stack, benchmark a plain single-engine setup: in our benchmarks of vLLM, SGLang and TensorRT-LLM for a Llama 3.1 70B FP8 API on a single H100, vLLM hit the lowest TTFT at 123ms and SGLang led throughput at 460 tokens/sec on batch size 64. Framework choice on one GPU can close much of the gap disaggregation targets.
When disaggregation isn't worth it
Disaggregation earns its keep at scale, with long prompts, or where prefill and decode genuinely starve each other. It is rarely worth it when:
Your traffic is low or bursty enough that two specialized pools sit idle. A single well-tuned engine is simpler and cheaper here.
Prompts are short, so prefill barely competes with decode.
You cannot guarantee block size synchronization across engines, since a mismatch breaks the KV transfer entirely.
Cold starts add another wrinkle: in our TensorRT-LLM guide we measured ~10-15s to load the model into GPU memory on a cold start, and that same tutorial reached ~1700 output tokens/sec (FP8) for Llama 3 8B on a single A10 without any disaggregation at all. If a single serverless GPU already meets your latency and throughput targets, splitting phases only adds moving parts. For picking the right chip in the first place, the 2026 GPU buyer's guide helps, and if you are weighing hyperscalers, see AWS alternatives for AI workloads.
Key takeaways
Disaggregated prefill/decode serving separates a compute-bound phase from a memory-bound one onto specialized, stateful GPU pods. It removes cross-phase interference but demands block size synchronization for KV handoff and KV-aware routing to place work near its cache. Reach for it when scale and prompt length justify it; otherwise a single well-chosen engine, sized on real inference cost and serverless GPU economics, is usually the better first move.
Frequently asked questions
- What does disaggregated prefill/decode serving actually separate?
- It puts the compute-bound prefill phase and the memory-bound decode phase on separate GPU pods, so each pool can be batched and scaled for its own bottleneck. As Modular notes, some pods specialize in prefill and others in decode, with the KV cache handed off between them.
- Why does prefill contend with decode on one GPU?
- Prefill saturates GPU cores with dense matrix multiplications while decode is bounded by KV cache memory, since every active request holds a slice of it. Running both together means long prompts stall decoding and large decode batches starve fresh prefills.
- What is the main constraint on KV cache handoff?
- Block sizes must match. vLLM warns that prefill and decode compute block sizes independently, but KV cache transfer requires both to match, so a mismatch breaks the transfer outright.
- Why does stateless routing fail in a disaggregated setup?
- Backends are no longer interchangeable; they hold large local KV caches that are expensive to rebuild. Round-robin routing sends work to pods that must recompute prefixes another pod already holds, so KV-aware routing that selects on cache residency is required.
- Which orchestrators support this?
- NVIDIA Dynamo/llm-d supports disaggregated prefill/decode with KV-aware routing. vLLM offers prefix caching per instance, and LocalAI v3 adds prefix-cache-aware routing across replicas.
Get started with Cerebrium
Deploy AI models on serverless GPUs in minutes, with no infrastructure to manage. Start for free and pay only for the compute you use.
Sources
- modular.com
“LLM Inference has two phases, and they stress hardware differently.”
Establishes the two-phase nature of inference.
- modular.com
“Some pods might specialize in prefill, others in decode.”
Pod specialization in disaggregated topology.
- modular.com
“Prefill processes the entire prompt in parallel. It's compute-bound. This means GPU cores are saturated doing dense [matrix multiplications].”
Prefill is compute-bound.
- vllm.ai
“KV cache capacity, since every active request holds a slice of it. KV cache size is thus the first question to answer for any topology, and it splits in two:”
Decode concurrency bounded by KV cache.
- modular.com
“They're GPU pods holding large, local KV caches in high-bandwidth, RAM or SSD memory. That state is expensive to rebuild, not uniformly available across the cluster, and often determines whether a request returns quickly or spends seconds recomputing previous work.”
Stateful, non-interchangeable pods and cache residency.
- vllm.ai
“One thing to keep in mind for disaggregated serving: prefill and decode compute their block sizes independently, but KV cache transfer requires both to match.”
Block size synchronization constraint for KV handoff.
- modular.com
“The backends here aren't interchangeable web servers.”
Why stateless routing breaks down.
- modular.com
“Cache state is the primary driver of prefill latency variance at scale. A router that selects pods based on cache residency eliminates prefill compute proportional to the shared prefix length for every hit.”
KV-aware routing benefit.
- nexlab.net
“**NVIDIA Dynamo** / **llm-d** | 8.1k / 4.6k | text | disaggregated prefill/decode, KV-aware routing | no | **KV-aware routing**”
Dynamo/llm-d orchestrator support.
- nexlab.net
“**vLLM** | 92k | text, vision, embeddings | TP/PP over Ray | no | prefix caching per instance | Prometheus metrics | no | no | no | yes (production-stack) | Linux, CUDA/ROCm/others | no”
vLLM prefix caching per instance.
- nexlab.net
“**LocalAI** | 49k | text, image, video, audio, embeddings, rerank | P2P federated, llama.cpp sharding, ds4 layer split | libp2p + shared token | prefix-cache-aware across replicas (v3) | per-key usage, users | no | no | no | Helm | Linux, macOS; Docker | **cosign**”
LocalAI v3 prefix-cache-aware routing.
- cerebrium.ai
“From our tests, it seems that vLLM is the best framework for use cases looking for the lowest TTFT of 123ms which is extremely quick if you take into consideration network latency.”
Cerebrium-measured 123ms TTFT.
- cerebrium.ai
“As for throughput use cases, it seems SGLang was a clear winner achieving a throughput of 460 tokens per second on a batch size of 64.”
Cerebrium-measured 460 tokens/sec throughput.
- cerebrium.ai
“This code will run on every cold start and takes roughly ~10-15s to load the model into GPU memory.”
Cerebrium cold-start model load time.
- cerebrium.ai
“In this tutorial we will achieve ~1700 output tokens per second (FP8)on a single Nvidia A10 instance however you can go up to ~4500 output tokens per second on a single Nvidia A100 40GB instance or even ~19,000 tokens on a H100.”
Cerebrium-measured ~1700 tokens/sec on A10.