LLM Serving Goodput Under Latency SLOs: A Guide
Connor Blier
Founding GTM
LLM Serving Goodput Under Latency SLOs: A Guide
Goodput under latency SLOs is the request rate an LLM endpoint sustains while requests still meet both their time-to-first-token (TTFT) and inter-token latency (ITL) targets. Unlike raw throughput, it only counts requests that honor the contract, making it the production metric buyers should plan capacity against.
Goodput under latency SLOs is the production metric, not the technique
If you are sizing an LLM deployment, the number that matters is not how many tokens a GPU can emit in a vacuum. It is how many requests per second you can serve while every one of them still meets its latency promise. That number is goodput, and the vLLM team defines it precisely: "the request rate you can sustain while requests still meet both their TTFT and ITL targets."
Goodput is frequently conflated with one technique that happens to raise it - disaggregated prefill/decode serving. That is a mistake. Disaggregation is a lever; goodput is the dial the lever moves. This guide defines the metric, separates it from throughput, pins down the SLO contract it depends on, shows how to measure it reproducibly, and then walks the levers that raise it.
Goodput vs throughput
Throughput asks: how many tokens or requests can this system emit per second at full tilt? Goodput asks a stricter question: how many of those requests actually landed inside the latency budget? A server can post a gaudy throughput number while quietly violating the SLO on half its requests under load - those requests are real work, but to a user they failed. Goodput discards them.
That distinction is why goodput is the buyer metric. Two endpoints can advertise identical throughput; the one with better tail-latency behavior will serve more usable traffic before the contract breaks. The gap shows up exactly where it hurts - at the tail, under concurrency.
The SLO contract: TTFT, ITL/TPOT, percentiles
Goodput is meaningless without the SLO it is measured against. For streaming LLMs that contract has two clauses:
Time-to-first-token (TTFT): how long before the first token appears. This drives perceived responsiveness. Unblocked, building adaptive routing across inference providers, records "the provider, token counts, duration, time to first token and task ID" on every call - TTFT is a first-class metric, not an afterthought.
Inter-token latency (ITL), also called TPOT: the gap between successive output tokens, which governs how smoothly text streams once it starts.
Both clauses must be expressed as percentiles, not averages. A p50 that looks healthy can hide a p99 that is five times worse - and the p99 is what your angriest users experience. Report goodput as the sustained rate at which p99 TTFT and p99 ITL both stay under target.
Provider choice matters here because the same model behaves differently across backends. As Unblocked observed, "Baseten, Fireworks and CoreWeave all serve GLM 5.2. They charge different prices, run at different speeds, and have outages at different times." Your goodput is a property of your stack, not just your model.
How to measure and report goodput reproducibly
A goodput number is only credible if the methodology travels with it. Pin down, at minimum:
Hardware and model. State the GPU and the exact model. In our benchmarking of vLLM, SGLang and TensorRT-LLM, we fixed this: "The following benchmark was done with 256 input tokens on 1xH100 of batch size 1. This was the total roundtrip time of a request made to the Cerebrium platform." Note that this is the total roundtrip, so network latency is included rather than hidden.
Arrival pattern and load sweep. Goodput is a curve, not a point. Drive the system with a realistic arrival process (Poisson is common) and sweep the request rate until the SLO breaks. The last rate that still satisfies the SLO is your goodput.
Input/output shape. Prompt length and output length change everything. A long-prompt, short-output profile stresses prefill; a short-prompt, long-output profile stresses decode.
On the Cerebrium platform, we measured a lowest TTFT of 123 ms with vLLM - "extremely quick if you take into consideration network latency." That number is only meaningful because the shape (256 input tokens, 1xH100, batch size 1, full roundtrip) is published alongside it. Do the same with your goodput.
Levers that raise goodput
Several techniques move the goodput curve. Disaggregated serving is one of them, not the whole story.
Disaggregated prefill/decode
Splitting the compute-bound prefill phase and the memory-bound decode phase onto separate GPU pods stops them from contending for the same hardware, which protects the tail. As the vLLM guide puts it, "A decode instance that only runs decode batches never has a long prefill stall its streams." The measured effect is stark: on two L40S GPUs with Qwen2.5-7B, "Collocated p99 jumps from 23 ms to 169 ms at 0.4 req/s and reaches 263 ms at 2 req/s. P/D p99 stays between 25 and 52 ms." At 0.4–0.6 req/s "the medians match and the collocated p99 is about six times higher" - same median, far tighter tail, so far more requests clear the SLO. The full method and curve are in our disaggregated prefill/decode serving guide.
Other levers worth pulling first
Engine and quantization. Framework choice alone shifts TTFT and throughput; see our vLLM vs SGLang vs TensorRT-LLM benchmark. With TensorRT-LLM in FP8 we hit ~1700 output tokens/sec on a single A10.
Prefix caching for multi-turn and shared-prompt traffic - covered in prefix caching for multi-turn agents.
Speculative decoding to lift decode throughput without hurting ITL, in speculative decoding throughput.
Load-aware routing so no single replica tips past its SLO, in token-load-aware routing.
Hardware selection, since the GPU sets the ceiling - compare options in H100 vs H200 inference throughput.
Applying goodput to capacity planning
Once goodput is a curve rather than a slogan, capacity planning becomes arithmetic. Divide peak offered load by the per-replica goodput at your SLO and you have the replica count - with no silent SLO violations baked in. This is also where goodput meets cost: higher goodput per GPU means fewer GPUs for the same promise. DistilLabs shows the payoff of tuning the whole stack this way - "we reduced their inference costs by 50%, and increased accuracy from 83% to 92%, while keeping latency and reliability consistent at production scale" (see the DistilLabs case study).
Factor in cold starts too: our TensorRT-LLM engine takes "roughly ~10-15s to load the model into GPU memory" on a cold start, so autoscaling against a goodput target has to account for warm-up, not just steady state. For the broader cost picture, see LLM inference cost at scale, and remember that the roundtrip in your SLO includes the wire - network latency in the inference pipeline often explains a goodput gap no engine tweak will close.
Treat goodput as the contract-aware throughput it is, publish the methodology beside the number, and let the techniques compete on how much of the curve they move.
Frequently asked questions
- What is goodput in LLM serving?
- Goodput is the request rate an endpoint can sustain while requests still meet both their time-to-first-token (TTFT) and inter-token latency (ITL) targets. It counts only requests that honor the latency contract, so it reflects usable capacity rather than raw emission rate.
- How is goodput different from throughput?
- Throughput counts every request or token the system emits at full load. Goodput discards any request that misses its latency SLO. A server can show high throughput while violating the SLO under load; goodput exposes that by only crediting requests inside the budget.
- Which latency metrics define the SLO for goodput?
- TTFT (time to first token) and ITL, also called TPOT (time per output token), both expressed as percentiles. Report goodput as the sustained request rate at which p99 TTFT and p99 ITL both stay under target.
- Does disaggregated serving define goodput?
- No. Disaggregated prefill/decode serving is one technique that can raise goodput, mainly by protecting tail latency. In measured tests it kept p99 inter-token latency between 25 and 52 ms where collocated serving reached 263 ms. Goodput is the metric; disaggregation is a lever.
- How should I report a goodput number so others can reproduce it?
- Pin the GPU and model, the arrival pattern and load sweep, the input and output token shape, and whether network roundtrip is included. For example, our TTFT figures were taken with 256 input tokens on 1xH100 at batch size 1 as a full roundtrip to the platform.
Get started with Cerebrium
Deploy AI models on serverless GPUs in minutes, with no infrastructure to manage. Start for free and pay only for the compute you use.
Sources
- vllm.ai
“What it buys is **goodput**: the request rate you can sustain while requests still meet both their TTFT and ITL targets.”
Definition of goodput used in the lead and opening section.
- vllm.ai
“Tail latency stays low as load climbs. A decode instance that only runs decode batches never has a long prefill stall its streams.”
Explains the tail-latency benefit of disaggregated serving.
- vllm.ai
“Collocated p99 jumps from 23 ms to 169 ms at 0.4 req/s and reaches 263 ms at 2 req/s. P/D p99 stays between 25 and 52 ms.”
Measured p99 ITL figures contrasting collocated and disaggregated serving.
- vllm.ai
“At 0.4–0.6 req/s the medians match and the collocated p99 is about six times higher.”
Shows median parity with large tail divergence, the core goodput argument.
- getunblocked.com
“We record every call in a ledger with the provider, token counts, duration, time to first token and task ID”
TTFT as a first-class measured metric in the SLO contract section.
- getunblocked.com
“Baseten, Fireworks and CoreWeave all serve GLM 5.2. They charge different prices, run at different speeds, and have outages at different times.”
Provider latency varies for the same model; goodput is stack-dependent.
- cerebrium.ai
“From our tests, it seems that vLLM is the best framework for use cases looking for the lowest TTFT of 123ms which is extremely quick if you take into consideration network latency.”
First-hand Cerebrium TTFT measurement used in the measurement section.
- cerebrium.ai
“The following benchmark was done with 256 input tokens on 1xH100 of batch size 1. This was the total roundtrip time of a request made to the Cerebrium platform.”
Methodology statement for reproducible reporting.
- cerebrium.ai
“In this tutorial we will achieve ~1700 output tokens per second (FP8)on a single Nvidia A10 instance”
First-hand throughput figure cited as an engine/quantization lever.
- cerebrium.ai
“This code will run on every cold start and takes roughly ~10-15s to load the model into GPU memory.”
Cold-start time factored into capacity planning.
- cerebrium.ai
“With Cerebrium, we reduced their inference costs by 50%, and increased accuracy from 83% to 92%, while keeping latency and reliability consistent at production scale.”
Cost payoff of raising goodput per GPU in capacity planning.