H100 vs H200 LLM Inference Throughput: A Buyer's Guide

Connor Blier
Founding GTM

A simple bar chart of five vertical bars increasing in height from left to right, with the tallest bar near the right side slightly shorter than its neighbor, sitting on a horizontal baseline, all rendered in solid black.

H100 vs H200 LLM Inference Throughput: A Buyer's Guide

For H100 vs H200 LLM inference throughput, the H200 wins on long-context, high-concurrency serving because it has 76% more memory at 43% higher bandwidth, not more compute. Short prompts and low concurrency see far smaller gains, so the H100 often remains the cheaper per-token choice.

Short answer

The H200 is not a faster chip than the H100. It is the same chip with more memory. NVIDIA ships the H200 with 141 GB of HBM3e at roughly 4.8 TB/s of bandwidth, which is "76% more VRAM and 43% more bandwidth than the H100 SXM5, which has 80 GB of HBM3 at 3.35 TB/s." The compute is identical: same CUDA cores, same FP16/BF16/FP8/INT8 throughput. So when you compare h100 vs h200 llm inference throughput, every extra token per second the H200 buys you comes from memory, not math.

That makes the decision unusually clean. If your workload is memory-bound (long context, big KV cache, high concurrency), the H200 can nearly double your throughput. If it is compute-bound or small (short prompts, low concurrency, a model that already fits comfortably in 80 GB), the two perform almost the same and the H100 is the smarter spend.

Why the difference is memory, not compute

Baseten puts it plainly: "The H200 GPU has the same compute as an H100 GPU, but with 76% more GPU memory (VRAM) at a 43% higher memory bandwidth." RunPod confirms the H200 SXM carries "CUDA Cores: 16,896 (identical to H100 SXM5)".

SpecH100 SXM5H200 SXM
Memory80 GB HBM3141 GB HBM3e
Bandwidth3.35 TB/s4.8 TB/s
CUDA cores16,89616,896 (identical)
FP16/BF16/FP8 computebaselineidentical

LLM decoding is memory-bound. Generating each new token means streaming the entire model's weights, plus the growing KV cache, out of VRAM and back. The arithmetic per token is trivial; the bottleneck is how fast you can move bytes. That is why bandwidth, not FLOPs, sets the ceiling on decode throughput, and why the H200's 43% bandwidth edge translates almost directly into faster generation on the workloads that saturate it.

The second lever is capacity. A model plus its KV cache has to fit in VRAM. When it does not, you either shard across more GPUs or evict sequences from the batch. The H200's extra 61 GB lets you keep more concurrent sequences resident, which is where the largest wins appear.

Throughput benchmarks: where the gains are largest

NVIDIA's headline numbers are the best case. It reports "Llama2 70B Inference 1.9XFaster" and "GPT-3 175B Inference 1.6XFaster" on the H200 versus the H100. Those are large, memory-hungry models, which is exactly where the bandwidth and capacity advantages compound.

Independent testing shows the same shape and clarifies the range. On Qwen3-32B in BF16 with vLLM, at a single user "Decode ran 1.2 to 1.3 times faster on the H200 at low concurrency" - a real but modest gain. Push the load up and the gap widens. On a 512-in / 512-out workload at 8 concurrent requests the same testing recorded the H200 at 401 tok/s against the H100's 302 tok/s ("512 in / 512 out | 8 | H100 | 302 ... 512 in / 512 out | 8 | H200 | 401").

The divergence becomes dramatic once context grows. At 8,192-token prompts with 8 concurrent requests, the H100 could only keep 3 of the 8 sequences resident while the H200 kept all 8, so throughput went from "8,192 in / 128 out | 8 | H100 | 60 ... 8,192 in / 128 out | 8 | H200 | 100" tok/s. And at the extreme, "At 16,384-token prompts the H200 was 2.2 times faster and 31% cheaper per token", because the H100 simply runs out of KV cache.

The pattern for h100 vs h200 llm inference throughput is consistent: gains shrink toward parity at short context and low concurrency, and grow toward 2x when context is long and the batch is deep. Our own benchmarking of vLLM, SGLang and TensorRT for a Llama 3.1 API shows how much the serving engine and batch size move the number too - in our benchmarks SGLang hit 460 tokens per second at batch size 64 on a single H100. If you are also weighing the next generation, see our note on B200 inference throughput and TTFT.

The cost-per-token tradeoff

More throughput is only cheaper if it outpaces the price premium. The H200 rents for more than the H100, so parity in throughput means the H100 wins on cost. That is why the 16,384-token case is decisive: the H200 was both faster and 31% cheaper per token, because the H100 was thrashing its KV cache. Below that regime, on short prompts where the two are within 20-30% of each other, the H100's lower hourly rate usually produces the lower cost per token.

The practical rule: cost-per-token, not tok/s, is the metric that pays your bill. We walk through how to reason about this at production volume in our guide to LLM inference cost at scale. Real deployments bear it out - DistilLabs cut inference costs by 50% largely by matching hardware and autoscaling to the actual workload shape rather than over-provisioning memory they did not need.

Software does as much as silicon

Before you buy bandwidth, make sure you are using the bandwidth you have. Optimized engines routinely deliver multiples that dwarf the H100-to-H200 gap. Modular's QUAIL engine reports that "Quail hits over a billion tokens processed per minute per H100 GPU (TPM/GPU), >10x faster than our vLLM baseline on the same hardware." That is a 10x swing on the same H100 - larger than anything a memory upgrade delivers.

Quantization, TensorRT-LLM, prefix caching and speculative decoding all reshape the memory-bound math. In our own testing with TensorRT-LLM we reached ~1700 output tokens per second in FP8 on a single A10, with a ~10-15s cold start to load the engine into GPU memory. Techniques like prefix caching for multi-turn agents and speculative decoding throughput gains often recover more headroom than a hardware tier change, and for memory-heavy architectures our note on mixture-of-experts VRAM footprint explains where capacity actually binds.

Decision framework for a Lead ML Engineer

Profile your workload first, then pick.

Choose the H200 when:

  • Your context windows are long (8K+ tokens) or your KV cache is large.

  • You serve high concurrency and want more sequences resident per GPU.

  • A model or its cache does not fit in 80 GB, forcing sharding on the H100.

  • Your batches are deep enough to saturate 4.8 TB/s of bandwidth.

Choose the H100 when:

  • Prompts are short and concurrency is low - you will land near parity.

  • Your model already fits comfortably in 80 GB with room for the cache.

  • You are compute-bound or latency-bound at batch size 1.

  • Cost per token, not peak throughput, is your governing constraint.

Whichever you choose, measure the full round trip on the platform you will bill against, not a vendor spec sheet. On Cerebrium's serverless GPUs you can benchmark H100 and H200 on your real traffic and switch tiers without rewriting your deployment, so the buyer's choice becomes a measurement rather than a guess.

Frequently asked questions

Is the H200 faster than the H100 for LLM inference?
Only on memory-bound workloads. The two chips have identical compute; the H200 adds 76% more VRAM at 43% higher bandwidth. NVIDIA reports up to 1.9x on Llama2 70B and 1.6x on GPT-3 175B, but independent testing shows only 1.2-1.3x at low concurrency and short context.
Why does memory matter more than compute for token throughput?
LLM decoding is memory-bound: generating each token streams the model weights and KV cache out of VRAM, so bandwidth sets the ceiling on decode speed. The arithmetic per token is small, which is why the H200's bandwidth advantage translates almost directly into throughput on saturated workloads.
When is the H100 the cheaper choice?
On short prompts and low concurrency where the two GPUs land within 20-30% of each other, the H100's lower hourly rate usually wins on cost per token. It is also the right pick when your model fits comfortably in 80 GB and you are latency-bound at batch size 1.
Can software close the gap between H100 and H200?
Often more than the hardware does. Modular's QUAIL engine reports over a billion tokens per minute per H100, more than 10x its vLLM baseline on the same hardware. Quantization, TensorRT-LLM, prefix caching and speculative decoding can recover more headroom than a memory-tier upgrade.

Get started with Cerebrium

Deploy AI models on serverless GPUs in minutes, with no infrastructure to manage. Start for free and pay only for the compute you use.

Sign up free

Sources

  1. runpod.io
    “This is 76% more VRAM and 43% more bandwidth than the H100 SXM5, which has 80 GB of HBM3 at 3.35 TB/s.”

    Spec comparison of H200 vs H100 memory and bandwidth.

  2. runpod.io
    “CUDA Cores: 16,896 (identical to H100 SXM5)”

    Confirms identical compute cores between H200 and H100.

  3. baseten.co
    “The H200 GPU has the same compute as an H100 GPU, but with 76% more GPU memory (VRAM) at a 43% higher memory bandwidth.”

    States compute is identical and only memory differs.

  4. nvidia.com
    “Llama2 70B Inference 1.9XFaster”

    NVIDIA's headline throughput claim for Llama2 70B.

  5. nvidia.com
    “GPT-3 175B Inference 1.6XFaster”

    NVIDIA's headline throughput claim for GPT-3 175B.

  6. jarvislabs.ai
    “Decode ran 1.2 to 1.3 times faster on the H200 at low concurrency”

    Independent low-concurrency decode measurement on Qwen3-32B.

  7. jarvislabs.ai
    “512 in / 512 out | 8 | H100 | 302 | 403 ms | 25.7 ms | 13.6 s | 8 | 512 in / 512 out | 8 | H200 | 401 | 337 ms | 19.3 ms | 10.2 s | 8”

    Throughput at 8 concurrent requests, 512 in/out.

  8. jarvislabs.ai
    “8,192 in / 128 out | 8 | H100 | 60 | 12.5 s | 34.6 ms | 17.3 s | 3 of 8 | 8,192 in / 128 out | 8 | H200 | 100 | 3.7 s | 51.2 ms | 10.2 s | 8”

    Long-context high-concurrency throughput and resident-sequence difference.

  9. jarvislabs.ai
    “At 16,384-token prompts the H200 was 2.2 times faster and 31% cheaper per token”

    Extreme-context case where H100 runs out of KV cache.

  10. modal.com
    “Quail hits over a billion tokens processed per minute per H100 GPU (TPM/GPU), >10x faster than our vLLM baseline on the same hardware.”

    Software optimization exceeding hardware-tier gains on the same H100.

  11. cerebrium.ai
    “As for throughput use cases, it seems SGLang was a clear winner achieving a throughput of 460 tokens per second on a batch size of 64.”

    Cerebrium first-hand H100 throughput benchmark.

  12. cerebrium.ai
    “In this tutorial we will achieve ~1700 output tokens per second (FP8)on a single Nvidia A10 instance however you can go up to ~4500 output tokens per second on a single Nvidia A100 40GB instance or even ~19,000 tokens on a H100.”

    Cerebrium first-hand TensorRT-LLM FP8 throughput.

  13. cerebrium.ai
    “With Cerebrium, we reduced their inference costs by 50%, and increased accuracy from 83% to 92%, while keeping latency and reliability consistent at production scale.”

    Customer cost-reduction case study from matching hardware to workload.


Related resources

See all
A single mechanical gear rendered as a settings-style cog with a ringed outer edge and a hollow circular center, symbolizing configurable infrastructure for running large language model inference.
Baseten Alternatives for Production LLM Inference
Two pairs of small rectangular processing blocks connected by short lines and arrows, representing separated stages of a computation pipeline feeding into one another.
Disaggregated Prefill/Decode LLM Serving: A Guide
A stylised speedometer-like dial with a clock hand, surrounded by eight small radiating spokes like a compass or gear, symbolising measuring load and directing traffic based on timing.
Token-Load-Aware Routing for LLM Serving