MI300X vs H200 LLM Inference Throughput

Connor Blier
Founding GTM

A single-colour bar chart icon: five vertical bars of increasing height rising left to right along a baseline, with an upward diagonal arrow above the tallest bars pointing to the upper right, symbolizing rising inference throughput.

MI300X vs H200 LLM Inference Throughput

On MI300X-vs-H200 LLM inference throughput, the gap is increasingly a software problem, not a silicon one. With a portable stack like Modular MAX, AMD's MI325X reaches throughput parity or better than vLLM on H200 for ShareGPT workloads, and the same container deploys across both vendors unchanged.

When a Lead ML Engineer asks whether MI300X can match H200 on LLM inference throughput, the honest answer in 2026 is: it mostly depends on your software. AMD's silicon has closed much of the raw-compute distance, and the remaining throughput differences are increasingly attributable to the maturity of the kernel and serving stack rather than the hardware itself. This guide walks through what the published benchmarks actually say, where the throughput is won, and what it means for a purchasing decision.

MI300X vs H200: why the throughput gap is a software problem, not a silicon one

For years the practical answer to "can we run on AMD?" was gated less by FLOPs than by the software ecosystem: hand-tuned kernels, vendor-locked runtimes, and serving frameworks that were optimized first for NVIDIA. That is the real shape of the MI300X vs H200 throughput conversation. When a portable, well-optimized stack runs on both, the vendor delta narrows sharply. The rest of this guide leans on published throughput numbers from that kind of stack, plus our own first-hand serverless benchmarks for context on what platform-level measurement looks like.

Throughput parity on named workloads: MAX vs vLLM on H200

The headline result comes from Modular's Platform 25.4 release: "Throughput parity or better on ShareGPT workloads running on MI325X when compared to vLLM on NVIDIA H200." One important caveat for honesty's sake: that parity claim is measured on the MI325X, a same-family sibling of the MI300X, not the MI300X part directly. Do not read it as a literal MI300X-equals-H200 statement. It is, however, strong directional evidence that AMD Instinct silicon can meet H200 on a real-world workload dataset when the serving software is competitive.

On the MI300X directly, the same release reports "Up to 32% better throughput for decode-heavy BF16 workloads when compared with vLLM on AMD MI300X" - i.e. MAX beating the vLLM baseline on the same AMD hardware. That is a software-vs-software comparison on identical silicon, which is exactly the point: the runtime you choose moves throughput by double digits.

If you want a sense of how much serving-framework choice matters generally, our own benchmarking of vLLM, SGLang and TensorRT on Llama 3.1 found SGLang a clear throughput winner at 460 tokens/sec on a batch size of 64. We measured that on a single H100; the framework delta on one GPU underscores why cross-vendor comparisons must hold the software constant.

The kernel layer: where AMD throughput is won

Drill one level down and the story is about kernels. Modular reports that "the Mojo kernel library implementation of matmul for BF16 outperforms equivalent hand-tuned kernels on MI300X, while maintaining portability to other hardware." This is the crux for anyone weighing AMD: historically you traded portability for performance, writing hand-tuned kernels per vendor. A portable kernel that beats the hand-tuned version on MI300X removes that trade-off. You get MI300X throughput without a one-vendor codebase.

Beyond inference: RL training numerical consistency on MI300X

Inference throughput is not the only axis Lead ML Engineers evaluate AMD on. Reinforcement-learning workflows demand that the training and rollout paths agree numerically, and that has been a real blocker on newer hardware. Here MI300X now has published validation: in an end-to-end "Qwen3-8B GRPO experiment on AMD Instinct MI300X, the vime + RL-Kernel strict path ran for 200 consecutive steps with mismatch_count = 0 and max_abs_diff = 0 throughout" - bitwise consistency between training (Megatron) and inference (vLLM) rollout across 8x MI300X via RL-Kernel v0.1.0. Zero mismatches over 200 steps is the kind of determinism that makes an RL pipeline trustworthy on AMD, not just a throughput demo.

One container, both vendors

The commercial unlock is deployment portability. The Modular 25.4 release ships a single container that runs across AMD MI300X/MI325X and NVIDIA GPUs with no code changes. For a purchasing decision that means you are no longer betting your codebase on one vendor's roadmap or spot availability; you build once and place the workload wherever capacity and price are best. That is the same logic behind choosing a flexible serverless platform and evaluating AWS alternatives for AI workloads: reduce lock-in, keep options open on hardware. It also feeds directly into LLM inference cost at scale, where the ability to shift vendors is a real lever on the bill.

Measurement discipline still matters. As we argue in our guide to speculative decoding throughput on serverless GPU, you should not trust vendor throughput specs - measure the full round trip on the platform you will bill against. Our own numbers are stated that way: on a TensorRT-LLM Llama 3 8B deployment we saw ~1700 output tokens/sec (FP8) on a single A10, with a ~10-15s cold start to load the model into GPU memory. Apply the same rigor to any MI300X-vs-H200 claim you evaluate.

Decision summary for Lead ML Engineers

  • If the question is raw silicon, treat MI300X and H200 as close enough that software decides the outcome. The published parity result is on MI325X, not MI300X directly - weight it accordingly.

  • On identical MI300X hardware, runtime choice alone moved decode-heavy BF16 throughput by up to 32%. Invest in the serving stack before the GPU.

  • Portable kernels can beat hand-tuned MI300X kernels, so you no longer pay a portability tax to run on AMD.

  • RL determinism on MI300X is now demonstrated (200 steps, zero mismatches), broadening AMD beyond pure inference.

  • Deploying one container across both vendors is the durable commercial win: it turns "MI300X or H200" from a lock-in decision into a placement decision.

For teams thinking about where this runs in production, see our notes on GPU inference for AI agent workloads, serving fine-tuned LLMs on serverless GPU, and the 2026 GPU Buyer's Guide.

Frequently asked questions

Does MI300X match H200 on LLM inference throughput?
The published parity result - throughput parity or better than vLLM on H200 for ShareGPT workloads - was measured on the MI325X, a same-family sibling of the MI300X, using Modular MAX. On the MI300X directly, MAX delivered up to 32% better throughput than vLLM for decode-heavy BF16 workloads, but that is a software comparison on the same AMD hardware, not a direct MI300X-vs-H200 parity claim.
Why is the MI300X vs H200 gap called a software problem?
Because the largest published throughput swings come from the serving runtime and kernel layer, not the silicon. On identical MI300X hardware, changing the runtime moved decode-heavy BF16 throughput by up to 32%, and a portable Mojo BF16 matmul kernel outperformed hand-tuned MI300X kernels.
Can I deploy the same code on both AMD and NVIDIA?
Yes. Modular Platform 25.4 ships a single container that runs across AMD MI300X/MI325X and NVIDIA GPUs with no code changes, which turns vendor selection into a placement decision rather than a lock-in commitment.
Is MI300X reliable for RL training, not just inference?
A published Qwen3-8B GRPO experiment on 8x AMD Instinct MI300X ran 200 consecutive steps with mismatch_count = 0 and max_abs_diff = 0 between Megatron training and vLLM rollout, via RL-Kernel v0.1.0 - bitwise consistency throughout.
How should I benchmark these GPUs for my own workload?
Do not trust vendor throughput specs. Measure the full round trip on the platform you will bill against, with a stated methodology (input tokens, batch size, GPU). Our own serverless benchmarks, such as 460 tokens/sec via SGLang at batch size 64 on an H100, are reported that way.

Get started with Cerebrium

Deploy AI models on serverless GPUs in minutes, with no infrastructure to manage. Start for free and pay only for the compute you use.

Sign up free

Sources

  1. modular.com
    “Throughput parity or better on ShareGPT workloads running on MI325X when compared to vLLM on NVIDIA H200.”

    Parity result cited in the parity section; explicitly noted as MI325X, not MI300X.

  2. modular.com
    “Up to 32% better throughput for decode-heavy BF16 workloads when compared with vLLM on AMD MI300X.”

    Software-vs-software throughput gain on identical MI300X hardware.

  3. modular.com
    “the Mojo kernel library implementation of `matmul` for `BF16` outperforms equivalent hand-tuned kernels on MI300X, while maintaining portability to other hardware.”

    Kernel-layer evidence that portability no longer costs performance on MI300X.

  4. vllm.ai
    “In an end-to-end Qwen3-8B GRPO experiment on AMD Instinct MI300X, the vime + RL-Kernel strict path ran for 200 consecutive steps with mismatch_count = 0 and max_abs_diff = 0 throughout.”

    RL numerical-consistency validation on MI300X.

  5. cerebrium.ai
    “As for throughput use cases, it seems SGLang was a clear winner achieving a throughput of 460 tokens per second on a batch size of 64.”

    First-hand Cerebrium throughput measurement illustrating framework impact on one H100.

  6. cerebrium.ai
    “This code will run on every cold start and takes roughly ~10-15s to load the model into GPU memory.”

    First-hand Cerebrium cold-start figure supporting the measurement-discipline point.

  7. cerebrium.ai
    “In this tutorial we will achieve ~1700 output tokens per second (FP8)on a single Nvidia A10 instance”

    First-hand Cerebrium throughput figure (A10, FP8) used as the platform-measurement example.


Related resources

See all
A hexagonal frame like a coin or badge with a plus sign at its centre, symbolizing a per-image cost unit.
FLUX Image API Cost Per Image: 2026 Provider Guide
A single mechanical gear rendered as a settings-style cog with a ringed outer edge and a hollow circular center, symbolizing configurable infrastructure for running large language model inference.
Baseten Alternatives for Production LLM Inference
A simple bar chart of five vertical bars increasing in height from left to right, with the tallest bar near the right side slightly shorter than its neighbor, sitting on a horizontal baseline, all rendered in solid black.
H100 vs H200 LLM Inference Throughput: A Buyer's Guide