4-Bit LLM Inference on H100: Throughput, VRAM & Cost

Connor Blier
Founding GTM

A single microchip rendered as a square outline with pins on all four sides, containing two short vertical bars inside representing a compact, low-bit numeric value.

4-Bit LLM Inference on H100: Throughput, VRAM & Cost

4-bit LLM inference on the H100 is how you fit models that FP16 cannot, then serve them cheaply. Quantization makes 405B-class models runnable on 80GB cards, pushes throughput past a billion tokens per minute per H100, and drops cost to under six cents per billion tokens.

TL;DR

4-bit LLM inference on the H100 is first a memory decision and second a cost decision. An FP16 copy of Llama 3.1 405B needs roughly 800GB of VRAM before any context is loaded, so it simply does not fit on 80GB cards without quantization. Drop to 4-bit or FP8 and the same model becomes servable, while high-throughput engines can push over a billion tokens per minute per H100 at under 6 cents per billion tokens. This guide walks through the VRAM math, the throughput trade-offs between 4-bit, FP8 and FP16, and when to pick each.

Why precision is now the first serving decision

A year ago, most teams served in 16-bit and moved on. That has changed. Across production Qwen3 endpoints, 16-bit builds fell from 95.9% to 85.8% over the studied period, while 8-bit rose from 2.0% to 10.9% and 4-bit from 3.1% to 8.8%. The shift is sharpest at the large end: 40.7% of deployments serving models above 70B parameters use a quantized build, versus 11.6% for models under 8B. In other words, the bigger the model, the less optional quantization becomes.

For a Lead ML Engineer, that means precision is no longer a tuning afterthought - it is the first thing you decide, because it dictates which GPU you need, how many of them, and what each token costs. If you are also weighing silicon, our H100 vs H200 inference throughput guide covers where extra memory and bandwidth actually pay off.

The VRAM math: what fits on one 80GB H100

The H100 caps at 80GB of VRAM, and that ceiling drives everything. As a rough rule, FP16 weights consume ~2 bytes per parameter, FP8 ~1 byte, and 4-bit ~0.5 bytes - before the KV cache and activations that context length adds on top.

The extreme case makes the point. Serving Llama 3.1 405B in FP16 would require ~800GB before context load, which even eight 80GB cards cannot cover. As that analysis puts it, you'll need 8-bit or 4-bit quantization if you want to use NVIDIA cards. At 4-bit the same model's footprint falls far enough to become tractable, though still multi-GPU. For 70B-class models, FP8 or 4-bit is what brings weights plus KV cache comfortably inside a single 80GB H100. If your model is a mixture-of-experts architecture, the footprint math differs again; see our MoE VRAM footprint guide.

Throughput: 4-bit vs FP8 vs FP16

Lower precision does not only save memory - it moves more tokens, because quantized weights mean less data to read from HBM per step and more room for larger batches. The headline here is dramatic: Modal's Quail engine, running quantized 4-bit and FP8 models, hits over a billion tokens per minute per H100, more than 10x faster than its vLLM baseline on the same hardware.

Those are throughput-maximized numbers on short requests. In our own benchmarking of Llama 3.1 70B in FP8 on a single H100, SGLang was the throughput winner at 460 tokens per second on a batch size of 64 - a realistic figure for larger prompts and heavier per-token work. The two numbers are not contradictory; they bracket the range you see as batch size, prompt length and engine change. The full method and engine-by-engine comparison are in our vLLM vs SGLang vs TensorRT-LLM benchmark, which notes the measurements were taken as total roundtrip to the platform, not vendor specs.

Smaller models in FP8 are faster still. In our TensorRT-LLM tutorial we measured ~1700 output tokens/sec for Llama 3 8B in FP8 on a single A10, with an H100 figure cited from NVIDIA's own published numbers - the full build is in Running Llama 3 8B with TensorRT-LLM. Techniques like speculative decoding stack further gains on top of quantization.

Cost per token

Throughput is the lever, but cost per token is the number your CFO reads. Because a quantized model moves more tokens through the same rented H100, the cost of each token falls proportionally. With Quail on quantized models, H100 inference comes out to under 6 cents per billion tokens on Modal. For the full picture of how batching, utilization and idle time compound into your monthly bill, see our guide to LLM inference cost at scale.

4-bit vs FP8: which precision when

Use 4-bit when memory is the binding constraint - when the model would not otherwise fit on your card count at all, as with 405B-class models, or when you want maximum batch size on a single H100. The trade-off is that aggressive 4-bit quantization carries more accuracy risk and often needs calibration or a quality check against your eval set.

Use FP8 when you want most of the memory and throughput win with a smaller accuracy hit and far less setup. FP8 is the default sweet spot for 70B-and-under models on the H100, which natively accelerates it. For teams serving fine-tuned weights, precision choice interacts with how you load adapters - see serving fine-tuned LLMs.

Getting started: pre-optimized FP8 models

The fastest path is to skip quantizing yourself. Llama-3 and Hermes have FP8 versions already configured to work with vLLM on HuggingFace, which removes most of the configuration burden. Pull a pre-optimized FP8 checkpoint, serve it with vLLM or SGLang, and measure before you reach for 4-bit.

Cold starts matter too when you autoscale: in our TensorRT-LLM build we measured a ~10-15s load of the engine into GPU memory per cold start - budget for it or keep a warm pool. DistilLabs did exactly this kind of precision-plus-autoscaling tuning on Cerebrium and cut inference costs by 50% while raising accuracy from 83% to 92%; the full story is in the DistilLabs case study.

Deploy it on Cerebrium

Cerebrium runs quantized LLM inference on serverless H100s with per-second billing, fast cold starts and OpenAI-compatible endpoints - so you can serve a 4-bit or FP8 model, autoscale it, and pay only for the tokens you actually move. If you are comparing providers first, start with our Modal alternatives and Baseten alternatives breakdowns, then deploy.

Frequently asked questions

Can I run Llama 3.1 405B on a single 80GB H100?
No. In FP16 the model needs roughly 800GB of VRAM before context load, so even eight 80GB cards fall short. 8-bit or 4-bit quantization is mandatory to serve 405B-class models on NVIDIA cards at all, and it still requires multiple GPUs.
How much faster is 4-bit/FP8 inference than FP16 on the H100?
It depends on batch size and prompt length. Modal's Quail engine, running quantized 4-bit and FP8 models, reaches over a billion tokens per minute per H100 - more than 10x its vLLM baseline on the same hardware. In our own 70B FP8 benchmark, SGLang reached 460 tokens/sec at batch size 64.
What does 4-bit H100 inference cost per token?
High-throughput quantized inference with Quail on Modal comes out to under 6 cents per billion tokens. Your real figure depends on utilization, batching and idle time.
Should I choose 4-bit or FP8?
Choose 4-bit when memory is the binding constraint or you need maximum batch size, accepting more accuracy risk. Choose FP8 for most 70B-and-under models on the H100 - it keeps most of the throughput and memory win with a smaller accuracy hit and far less setup.
Is quantization widely used in production?
Increasingly yes, especially for large models. Across production Qwen3 endpoints 4-bit rose from 3.1% to 8.8% and 8-bit from 2.0% to 10.9%, and 40.7% of deployments serving 70B+ models use a quantized build versus 11.6% for sub-8B models.

Get started with Cerebrium

Deploy AI models on serverless GPUs in minutes, with no infrastructure to manage. Start for free and pay only for the compute you use.

Sign up free

Sources

  1. runpod.io
    “Among Qwen3 endpoints on Runpod, the share running 16-bit builds fell from **95.9%** to **85.8%** over the period we studied. Over the same period, 8-bit builds rose from 2.0% to **10.9%** and 4-bit from 3.1% to **8.8%**.”

    Adoption metrics for quantization precision on production inference endpoints.

  2. runpod.io
    “**40.7%** of Pods serving models above 70B parameters use a quantized build, compared with 11.6% of Pods serving models under 8B.”

    Quantization adoption correlates with model size.

  3. runpod.io
    “You'll need to run the model with 8-bit or 4-bit quantization if you want to use NVidia cards, as they currently cap out at 80GB and Runpod maxes out at 8 cards, which isn't enough for the ~800GB before context load that a fp16 405b would require.”

    FP16 405B needs ~800GB; quantization required on 80GB cards.

  4. runpod.io
    “The Llama-3 and Hermes models have FP8 versions already configured to work with vLLM, so we would recommend using those for ease of configuration and implementation.”

    Pre-optimized FP8 models available for vLLM.

  5. modal.com
    “Quail hits over a billion tokens processed per minute per H100 GPU (TPM/GPU), >10x faster than our vLLM baseline on the same hardware.”

    Peak throughput benchmark for quantized H100 inference.

  6. modal.com
    “On Modal, that comes out to under 6¢ per billion tokens.”

    Per-token cost for high-throughput quantized H100 inference.

  7. cerebrium.ai
    “As for throughput use cases, it seems SGLang was a clear winner achieving a throughput of 460 tokens per second on a batch size of 64.”

    Cerebrium-measured peak throughput, Llama 3.1 70B FP8 on a single H100.

  8. cerebrium.ai
    “In this tutorial we will achieve ~1700 output tokens per second (FP8)on a single Nvidia A10 instance however you can go up to ~4500 output tokens per second on a single Nvidia A100 40GB instance or even ~19,000 tokens on a H100.”

    Cerebrium-measured FP8 throughput for Llama 3 8B.

  9. cerebrium.ai
    “This code will run on every cold start and takes roughly ~10-15s to load the model into GPU memory.”

    Cerebrium cold-start model-load time.

  10. cerebrium.ai
    “With Cerebrium, we reduced their inference costs by 50%, and increased accuracy from 83% to 92%, while keeping latency and reliability consistent at production scale.”

    DistilLabs case study headline result.


Related resources

See all
A large four-pointed sparkle star at the centre of the frame, surrounded by two smaller sparkle stars of the same shape near its lower corners, all rendered as solid black shapes suggesting bursts of instant activation.
Modal vs RunPod: LLM Inference Cost & Cold Starts
A monitor or GPU-card outline containing a grid of differently sized rectangular tiles packed together like windows, representing several separate models sharing the space of one processor.
Multiple Models on One GPU: Pack or Dedicate?
A single settings-gear icon combined with a plus sign at its centre, symbolising a configurable, extensible API endpoint.
OpenAI-Compatible Endpoints for Open-Source LLMs