Llama 405B Serverless GPU Inference Cost: An Honest Model

Connor Blier
Founding GTM

A single stylised robot or AI head icon, face-on, shown as a rounded square outline with two circular eyes, a triangular nose marker, and small limb-like attachment points on either side, rendered as a flat black line drawing.

Llama 405B Serverless GPU Inference Cost: An Honest Model

Llama 405B serverless GPU inference cost is driven less by a headline per-token price than by VRAM footprint, quantization, cold starts, idle billing and concurrency. At roughly 800GB in fp16, the model forces 8-bit or 4-bit quantization and multi-GPU nodes, so your real bill depends on utilization, not list rate.

The short answer on Llama 405B serverless GPU inference cost

There is no single number. The llama 405b serverless gpu inference cost you actually pay is the product of how much hardware the model demands, how you are billed for it, and how busy you keep it. A per-million-token headline looks clean on a pricing page, but for a 405-billion-parameter model it hides the three things that move your bill the most: the VRAM footprint that dictates your node shape, the billing granularity that decides whether idle time costs you, and the concurrency that determines how many tokens each expensive GPU-second actually produces.

This guide walks the honest cost model for an ML engineer evaluating serverless for Llama 405B, and is deliberate about where we can cite a measured number and where we cannot.

The VRAM reality sets the floor

Before you compare any price, you have to fit the model. Llama 405B is large enough that full precision is impractical on current single-node serverless hardware. RunPod states the constraint plainly: "You'll need to run the model with 8-bit or 4-bit quantization if you want to use NVidia cards, as they currently cap out at 80GB and Runpod maxes out at 8 cards, which isn't enough for the ~800GB before context load that a fp16 405b would require."

That single fact drives most of the cost conversation. A ~800GB fp16 footprint exceeds an 8x80GB node, so you are choosing between quantization (to fit on fewer cards) or multi-node serving (which adds interconnect requirements). For the 4-bit path and why it changes both memory and throughput, see our deeper write-up on 4-bit LLM inference on H100. If you are weighing node shapes and accelerators, the 2026 GPU Buyer's Guide and our H100 vs H200 throughput comparison are the right next reads.

Cost-per-token vs cost-per-GPU-second

Two pricing lenses dominate, and they answer different questions.

Cost-per-GPU-second is what most serverless GPU platforms meter. It is honest about the resource but says nothing about output until you know your throughput. Billing granularity matters here: Modal describes the difference against reserved models as "on-demand cluster access means paying by the hour or taking a reservation. On Modal, you run whenever you want and only pay for what you use." Per-second billing is strictly better for spiky traffic because you are not paying for whole idle hours.

Cost-per-token is what your product economics care about. To convert GPU-seconds into tokens you need measured throughput. We don't have a published 405B throughput figure, so we won't invent one - but as a directional anchor, in our benchmarks for Llama 3.1 70B FP8 on a single H100, SGLang reached 460 tokens per second at a batch size of 64. A 405B model will produce fewer tokens per GPU-second than a 70B one, so treat that as a ceiling, not a target. The full method and framework comparison live in our vLLM, SGLang and TensorRT benchmark.

The positioning argument for open models still holds: RunPod states that for a model this size, "It is far more cost effective per token to run a model of this size on Runpod" than paying closed-source API rates. That is the commercial case for self-serving Llama 405B at all - but it only materializes if your utilization is high.

The hidden drivers: cold starts, idle time and concurrency

Cold starts

Loading a very large model into GPU memory is not instant. For a far smaller model - Llama 3 8B on TensorRT-LLM - we measured a cold start where "This code will run on every cold start and takes roughly ~10-15s to load the model into GPU memory." A 405B model, even quantized across multiple cards, loads substantially more weight than 8B, so budget for longer. Every cold start is GPU-seconds you pay for before serving a single token, which is why cold-start frequency is a real cost line, not just a latency concern. Our notes on serverless GPU cold starts go deeper on mitigation.

Idle billing

The serverless premise is that you avoid paying for an idle accelerator between requests. That is exactly where serverless beats a reserved 8-GPU node for bursty traffic - and exactly where it loses for steady, saturated traffic. We cover the crossover in reserved GPU capacity for production LLM inference.

Concurrency and utilization

Utilization is the lever that turns a scary per-GPU-second rate into a competitive per-token cost. Higher batch sizes amortize fixed GPU cost across more tokens. Platform-level packing helps too: Anyscale claims you can "Deploy Ray Serve with up to 50% fewer nodes using Anyscale Replica Compaction" through automatic replica migration - fewer nodes for the same work is a direct cost reduction. For the scaling math at volume, see LLM inference cost at scale.

When serverless fits - and when dedicated wins

Serverless Llama 405B makes sense when traffic is variable, when you want per-second billing to track real usage, and when you cannot justify a permanently reserved multi-GPU node. Dedicated or reserved capacity wins once a node is saturated around the clock, because at that point there is no idle time for serverless to save you and reserved pricing is lower per hour.

The decision is a utilization calculation, not a vendor preference. The managed-platform case is that engineering effort - quantization, multi-GPU serving, autoscaling, cold-start control - is itself a cost. On that front, Cerebrium's customer DistilLabs reported that "With Cerebrium, we reduced their inference costs by 50%, and increased accuracy from 83% to 92%, while keeping latency and reliability consistent at production scale" - a reminder that configuration and autoscaling, not just list price, determine what you spend. If you are also comparing providers, our Modal vs RunPod cost and cold starts breakdown is the practical companion to this guide.

How to estimate your own number

  1. Pick a quantization (8-bit or 4-bit) and confirm it fits your intended node shape - this fixes your GPU-second rate.

  2. Benchmark tokens-per-second at your real batch size and context length; do not borrow a smaller model's throughput.

  3. Divide GPU-second cost by tokens-per-second to get cost-per-token at full load.

  4. Multiply by expected idle and cold-start overhead to get your effective, real-world cost.

That chain - rate, throughput, utilization - is the whole model. Anyone quoting a flat per-million-token price for Llama 405B without it is quoting a figure that assumes perfect utilization.

Frequently asked questions

Why can't I run Llama 405B in fp16 on a single serverless node?
Because the full-precision model needs roughly 800GB of VRAM before context load. RunPod notes NVIDIA cards cap at 80GB and its nodes max out at 8 cards, so you must use 8-bit or 4-bit quantization or serve across multiple nodes.
Should I compare Llama 405B cost per token or per GPU-second?
Both. Platforms bill per GPU-second, but your product economics depend on cost per token. Convert between them using measured throughput at your real batch size and context length - never a smaller model's numbers.
Do cold starts really affect cost?
Yes. Loading weights into GPU memory is billable GPU time before any token is served. We measured ~10-15s just to load Llama 3 8B on TensorRT-LLM; a 405B model loads far more weight, so frequent cold starts add up.
When does dedicated capacity beat serverless for Llama 405B?
Once a node runs saturated around the clock. Serverless saves you idle time; if there is no idle time, reserved pricing is cheaper per hour. The crossover is a utilization calculation.

Get started with Cerebrium

Deploy AI models on serverless GPUs in minutes, with no infrastructure to manage. Start for free and pay only for the compute you use.

Sign up free

Sources

  1. runpod.io
    “You'll need to run the model with 8-bit or 4-bit quantization if you want to use NVidia cards, as they currently cap out at 80GB and Runpod maxes out at 8 cards, which isn't enough for the ~800GB before context load that a fp16 405b would require.”

    Sets the VRAM floor and quantization requirement for Llama 405B serverless deployment.

  2. runpod.io
    “It is far more cost effective per token to run a model of this size on Runpod.”

    Supports the commercial case for self-serving open models vs closed-source APIs.

  3. modal.com
    “Elsewhere, on-demand cluster access means paying by the hour or taking a reservation. On Modal, you run whenever you want and only pay for what you use.”

    Illustrates per-second serverless billing vs hourly/reservation models.

  4. anyscale.com
    “Deploy Ray Serve with up to 50% fewer nodes using Anyscale Replica Compaction”

    Node-reduction as a utilization-driven cost lever.

  5. cerebrium.ai
    “This code will run on every cold start and takes roughly ~10-15s to load the model into GPU memory.”

    First-hand Cerebrium cold-start measurement used as a directional anchor.

  6. cerebrium.ai
    “As for throughput use cases, it seems SGLang was a clear winner achieving a throughput of 460 tokens per second on a batch size of 64.”

    First-hand throughput figure used as a ceiling reference for converting GPU-seconds to tokens.

  7. cerebrium.ai
    “With Cerebrium, we reduced their inference costs by 50%, and increased accuracy from 83% to 92%, while keeping latency and reliability consistent at production scale.”

    Shows configuration/autoscaling, not list price, drives real spend.


Related resources

See all
A single server-rack module with three stacked slots, flanked by two arrows: one pointing up-left into the rack and one pointing down-right away from it, suggesting replicas scaling up and down automatically.
Autoscaling Llama 70B: Cutting Replica Startup Latency
A hexagonal network node outline, like a chip or module icon, with a simple letter-H-shaped bracket motif at its centre, drawn as a single black line-art mark on a transparent background.
CoreWeave Alternatives for Production LLM Inference (2026)
A speedometer-style gauge with a needle pointed toward the upper right red zone, symbolizing a system running against a latency limit.
LLM Serving Goodput Under Latency SLOs: A Guide