Multiple Models on One GPU: Pack or Dedicate?

Connor Blier
Founding GTM

A monitor or GPU-card outline containing a grid of differently sized rectangular tiles packed together like windows, representing several separate models sharing the space of one processor.

Multiple Models on One GPU: Pack or Dedicate?

Running multiple models on one GPU means packing several models into a single accelerator and swapping or multiplexing between them, instead of dedicating a whole GPU to each. Pack when models are small or bursty and share hardware; dedicate when one model saturates the GPU on its own.

TL;DR

There are two ways to size an inference fleet. GPU-per-model gives every model its own accelerator: simple, isolated, and the safe default. Model-packing (multiplexing) puts several models on one GPU and swaps between them, reclaiming the capacity a dedicated GPU leaves idle. The safe default has a hidden tax: resource fragmentation. If your models are small, bursty, or numerous, packing usually wins on cost. If a single model saturates a GPU on its own at steady load, dedicate. This guide gives you the decision framework and the proof points behind it.

The two architectures

GPU-per-model is one deployment, one model, one (or more) whole GPUs. Each model scales independently, failures are isolated, and there is nothing to reason about beyond the model itself. It is the pattern almost every team starts with, and for good reason.

Model-packing / multiplexing loads multiple models into one GPU's memory and time-slices the compute between them. When a request arrives for model B while model A is loaded, the stack swaps context and serves it. IonRouter describes exactly this approach: "Our custom inference stack multiplexes models on a single GPU, swaps in ms, and adapts to traffic in real time." The unit of billing stops being the model and becomes the GPU-second, which is what changes the economics.

Why GPU-per-model wastes money: resource fragmentation

The problem with one-GPU-per-model is not the GPU you use, it is the GPU you reserve but don't use. Real traffic is bursty. A model provisioned for peak sits mostly idle at the trough, and its GPU cannot be lent to a busier neighbour because it belongs to a different deployment. Multiply that across a dozen models and you are paying for a fleet that is, on average, half-empty.

Anyscale names this precisely: "Resource fragmentation occurs when scaling activities lead to uneven resource utilization across nodes." Every scale-up and scale-down leaves stranded capacity behind. The dedicated-GPU model doesn't overspend on any one request; it overspends on the headroom it can never reclaim. If you are trying to bring down your LLM inference cost at scale, fragmentation is usually the first place to look.

How model-packing reclaims the GPU

Packing works when swapping between models is cheap enough that idle headroom disappears. The bar is sub-second, and increasingly sub-millisecond. IonRouter reports model "swaps in ms" and, in a production case, "Five vision-language models on a single GPU - 2,700 video clips, concurrent users, <1s cold starts." Five models where a dedicated fleet would have provisioned five GPUs.

The enabling technology is fast loading and fast restore. Slow cold starts are what make packing feel risky, because every swap looks like a fresh boot. That is exactly the tax memory snapshots attack. In our benchmarks, Cerebrium snapshots reduced cold starts by an average of 71% versus running the same workloads without snapshots, with reductions as high as 88% on vLLM - the full method and curves are in our write-up on reducing GPU cold starts with memory snapshots. When restoring a model is near-instant, keeping many models on one GPU stops costing you latency. We go deeper on the mechanism in our write-up on serverless GPU cold starts, and on the node-boot side we reduced machine boot time by 83% by reworking how platform images load.

Does packing cost throughput?

The fear is that sharing a GPU throttles each model. In practice a well-engineered stack can push a shared GPU harder than a naive dedicated one. Modal's Quail is the extreme proof: "On one multi-join query where planning is particularly important, Quail hits over a billion tokens processed per minute per H100 GPU (TPM/GPU), >10x faster than our vLLM baseline on the same hardware." A single H100, saturated, does not need a second GPU to be fast - it needs to not sit idle.

The lesson generalises. Throughput is lost to underutilization, not to sharing. If your models rarely max out the card, the throughput they leave on the table is exactly what a co-tenant reclaims. This is closely related to how H100 vs H200 inference throughput plays out: the ceiling only matters if you can reach it.

Operating a multi-model GPU in production

Packing is an operations discipline, not just an allocation trick. Two capabilities make it safe.

Heterogeneous co-location. Ray Serve lets models of different sizes share a service: "Ray Serve makes it possible to create deployments for Llama-3-8B and Llama-3-70B on the same Service with different resource requirements (1 GPU and 4 GPU per replica respectively)." You are not forced to pack identical models.

Automatic defragmentation. As replicas scale up and down, they scatter. Anyscale's Replica Compaction addresses this directly: "With Replica Compaction, Anyscale will automatically migrate replicas into fewer nodes in order to optimize resource use and reduce costs." Consolidation is what turns theoretical density into a smaller bill.

Roll new models onto a shared GPU the same way you would any endpoint - see our guide to canary rollout and rollback on a serverless GPU endpoint. Routing matters too: token-load-aware routing keeps a busy co-tenant from starving its neighbours. Cerebrium removes much of this operational drag; in one migration we describe cutting GPU costs nearly in half while reducing engineering overhead by more than 90%.

Decision framework: pack or dedicate

SignalPack (multiplex)Dedicate (GPU-per-model)
Model size vs GPUFits with room to spareSaturates the card alone
Traffic patternBursty, spiky, unevenSteady, sustained, high
Number of modelsMany small/mediumFew large
Latency budgetTolerant of ms-scale swapsHard, sub-swap SLAs
Isolation needsShared tenancy acceptableStrict blast-radius limits
Cost pressureHigh - reclaiming idle timeUtilization already high

Rules of thumb. Pack when several models each use a fraction of a GPU and their peaks don't coincide - the classic fleet of small vision, embedding, or fine-tuned models. Dedicate when a single model runs hot enough to keep a GPU busy on its own, or when tenancy isolation is a hard requirement. When you are between the two, packing plus fast swaps and automatic compaction is the lower-risk bet, because you can always split a hot model back out. For fine-tuned variants specifically, see serving fine-tuned LLMs on serverless GPU; for the smaller end of the fleet, concurrent voice sessions per GPU shows how far density can go.

The cheapest GPU is the one you already own, fully used. Cerebrium's serverless GPU platform is built to keep it that way, with 2-4 second cold starts and per-second billing described in our AWS alternatives guide.

Frequently asked questions

Can you really run multiple models on one GPU for inference?
Yes. A multiplexing inference stack loads several models into one GPU and swaps between them. IonRouter reports millisecond model swaps and, in production, five vision-language models on a single GPU serving 2,700 video clips with sub-second cold starts.
Does packing models onto one GPU hurt throughput?
Not if the GPU was underutilized to begin with. Throughput is lost to idle capacity, not to sharing. Modal's Quail hits over a billion tokens per minute on a single H100 - more than 10x its vLLM baseline on the same hardware - showing one saturated GPU can outperform a fleet of idle ones.
What is GPU resource fragmentation?
It is the stranded capacity left behind as model replicas scale up and down. Anyscale defines it as uneven resource utilization across nodes. Dedicated GPU-per-model fleets accumulate it constantly, which is why they overspend even when no single request is expensive.
When should I dedicate a GPU per model instead of packing?
Dedicate when a single model saturates the GPU at steady load, when latency SLAs leave no room for swap overhead, or when strict tenant isolation is required. Otherwise, packing with fast swaps and automatic compaction usually reclaims more capacity at lower cost.

Get started with Cerebrium

Deploy AI models on serverless GPUs in minutes, with no infrastructure to manage. Start for free and pay only for the compute you use.

Sign up free

Sources

  1. ionrouter.io
    “Our custom inference stack multiplexes models on a single GPU, swaps in ms, and adapts to traffic in real time.”

    IonRouter describes its multiplexing approach to running multiple models on one GPU.

  2. ionrouter.io
    “Five vision-language models on a single GPU — 2,700 video clips, concurrent users, <1s cold starts.”

    IonRouter production density example for packing models onto one GPU.

  3. modal.com
    “On one multi-join query where planning is particularly important, Quail hits over a billion tokens processed per minute per H100 GPU (TPM/GPU), >10x faster than our vLLM baseline on the same hardware.”

    Proof that a single saturated GPU delivers extreme throughput.

  4. anyscale.com
    “Ray Serve makes it possible to create deployments for Llama-3-8B and Llama-3-70B on the same Service with different resource requirements (1 GPU and 4 GPU per replica respectively).”

    Ray Serve supports heterogeneous multi-model deployments.

  5. anyscale.com
    “Resource fragmentation occurs when scaling activities lead to uneven resource utilization across nodes.”

    Definition of resource fragmentation, the hidden tax of GPU-per-model.

  6. anyscale.com
    “With Replica Compaction, Anyscale will automatically migrate replicas into fewer nodes in order to optimize resource use and reduce costs.”

    Replica Compaction consolidates replicas to reclaim fragmented capacity.

  7. cerebrium.ai
    “Across the benchmark suite, Cerebrium snapshots reduced cold starts by an average of 71% compared to running the same workloads on Cerebrium without snapshots, with reductions as high as 88% on vLLM.”

    Cerebrium first-hand benchmark showing fast restore that makes packing viable.

  8. cerebrium.ai
    “In this post, we show how we reworked node startup and initialization on AWS to reduce machine boot time by 83%, cut the long tail of cold starts, and lower the amount of excess capacity we needed to keep running, improving overall utilization.”

    Cerebrium's 83% machine boot time reduction supporting higher utilization.

  9. cerebrium.ai
    “Cerebrium removes the operational drag of managing multiple systems while cutting GPU costs nearly in half and reducing engineering overhead by more than **90%** — letting teams focus entirely on building, not babysitting infrastructure.”

    Cerebrium first-hand cost and overhead reduction figures.


Related resources

See all
A single settings-gear icon combined with a plus sign at its centre, symbolising a configurable, extensible API endpoint.
OpenAI-Compatible Endpoints for Open-Source LLMs
A hexagonal frame like a coin or badge with a plus sign at its centre, symbolizing a per-image cost unit.
FLUX Image API Cost Per Image: 2026 Provider Guide
A single-colour bar chart icon: five vertical bars of increasing height rising left to right along a baseline, with an upward diagonal arrow above the tallest bars pointing to the upper right, symbolizing rising inference throughput.
MI300X vs H200 LLM Inference Throughput