Multi-Node LLM Inference on Serverless Clusters

Connor Blier
Founding GTM

A diagram of five small server-rack boxes arranged in a cross pattern \u2014 one in the center and four at the corners \u2014 connected to each other by thin lines, resembling a distributed computing cluster.

Multi-Node LLM Inference on Serverless Clusters

Multi-node LLM inference on serverless clusters lets you run models too large for one node without operating the GPU fabric yourself. The real decision is not whether to go multi-node, but whether to build and babysit the cluster or consume it as a billed-by-the-second serverless primitive.

Multi-node LLM inference on serverless clusters lets you serve models that no longer fit in a single node's memory without standing up and operating your own GPU fabric. For an engineering leader, the decision is rarely whether to go multi-node - frontier and large open-weight models force your hand. The decision is whether your team should operate the interconnect, scheduler, and autoscaler, or consume them as a managed, per-second serverless primitive. This guide frames that build-vs-buy call in VP terms: cost, reliability, and where your engineers spend their week.

Why multi-node inference is now unavoidable

The largest open-weight and frontier models exceed the memory of a single GPU node. Once weights plus KV cache cross that ceiling, you must shard the model across nodes and run tensor, pipeline, or expert parallelism - which only works if those nodes talk to each other over a high-bandwidth, low-latency fabric. This is no longer an exotic training-only requirement: inference needs it too, specifically for KV cache transfer when you split prefill and decode across machines.

That is the uncomfortable part for a VP of Engineering. Single-node serverless was a clean abstraction. Multi-node reintroduces the hard systems problem - scheduling, networking, failure domains - that serverless was supposed to hide. The question is who absorbs it.

The hidden cost of the DIY cluster path

Building your own multi-node inference cluster looks like a procurement exercise and turns into a staffing commitment. The line items that do not appear in the GPU quote:

  • Interconnect engineering. RDMA over InfiniBand or RoCE has to be configured, tuned, and kept working for PyTorch and NCCL. This is specialist work, and it is load-bearing for every request.

  • Gang scheduling. You need all N nodes allocated together or the job cannot start. Off-the-shelf schedulers assume a single-node unit; you will be extending one.

  • Idle burn on reservations. Reserved or hourly cluster capacity bills whether or not traffic is flowing. Fragmented, partially-used nodes are pure waste.

  • Autoscaling and compaction. Scaling each model independently fragments the cluster; reclaiming that capacity is its own ongoing project.

  • On-call surface. Every layer above is now something your team pages on at 3 a.m.

Much of this overlaps with the operational burden we described in Serverless Training for LLMs: what it actually means - the inference version simply has tighter latency SLOs. If you are weighing hyperscaler DIY against managed options, our AWS alternatives for AI workloads breakdown is a useful companion.

What "serverless" changes for multi-node

Serverless does not remove the systems problem; it moves ownership of it to the platform. Three mechanics matter most.

Gang scheduling and atomic N-node allocation

A multi-node job is useless until every node is ready. The scheduler's atomic unit has to expand from one node to N. As the Modal Clusters announcement puts it, "the atomic unit for scheduling expands from a single node to N nodes." On a serverless platform this is the vendor's problem, not your Kubernetes operator's.

Automatic RDMA for PyTorch and NCCL

High-bandwidth node-to-node communication is the difference between a working shard and a stalled one. Managed clusters can wire this up for you - communication "via InfiniBand verbs at up to 6.4 Tbps, automatically configured for PyTorch and NCCL." The same fabric serves both training's all-reduce and inference's KV cache transfer during prefill-decode disaggregation, so you are not maintaining two networking stacks.

Per-second billing on shared serverless capacity

The economic unlock is that you draw nodes from a shared pool and pay only while you hold them - "your code running on a cluster within seconds of requesting it, billed by the second." That is structurally different from a reserved block you pay for around the clock. If reserved capacity is still on your table for baseline load, weigh it against on-demand in our note on reserved GPU capacity for production LLM inference, and model the totals with LLM inference cost at scale.

Architectural levers that map to VP metrics

Two techniques convert the multi-node primitive into cost and latency wins you can defend in a board deck.

Prefill-decode disaggregation

Prefill (compute-bound) and decode (memory-bound) have opposite hardware appetites. Splitting them across nodes and streaming the KV cache between them lets you right-size each stage independently - better time-to-first-token without over-provisioning decode. The full method is in disaggregated prefill-decode LLM serving. VP metric: p99 latency per dollar.

Replica compaction

When each deployment scales on its own, "resource fragmentation is inevitable as traffic changes and the cluster scales up and down." Compaction repacks replicas onto fewer nodes so you release the stragglers. On per-second serverless billing, released nodes stop costing money immediately. VP metric: utilization, i.e. spend you can actually cut. Teams running many models per node should also read multiple models on one GPU.

Build vs. buy: a decision framework

DimensionOperate (DIY cluster)Consume (serverless clusters)
Node allocationYou build gang schedulingAtomic N-node, managed
RDMA / NCCLYou configure and tuneAuto-configured, up to 6.4 Tbps/node
BillingReserved / hourly, idle burnPer-second, shared pool
FragmentationYour compaction projectPlatform compaction
Time-to-first-clusterWeeks to procure and wireSeconds to acquire
Team focusInfra on-callProduct and models

The honest read: build only if multi-node operation is itself a competitive advantage for your company. For almost everyone else, consuming the fabric frees the exact engineers you most want on the product.

Proof: production deployments

Serverless multi-node is already running real workloads, not demos. Per Modal's Runtime update, "Modal Clusters are already powering post-training at Decagon, pre-training at 1x, and real-time inference at Runway." That spread - post-training, pre-training, and live inference - is the signal a VP should care about: the same primitive covers the full model lifecycle.

On Cerebrium's own serverless GPUs, the single-node numbers that compound across a cluster are concrete. In our benchmarks of Llama 3.1 70B FP8 on one H100, SGLang reached 460 tokens per second at batch size 64 - see the full comparison in benchmarking vLLM, SGLang and TensorRT. And it translates to customer outcomes: we measured a 50% reduction in inference costs for DistilLabs while accuracy rose from 83% to 92% at production scale, detailed in how DistilLabs cut inference costs 50%.

Evaluation checklist

Before you sign anything, confirm the platform:

  • Allocates all N nodes atomically (gang scheduling), not best-effort.

  • Auto-configures RDMA for PyTorch and NCCL - ask for the per-node bandwidth figure.

  • Bills per second from shared capacity, with no idle reservation floor.

  • Supports prefill-decode disaggregation with KV cache transfer across nodes.

  • Compacts replicas to reclaim fragmented capacity automatically.

  • Publishes real cold-start times (ours load a TensorRT-LLM engine into GPU memory in ~10-15s).

  • Has named production references across inference and training.

FAQ

Is multi-node inference only for the very largest models? Primarily, yes - it becomes necessary once weights and KV cache exceed a single node's memory. Below that threshold, single-node serverless is simpler and cheaper.

What makes the inter-node network so important? Sharded models exchange huge tensors every step. Without high-bandwidth RDMA - up to 6.4 Tbps per node - the interconnect becomes the bottleneck and throughput collapses.

Does serverless remove the systems complexity entirely? No. It transfers ownership of scheduling, networking, and compaction to the platform so your team doesn't operate them. The complexity still exists; you just stop paying salaries to manage it.

How does per-second billing beat reserved capacity? You pay only while you hold the nodes and release them in seconds, instead of paying around the clock for a reserved block that sits partly idle.

Can one cluster serve both training and inference? Yes. The same RDMA fabric handles DDP all-reduce and model parallelism for training and KV cache transfer for inference, so you maintain one stack.

How do I prove the cost case to leadership? Start from utilization. Compaction plus per-second billing turns fragmented idle nodes into recoverable spend; pair that with disaggregation for latency-per-dollar. We measured a 50% cost cut for one production customer this way.

Recommendation

Multi-node is non-optional; operating the cluster is optional. Unless running GPU fabric is your differentiator, consume multi-node LLM inference on serverless clusters and point your engineers at the product. If you want to see where the single-node throughput, cold-start, and cost numbers land before you scale out, start with our Llama 3.1 benchmarks and the serverless platform overview - then talk to us about where multi-node fits your roadmap.

Frequently asked questions

Is multi-node inference only for the very largest models?
Primarily, yes - it becomes necessary once weights and KV cache exceed a single node's memory. Below that threshold, single-node serverless is simpler and cheaper.
What makes the inter-node network so important?
Sharded models exchange huge tensors every step. Without high-bandwidth RDMA - up to 6.4 Tbps per node - the interconnect becomes the bottleneck and throughput collapses.
Does serverless remove the systems complexity entirely?
No. It transfers ownership of scheduling, networking, and compaction to the platform so your team doesn't operate them. The complexity still exists; you just stop paying salaries to manage it.
How does per-second billing beat reserved capacity?
You pay only while you hold the nodes and release them in seconds, instead of paying around the clock for a reserved block that sits partly idle.
Can one cluster serve both training and inference?
Yes. The same RDMA fabric handles DDP all-reduce and model parallelism for training and KV cache transfer for inference, so you maintain one stack.
How do I prove the cost case to leadership?
Start from utilization. Compaction plus per-second billing turns fragmented idle nodes into recoverable spend; pair that with disaggregation for latency-per-dollar. We measured a 50% cost cut for one production customer this way.

Get started with Cerebrium

Deploy AI models on serverless GPUs in minutes, with no infrastructure to manage. Start for free and pay only for the compute you use.

Sign up free

Sources

  1. modal.com
    “Communication between nodes via InfiniBand verbs at up to 6.4 Tbps, automatically configured for PyTorch and NCCL”

    RDMA networking capability for the automatic RDMA section and the comparison table.

  2. modal.com
    “the atomic unit for scheduling expands from a single node to N nodes.”

    Gang scheduling section — atomic N-node allocation.

  3. modal.com
    “Your code running on a cluster within seconds of requesting it, billed by the second”

    Per-second billing on shared serverless capacity section.

  4. modal.com
    “training requires it for DDP all-reduce and all types of model parallelism; inference requires it for KV cache transfer when doing PD disaggregation”

    Supports the FAQ answer on one fabric serving both training and inference.

  5. modal.com
    “Modal Clusters are already powering post-training at Decagon, pre-training at 1x, and real-time inference at Runway.”

    Proof section — named production deployments.

  6. anyscale.com
    “Because scaling only considers a single deployment at a time, resource fragmentation is inevitable as traffic changes and the cluster scales up and down.”

    Replica compaction architectural lever section.

  7. cerebrium.ai
    “As for throughput use cases, it seems SGLang was a clear winner achieving a throughput of 460 tokens per second on a batch size of 64.”

    First-hand Cerebrium throughput number in the proof section.

  8. cerebrium.ai
    “With Cerebrium, we reduced their inference costs by 50%, and increased accuracy from 83% to 92%, while keeping latency and reliability consistent at production scale.”

    First-hand customer cost outcome cited in proof section and FAQ.

  9. cerebrium.ai
    “This code will run on every cold start and takes roughly ~10-15s to load the model into GPU memory.”

    Cold-start figure cited in the evaluation checklist.


Related resources

See all
A set of concentric rings forming a stylised target or node icon, with four small triangular tick marks pointing inward from the top, bottom, left, and right, suggesting multiple interchangeable endpoints converging on one core.
Replicate Alternatives for Production Image & LLM Serving
A single stylised robot or AI head icon, face-on, shown as a rounded square outline with two circular eyes, a triangular nose marker, and small limb-like attachment points on either side, rendered as a flat black line drawing.
Llama 405B Serverless GPU Inference Cost: An Honest Model
A single server-rack module with three stacked slots, flanked by two arrows: one pointing up-left into the rack and one pointing down-right away from it, suggesting replicas scaling up and down automatically.
Autoscaling Llama 70B: Cutting Replica Startup Latency