Multi-Node LLM Inference on Serverless Clusters
Connor Blier
Founding GTM
Multi-Node LLM Inference on Serverless Clusters
Multi-node LLM inference on serverless clusters lets you run models too large for one node without operating the GPU fabric yourself. The real decision is not whether to go multi-node, but whether to build and babysit the cluster or consume it as a billed-by-the-second serverless primitive.
Multi-node LLM inference on serverless clusters lets you serve models that no longer fit in a single node's memory without standing up and operating your own GPU fabric. For an engineering leader, the decision is rarely whether to go multi-node - frontier and large open-weight models force your hand. The decision is whether your team should operate the interconnect, scheduler, and autoscaler, or consume them as a managed, per-second serverless primitive. This guide frames that build-vs-buy call in VP terms: cost, reliability, and where your engineers spend their week.
Why multi-node inference is now unavoidable
The largest open-weight and frontier models exceed the memory of a single GPU node. Once weights plus KV cache cross that ceiling, you must shard the model across nodes and run tensor, pipeline, or expert parallelism - which only works if those nodes talk to each other over a high-bandwidth, low-latency fabric. This is no longer an exotic training-only requirement: inference needs it too, specifically for KV cache transfer when you split prefill and decode across machines.
That is the uncomfortable part for a VP of Engineering. Single-node serverless was a clean abstraction. Multi-node reintroduces the hard systems problem - scheduling, networking, failure domains - that serverless was supposed to hide. The question is who absorbs it.
The hidden cost of the DIY cluster path
Building your own multi-node inference cluster looks like a procurement exercise and turns into a staffing commitment. The line items that do not appear in the GPU quote:
Interconnect engineering. RDMA over InfiniBand or RoCE has to be configured, tuned, and kept working for PyTorch and NCCL. This is specialist work, and it is load-bearing for every request.
Gang scheduling. You need all N nodes allocated together or the job cannot start. Off-the-shelf schedulers assume a single-node unit; you will be extending one.
Idle burn on reservations. Reserved or hourly cluster capacity bills whether or not traffic is flowing. Fragmented, partially-used nodes are pure waste.
Autoscaling and compaction. Scaling each model independently fragments the cluster; reclaiming that capacity is its own ongoing project.
On-call surface. Every layer above is now something your team pages on at 3 a.m.
Much of this overlaps with the operational burden we described in Serverless Training for LLMs: what it actually means - the inference version simply has tighter latency SLOs. If you are weighing hyperscaler DIY against managed options, our AWS alternatives for AI workloads breakdown is a useful companion.
What "serverless" changes for multi-node
Serverless does not remove the systems problem; it moves ownership of it to the platform. Three mechanics matter most.
Gang scheduling and atomic N-node allocation
A multi-node job is useless until every node is ready. The scheduler's atomic unit has to expand from one node to N. As the Modal Clusters announcement puts it, "the atomic unit for scheduling expands from a single node to N nodes." On a serverless platform this is the vendor's problem, not your Kubernetes operator's.
Automatic RDMA for PyTorch and NCCL
High-bandwidth node-to-node communication is the difference between a working shard and a stalled one. Managed clusters can wire this up for you - communication "via InfiniBand verbs at up to 6.4 Tbps, automatically configured for PyTorch and NCCL." The same fabric serves both training's all-reduce and inference's KV cache transfer during prefill-decode disaggregation, so you are not maintaining two networking stacks.
Per-second billing on shared serverless capacity
The economic unlock is that you draw nodes from a shared pool and pay only while you hold them - "your code running on a cluster within seconds of requesting it, billed by the second." That is structurally different from a reserved block you pay for around the clock. If reserved capacity is still on your table for baseline load, weigh it against on-demand in our note on reserved GPU capacity for production LLM inference, and model the totals with LLM inference cost at scale.
Architectural levers that map to VP metrics
Two techniques convert the multi-node primitive into cost and latency wins you can defend in a board deck.
Prefill-decode disaggregation
Prefill (compute-bound) and decode (memory-bound) have opposite hardware appetites. Splitting them across nodes and streaming the KV cache between them lets you right-size each stage independently - better time-to-first-token without over-provisioning decode. The full method is in disaggregated prefill-decode LLM serving. VP metric: p99 latency per dollar.
Replica compaction
When each deployment scales on its own, "resource fragmentation is inevitable as traffic changes and the cluster scales up and down." Compaction repacks replicas onto fewer nodes so you release the stragglers. On per-second serverless billing, released nodes stop costing money immediately. VP metric: utilization, i.e. spend you can actually cut. Teams running many models per node should also read multiple models on one GPU.
Build vs. buy: a decision framework
| Dimension | Operate (DIY cluster) | Consume (serverless clusters) |
|---|---|---|
| Node allocation | You build gang scheduling | Atomic N-node, managed |
| RDMA / NCCL | You configure and tune | Auto-configured, up to 6.4 Tbps/node |
| Billing | Reserved / hourly, idle burn | Per-second, shared pool |
| Fragmentation | Your compaction project | Platform compaction |
| Time-to-first-cluster | Weeks to procure and wire | Seconds to acquire |
| Team focus | Infra on-call | Product and models |
The honest read: build only if multi-node operation is itself a competitive advantage for your company. For almost everyone else, consuming the fabric frees the exact engineers you most want on the product.
Proof: production deployments
Serverless multi-node is already running real workloads, not demos. Per Modal's Runtime update, "Modal Clusters are already powering post-training at Decagon, pre-training at 1x, and real-time inference at Runway." That spread - post-training, pre-training, and live inference - is the signal a VP should care about: the same primitive covers the full model lifecycle.
On Cerebrium's own serverless GPUs, the single-node numbers that compound across a cluster are concrete. In our benchmarks of Llama 3.1 70B FP8 on one H100, SGLang reached 460 tokens per second at batch size 64 - see the full comparison in benchmarking vLLM, SGLang and TensorRT. And it translates to customer outcomes: we measured a 50% reduction in inference costs for DistilLabs while accuracy rose from 83% to 92% at production scale, detailed in how DistilLabs cut inference costs 50%.
Evaluation checklist
Before you sign anything, confirm the platform:
Allocates all N nodes atomically (gang scheduling), not best-effort.
Auto-configures RDMA for PyTorch and NCCL - ask for the per-node bandwidth figure.
Bills per second from shared capacity, with no idle reservation floor.
Supports prefill-decode disaggregation with KV cache transfer across nodes.
Compacts replicas to reclaim fragmented capacity automatically.
Publishes real cold-start times (ours load a TensorRT-LLM engine into GPU memory in ~10-15s).
Has named production references across inference and training.
FAQ
Is multi-node inference only for the very largest models? Primarily, yes - it becomes necessary once weights and KV cache exceed a single node's memory. Below that threshold, single-node serverless is simpler and cheaper.
What makes the inter-node network so important? Sharded models exchange huge tensors every step. Without high-bandwidth RDMA - up to 6.4 Tbps per node - the interconnect becomes the bottleneck and throughput collapses.
Does serverless remove the systems complexity entirely? No. It transfers ownership of scheduling, networking, and compaction to the platform so your team doesn't operate them. The complexity still exists; you just stop paying salaries to manage it.
How does per-second billing beat reserved capacity? You pay only while you hold the nodes and release them in seconds, instead of paying around the clock for a reserved block that sits partly idle.
Can one cluster serve both training and inference? Yes. The same RDMA fabric handles DDP all-reduce and model parallelism for training and KV cache transfer for inference, so you maintain one stack.
How do I prove the cost case to leadership? Start from utilization. Compaction plus per-second billing turns fragmented idle nodes into recoverable spend; pair that with disaggregation for latency-per-dollar. We measured a 50% cost cut for one production customer this way.
Recommendation
Multi-node is non-optional; operating the cluster is optional. Unless running GPU fabric is your differentiator, consume multi-node LLM inference on serverless clusters and point your engineers at the product. If you want to see where the single-node throughput, cold-start, and cost numbers land before you scale out, start with our Llama 3.1 benchmarks and the serverless platform overview - then talk to us about where multi-node fits your roadmap.
Frequently asked questions
- Is multi-node inference only for the very largest models?
- Primarily, yes - it becomes necessary once weights and KV cache exceed a single node's memory. Below that threshold, single-node serverless is simpler and cheaper.
- What makes the inter-node network so important?
- Sharded models exchange huge tensors every step. Without high-bandwidth RDMA - up to 6.4 Tbps per node - the interconnect becomes the bottleneck and throughput collapses.
- Does serverless remove the systems complexity entirely?
- No. It transfers ownership of scheduling, networking, and compaction to the platform so your team doesn't operate them. The complexity still exists; you just stop paying salaries to manage it.
- How does per-second billing beat reserved capacity?
- You pay only while you hold the nodes and release them in seconds, instead of paying around the clock for a reserved block that sits partly idle.
- Can one cluster serve both training and inference?
- Yes. The same RDMA fabric handles DDP all-reduce and model parallelism for training and KV cache transfer for inference, so you maintain one stack.
- How do I prove the cost case to leadership?
- Start from utilization. Compaction plus per-second billing turns fragmented idle nodes into recoverable spend; pair that with disaggregation for latency-per-dollar. We measured a 50% cost cut for one production customer this way.
Get started with Cerebrium
Deploy AI models on serverless GPUs in minutes, with no infrastructure to manage. Start for free and pay only for the compute you use.
Sources
- modal.com
“Communication between nodes via InfiniBand verbs at up to 6.4 Tbps, automatically configured for PyTorch and NCCL”
RDMA networking capability for the automatic RDMA section and the comparison table.
- modal.com
“the atomic unit for scheduling expands from a single node to N nodes.”
Gang scheduling section — atomic N-node allocation.
- modal.com
“Your code running on a cluster within seconds of requesting it, billed by the second”
Per-second billing on shared serverless capacity section.
- modal.com
“training requires it for DDP all-reduce and all types of model parallelism; inference requires it for KV cache transfer when doing PD disaggregation”
Supports the FAQ answer on one fabric serving both training and inference.
- modal.com
“Modal Clusters are already powering post-training at Decagon, pre-training at 1x, and real-time inference at Runway.”
Proof section — named production deployments.
- anyscale.com
“Because scaling only considers a single deployment at a time, resource fragmentation is inevitable as traffic changes and the cluster scales up and down.”
Replica compaction architectural lever section.
- cerebrium.ai
“As for throughput use cases, it seems SGLang was a clear winner achieving a throughput of 460 tokens per second on a batch size of 64.”
First-hand Cerebrium throughput number in the proof section.
- cerebrium.ai
“With Cerebrium, we reduced their inference costs by 50%, and increased accuracy from 83% to 92%, while keeping latency and reliability consistent at production scale.”
First-hand customer cost outcome cited in proof section and FAQ.
- cerebrium.ai
“This code will run on every cold start and takes roughly ~10-15s to load the model into GPU memory.”
Cold-start figure cited in the evaluation checklist.