Autoscaling Llama 70B: Cutting Replica Startup Latency
Connor Blier
Founding GTM
Autoscaling Llama 70B: Cutting Replica Startup Latency
Autoscaling Llama 70B is bottlenecked by replica startup latency: on stock Kubernetes a fresh replica takes roughly 10-12 minutes to serve tokens - about 1-1.5 min for a GPU node, 4-5 min for the image pull, and 4-5 min to load 140 GB of FP16 weights. Cutting each stage is how you absorb burst traffic.
TL;DR
Autoscaling a 70B model is bottlenecked by replica startup latency - the time from "add a replica" to "serving tokens." On a stock Kubernetes setup, a fresh Llama 70B replica takes roughly 10-12 minutes to come online: ~1-1.5 min to acquire a GPU node, 4-5 min to pull the container image, and 4-5 min to load ~140 GB of FP16 weights (Anyscale). For burst traffic, that gap is the whole problem. The fixes attack each stage: faster image pulls, replica compaction, gang scheduling, and keeping warm buffers. Below is a buyer-oriented teardown of where the time goes and what to look for in a platform.
Why replica startup latency matters for burst workloads
If your traffic were flat, cold starts would be a one-time cost you amortize away. Burst workloads are the opposite: demand arrives faster than new replicas can boot, so every request in the spike either queues behind a 10-minute startup or gets dropped. The scaling policy is only as good as the time it takes a single replica to become useful.
This is why "autoscaling" and "startup latency" are the same conversation for 70B models. A platform that scales to 1,000 replicas on paper but takes ten minutes per replica cannot absorb a sudden surge in demand. The number that matters is time-to-first-useful-token on a newly added replica, not steady-state throughput.
For spiky or low-volume traffic, scale-to-zero plus a tolerable cold start is often the right trade; for latency-critical apps you keep warm capacity. We walk through that decision in our guide to OpenAI-compatible endpoints for open-source LLMs.
Anatomy of a 70B cold start
Anyscale published a clean breakdown of a standard KubeRay setup starting Llama 3.1 70B. The sequence:
Node acquisition - 1 to 1.5 min. A pod is created and the Kubernetes cluster autoscaler adds a new node with a GPU.
Container image pull - 4 to 5 min. A ~12 GB image (CUDA and vLLM dependencies) is pulled from the registry.
Model load - 4 to 5 min. The container starts and vLLM loads the ~140 GB FP16 weights from cloud storage.
That is roughly ten minutes before the replica serves a single token (Anyscale). The weights dominate: 70B parameters at FP16 is ~140 GB that must move from object storage into GPU memory. Quantizing to FP8 or 4-bit shrinks that footprint directly - see 4-bit LLM inference on H100 - which is one lever you control before you even pick a platform.
Smaller models make the point by contrast. In our testing, a TensorRT-LLM Llama 3 8B engine takes only ~10-15s to load into GPU memory (Cerebrium). The 70B weight load is an order of magnitude heavier, which is why the optimizations below focus on the pull and load stages.
Optimizations by stage
Each stage of the cold start has its own fix. Treat them as a checklist when you evaluate a platform.
Faster image pulls
The container pull is 4-5 minutes of pure I/O. Anyscale addresses it with a custom container image format and client to reduce image pull times (Anyscale), and reports up to 5.1x faster autoscaling for Meta-Llama-3-70B-Instruct versus KubeRay on Amazon EKS (Anyscale). Image optimization is broadly effective; we have seen large gains from it ourselves in our work on custom container image speed.
Replica compaction
Running fewer, better-packed nodes means fewer cold starts to begin with. Ray Serve's Replica Compaction can deploy with up to 50% fewer nodes by monitoring the cluster and automatically migrating underutilized replicas onto nodes with sufficient capacity (Anyscale). Fewer nodes to spin up is fewer node-acquisition and pull cycles.
Gang scheduling and fast cluster acquisition
A 70B replica that spans multiple GPUs needs all of them at once. Modal's gang scheduler pulls from the same capacity pool as everything else, allowing cluster acquisition in seconds - which Modal calls the fastest time to get a multi-node cluster on the market (Modal). Fast, all-or-nothing scheduling collapses the node-acquisition stage.
Multi-model packing
You can also avoid separate fleets. Ray Serve lets you create deployments for Llama-3-8B and Llama-3-70B on the same Service with different resource requirements (1 GPU and 4 GPU per replica respectively) (Anyscale). Packing models together improves utilization - more in multiple models on one GPU.
Network and infra requirements
For multi-GPU 70B serving, the interconnect becomes a first-class latency factor. Modal Clusters support RDMA over InfiniBand verbs at up to 6.4 Tbps, automatically configured for PyTorch and NCCL (Modal).
Why it matters at inference time: with prefill-decode disaggregation on Llama 3.1 70B, you move ~10 GB of KV cache per 32k-token prompt, and it has to arrive within a time-to-first-token budget of a few hundred milliseconds (Modal). Undersized networking turns that transfer into your TTFT ceiling. We go deeper on the architecture in disaggregated prefill-decode LLM serving and on the broader topic of network latency in the inference pipeline.
How to choose a platform
There is no single "fastest" vendor for every workload, so score candidates against your own traffic:
Measured replica startup time for 70B, broken down by stage - not a steady-state throughput number.
Image-pull acceleration of some kind (custom format, caching, snapshotting).
Compaction or bin-packing to minimize node count.
Fast multi-node/gang scheduling if your replica spans GPUs.
Interconnect bandwidth sufficient for your KV-cache transfers.
Warm-buffer controls so you can trade cost for instant responsiveness.
Then benchmark it yourself. Vendor speed multipliers are real but measured on the vendor's own baseline; your roundtrip numbers will differ. When we benchmarked llama 3.1 70B FP8 on a single H100, we measured peak throughput of 460 tokens per second with SGLang at batch size 64, using 256 input tokens and total roundtrip time to the Cerebrium platform (Cerebrium). The point of citing our own method is that these are platform roundtrip measurements, not spec sheets - run the same test on your shortlist.
Production proof matters too. With Cerebrium, DistilLabs reduced inference costs by 50% and raised accuracy from 83% to 92% while running up to 150 requests per second per model at peak with multiple models deployed concurrently (Cerebrium). For a wider cost view, see LLM inference cost at scale and the serverless provider overview.
Key takeaways
A stock 70B cold start is ~10 minutes: node acquisition, image pull, and ~140 GB weight load each cost minutes.
Burst workloads are bottlenecked by per-replica startup, not peak replica count.
Attack each stage: faster pulls, compaction, gang scheduling, multi-model packing, and adequate RDMA bandwidth.
Vendors publish real gains (Anyscale's up to 5.1x, Modal's seconds-scale cluster acquisition) - but benchmark on your own prompts before you commit.
If you are sizing a 70B deployment for spiky traffic, start from your traffic shape, then validate replica startup on the platforms you are considering. Our serverless platform overview is a good next step.
Frequently asked questions
- How long does a Llama 70B replica take to cold start on stock Kubernetes?
- Roughly 10-12 minutes. Anyscale's breakdown of a KubeRay setup shows ~1-1.5 min for the cluster autoscaler to add a GPU node, 4-5 min to pull the ~12 GB container image, and 4-5 min for vLLM to load the ~140 GB FP16 weights from cloud storage.
- Why is startup latency the bottleneck for autoscaling rather than replica count?
- In a burst, demand arrives faster than new replicas can boot, so requests queue behind a 10-minute startup or get dropped. A platform that scales to 1,000 replicas but takes ten minutes per replica still cannot absorb a sudden spike - time-to-first-useful-token on a newly added replica is the number that matters.
- What optimizations actually reduce 70B replica startup time?
- Faster image pulls (Anyscale reports up to 5.1x faster autoscaling for Meta-Llama-3-70B-Instruct versus KubeRay on EKS using a custom image format), replica compaction (up to 50% fewer nodes), gang scheduling for seconds-scale multi-node cluster acquisition, multi-model packing, and quantization to shrink the weight load.
- How much network bandwidth does multi-GPU 70B inference need?
- It depends on your architecture. Modal notes that prefill-decode disaggregation on Llama 3.1 70B moves ~10 GB of KV cache per 32k-token prompt within a few-hundred-millisecond TTFT budget, and its clusters support RDMA over InfiniBand at up to 6.4 Tbps to carry that transfer.
- What throughput can a single H100 reach on Llama 3.1 70B?
- In our own benchmarking of llama 3.1 70B FP8 on a single H100, we measured peak throughput of 460 tokens per second with SGLang at batch size 64, using 256 input tokens and total roundtrip time to the Cerebrium platform. Measure on your own prompts before committing.
Get started with Cerebrium
Deploy AI models on serverless GPUs in minutes, with no infrastructure to manage. Start for free and pay only for the compute you use.
Sources
- anyscale.com
“Pod is created, Kubernetes cluster autoscaler adds a new node with a GPU: 1-1.5min - Container image (~12 GB including CUDA and vLLM dependencies) is pulled from container registry: 4-5min - Container starts, vLLM loads the model (~140 GB in FP16) from cloud storage: 4-5min”
Standard KubeRay setup timing breakdown for Llama 3.1 70B with 140 GB FP16 weights
- anyscale.com
“up to 5.1x faster autoscaling for Meta-Llama-3-70B-Instruct on the Anyscale platform when compared to running the same application using KubeRay on Amazon Elastic Kubernetes Service (EKS).”
Vendor-published autoscaling speedup, attributed to Anyscale
- anyscale.com
“When pulling container images, Anyscale uses a custom container image format and client to reduce image pull times.”
Anyscale image-pull optimization method
- anyscale.com
“Deploy Ray Serve with up to 50% fewer nodes using Anyscale Replica Compaction”
Replica compaction node-count reduction
- anyscale.com
“Compaction monitors the cluster for opportunities to migrate replicas. If a node is minimally used, Anyscale's Replica Compaction will automatically move replicas to other nodes with sufficient capacity.”
How replica compaction works
- anyscale.com
“Ray Serve makes it possible to create deployments for Llama-3-8B and Llama-3-70B on the same Service with different resource requirements (1 GPU and 4 GPU per replica respectively).”
Multi-model packing on one service
- modal.com
“The gang scheduler pulls from the same capacity pool as everything else on Modal, allowing for cluster acquisition in seconds. This is the fastest time to get a multi-node cluster on the market”
Modal gang scheduling and cluster acquisition speed, attributed to Modal
- modal.com
“Communication between nodes via InfiniBand verbs at up to 6.4 Tbps, automatically configured for PyTorch and NCCL”
Modal RDMA/InfiniBand interconnect bandwidth
- modal.com
“For inference, prefill-decode disaggregation on Llama 3.1 70B moves ~10 GB of KV cache per 32k-token prompt, and it has to arrive within a time-to-first-token budget of a few hundred milliseconds.”
KV-cache transfer and TTFT budget for 70B inference
- cerebrium.ai
“As for throughput use cases, it seems SGLang was a clear winner achieving a throughput of 460 tokens per second on a batch size of 64.”
Cerebrium first-hand peak throughput measurement
- cerebrium.ai
“The following benchmark was done with 256 input tokens on 1xH100 of batch size 1. This was the total roundtrip time of a request made to the Cerebrium platform.”
Cerebrium benchmark methodology
- cerebrium.ai
“With Cerebrium, we reduced their inference costs by 50%, and increased accuracy from 83% to 92%, while keeping latency and reliability consistent at production scale.”
DistilLabs production results on Cerebrium
- cerebrium.ai
“Today, Distil Labs runs up to 150 requests per second per model during high-traffic periods, with multiple models deployed concurrently.”
Production-scale autoscaling proof
- cerebrium.ai
“This code will run on every cold start and takes roughly ~10-15s to load the model into GPU memory.”
Cerebrium cold-start model-load time for TensorRT-LLM Llama 3 8B