Modal vs RunPod: LLM Inference Cost & Cold Starts

Connor Blier
Founding GTM

A large four-pointed sparkle star at the centre of the frame, surrounded by two smaller sparkle stars of the same shape near its lower corners, all rendered as solid black shapes suggesting bursts of instant activation.

Modal vs RunPod: LLM Inference Cost & Cold Starts

There is no single cheaper winner between Modal and RunPod for LLM inference. RunPod posts lower headline GPU rates (H100 PCIe at $2.89/hr, A100 at $1.59/hr), while Modal charges more per hour but bills per-second with no minimum. Your real cost is decided by traffic shape and cold-start tail, not the sticker rate.

The question "Modal vs RunPod for LLM inference cost and cold starts" almost never has the answer engineers expect. The GPU rate is the number everyone screenshots, but it is rarely the number that dominates the bill. What dominates is how the platform bills idle time, how quickly a container comes back from zero, and how spiky your traffic is. Get those three right and a "more expensive" GPU can be the cheaper deployment.

This guide keeps it neutral and calculator-first: how each platform bills, the head-to-head rates, cold starts by model size, a worked total-cost model, and a workload-keyed verdict at the end.

TL;DR verdict - pick X when

PickWhen your workload looks like
RunPodSteady, high-utilization traffic where a Pod runs near 24/7, or you want the lowest headline GPU rate and are comfortable managing your own containers and scaling.
ModalSpiky or bursty traffic that benefits from true scale-to-zero and per-second billing, or large models where GPU-memory snapshotting collapses the cold-start tail.
Either (benchmark first)Latency-critical endpoints - the cold-start tail, not the median, decides whether scale-to-zero is safe for you.

The rest of this article is the evidence behind that table.

How each platform bills

Billing model matters more than the hourly rate for anything that isn't pinned at 100% utilization.

Modal is genuinely per-second: its pricing has no minimum billing increment. RunPod Serverless is also per-second, but rounded up to the nearest second - you're billed from when a worker starts until it fully stops. RunPod also sells dedicated Pods (always-on VMs with GPUs), which are billed for as long as the Pod is running whether or not it serves a request.

That split is the whole game. A Pod at a low hourly rate is cheap only if you keep it busy; the moment it idles, you pay full rate for nothing. A per-second serverless function that scales to zero pays nothing between requests but pays a cold-start penalty on the way back up.

GPU pricing head-to-head

GPUModalRunPod
H100$0.001097/sec (~$3.95/hr) SXM5$2.89/hr (H100 PCIe, Secure Cloud Pod)
A100 80GB$0.000694/sec (~$2.50/hr)$1.59/hr (A100 PCIe Pod)

On paper RunPod wins the rate card: its H100 PCIe is $2.89/hr and its A100 is $1.59/hr, versus Modal's $0.001097/sec H100 and $0.000694/sec A100. Note the two H100s aren't the same board - RunPod quotes PCIe, Modal quotes SXM5 - so this is not an apples-to-apples silicon comparison. It's a comparison of what each platform charges you to run an LLM. For a deeper GPU-by-GPU view see our H100 vs H200 throughput analysis and the 2026 GPU Buyer's Guide.

Cold starts by model size and tail

Cold start is the number to compare, because it decides whether scale-to-zero is even available to you. It also splits into two very different regimes.

Container boot is fast on both. Modal reports that containers boot in about one second before your initialization runs. RunPod's FlashBoot advertises sub-200ms cold starts on its product page, and its engineering data shows cold starts as low as 500ms with 95% under 2.3 seconds and 90% under 2s.

Model load is where large LLMs hurt. Booting a container in a second means nothing if you then spend two minutes pulling weights into VRAM. This is the tail that kills naive scale-to-zero. Modal's GPU-memory snapshotting attacks it directly: on Ministral 3, Modal measured a ~10x reduction in median cold start, from ~118s to ~12s. RunPod's FlashBoot caches to keep workers warm and cut the same tail.

For calibration from our own stack: in our testing a TensorRT-LLM Llama 3 8B engine takes ~10-15s to load the model into GPU memory on every cold start - a useful order-of-magnitude for what "model load" costs before any snapshotting trick. If cold starts are your gating concern, our serverless GPU cold starts for voice AI guide walks through the latency-critical case in detail.

TCO model - a worked example

Use this formula rather than the rate card:

monthly_cost = active_GPU_seconds * per_second_rate
             + idle_GPU_seconds  * idle_rate
             + cold_start_seconds * per_second_rate * cold_starts

On a per-second serverless model (Modal, or RunPod Serverless) idle_rate trends to zero because you scale to zero. On a Pod, idle_rate equals the full hourly rate - you pay for every second the box is up.

Scenario A - steady 24/7 traffic, one H100 pinned near 100%. A RunPod Pod at $2.89/hr runs ~$2,081/month. Modal serverless at ~$3.95/hr for the same continuous utilization is ~$2,845/month. Here the always-on Pod wins clearly: there is no idle to reclaim and no cold starts to pay for.

Scenario B - spiky traffic averaging 50 GPUs against a 75-GPU peak. This is Modal's published example: a fixed 75-GPU fleet at $3/GPU-hr costs $5,400/day, while Modal serverless averaging 50 GPUs costs $4,740/day even at the higher $3.95 rate - because you only pay for the GPUs you're actually using. The higher per-hour rate loses to scale-to-zero once utilization drops below roughly the rate ratio.

The crossover is the point: the cheaper deployment flips at the utilization line, not the rate line. Model your own duty cycle before choosing. Our LLM inference cost at scale breakdown carries this further.

Autoscaling & warm-pool controls

Both platforms scale to zero, and both let you buy your way out of cold starts by keeping workers warm.

Modal keeps containers idle for a short window before shutting down - 60 seconds by default, configurable via scaledown_window. RunPod Serverless uses a configurable idle timeout, defaulting to 5 seconds. A longer idle window trades money for fewer cold starts; a shorter one does the reverse. This dial, plus a min-warm-worker count, is where you actually control the cost/latency tradeoff - see how one team tuned it in the DistilLabs case study, where production-grade autoscaling delivered 50% lower inference costs while raising accuracy from 83% to 92%.

Developer experience & migration friction

Modal is code-first: you define images, functions and GPUs in Python and deploy from the SDK, which suits teams that want infrastructure as code. RunPod gives you both a serverless endpoint model and raw Pods, which suits teams that want container-level control and the lowest rate. Migration friction is mostly about how much of your serving stack you hand-roll versus let the platform manage.

Both are worth benchmarking against a managed alternative. In our benchmarks of serving frameworks on a single H100, vLLM delivered the lowest TTFT at 123ms and SGLang the highest throughput at 460 tokens/sec at batch size 64 - a reminder that framework choice can move latency and cost as much as the platform does. If you're weighing more than these two, see our roundups of Modal alternatives and AWS alternatives for AI workloads.

Decision framework recap

  1. Measure your duty cycle first. Above ~70-80% sustained utilization, an always-on RunPod Pod's lower rate usually wins.

  2. If traffic is spiky, price scale-to-zero. Per-second billing plus scale-to-zero can beat a lower rate once idle time appears.

  3. Size the cold-start tail, not the median. Large models need snapshotting or warm pools; budget the seconds you'll pay on every cold start.

  4. Tune the idle window to trade money against cold starts deliberately.

  5. Benchmark on your own model - vendor cold-start numbers are best case.

Frequently asked questions

Is Modal or RunPod cheaper for LLM inference?
Neither wins outright. RunPod posts lower headline rates ($2.89/hr H100 PCIe, $1.59/hr A100), which win on steady 24/7 traffic. Modal's per-second billing with scale-to-zero wins on spiky workloads, where you avoid paying for idle GPUs even at its higher $3.95/hr H100 rate.
How do Modal and RunPod bill differently?
Modal bills per-second with no published minimum increment. RunPod Serverless also bills per-second but rounds up to the nearest second, from worker start to full stop. RunPod additionally sells always-on Pods billed for their entire runtime, which is cheap only at high utilization.
Which has faster cold starts, Modal or RunPod?
Container boot is fast on both: Modal boots containers in about one second, and RunPod's FlashBoot advertises sub-200ms starts, with 95% under 2.3 seconds. The bigger factor is model load. Modal's GPU-memory snapshotting cut one model's median cold start from ~118s to ~12s.
Does the GPU rate decide my total cost?
No. Total cost is active seconds times rate, plus idle time, plus cold-start seconds. On always-on Pods idle equals full rate; on per-second serverless it trends to zero. Your traffic shape and utilization, not the sticker rate, determine which platform is actually cheaper.
When should I choose an always-on Pod over serverless?
Choose a Pod when utilization is high and sustained - traffic that keeps a GPU busy near 24/7. At that duty cycle there is no idle to reclaim and no cold starts to pay for, so RunPod's lower hourly Pod rate produces the lowest bill.
How can I reduce cold starts without paying for idle GPUs?
Tune the idle timeout and keep a small warm pool. Modal defaults to a 60-second scaledown window (configurable); RunPod defaults to a 5-second idle timeout. Longer windows and minimum warm workers cut cold starts at the cost of some idle spend - a deliberate tradeoff you control.

Get started with Cerebrium

Deploy AI models on serverless GPUs in minutes, with no infrastructure to manage. Start for free and pay only for the compute you use.

Sign up free

Sources

  1. modal.com
    “Nvidia H100 SXM5 $0.001097 / sec”

    Modal H100 per-second rate.

  2. runpod.io
    “H100 PCIe 80 GB VRAM 188 GB RAM 16 vCPUs $2.89/hr”

    RunPod H100 PCIe Pod rate.

  3. modal.com
    “Nvidia A100, 80 GB $0.000694 / sec”

    Modal A100 per-second rate.

  4. runpod.io
    “A100 PCIe 80 GB VRAM 117 GB RAM 8 vCPUs $1.59/hr”

    RunPod A100 PCIe Pod rate.

  5. spheron.network
    “Modal's GPU pricing model is genuinely per-second, with no minimum billing increment published on its pricing page.”

    Modal per-second billing with no minimum.

  6. docs.runpod.io
    “Serverless offers pay-per-second pricing with no upfront costs. You're billed from when a worker starts until it fully stops, rounded up to the nearest second.”

    RunPod Serverless per-second, rounded up.

  7. runpod.io
    “We have seen cold-starts as low as 500ms.”

    FlashBoot minimum cold start.

  8. runpod.io
    “95% of our cold-starts are less than 2.3 seconds, and 90% are less than 2s!”

    FlashBoot cold-start tail distribution.

  9. runpod.io
    “Runpod Serverless runs AI inference with sub-200ms FlashBoot cold starts, per-second billing, and scale to zero.”

    RunPod advertised sub-200ms cold start.

  10. modal.com
    “Containers boot in about one second.”

    Modal container boot time.

  11. modal.com
    “We tested this on the 3B version of Ministral 3 and saw an almost 10x reduction in median cold start time, from ~118s to ~12s.”

    Modal snapshotting cold-start reduction.

  12. modal.com
    “Modal containers will remain idle for a short period before shutting down. By default, the maximum idle time is 60 seconds. You can configure this by setting the `scaledown_window` on the [@function](https://modal.com/docs/sdk/py/latest/App#function) decorator.”

    Modal scaledown window control.

  13. docs.runpod.io
    “Idle timeout duration: The time a worker remains active (running) after completing a request, waiting for additional requests before scaling down (default: 5 seconds).”

    RunPod idle timeout default.

  14. modal.com
    “Traditional cloud: $5,400 75 GPUs * 24 hrs * $3 / GPU-hr Modal serverless cloud: $4,740 Avg 50 GPUs * 24 hrs * $3.95 / GPU-hr”

    Modal spiky-workload cost example.

  15. cerebrium.ai
    “This code will run on every cold start and takes roughly ~10-15s to load the model into GPU memory.”

    Cerebrium first-hand model-load cold-start time.

  16. cerebrium.ai
    “With Cerebrium, we reduced their inference costs by 50%, and increased accuracy from 83% to 92%, while keeping latency and reliability consistent at production scale.”

    DistilLabs autoscaling cost/accuracy result.

  17. cerebrium.ai
    “From our tests, it seems that vLLM is the best framework for use cases looking for the lowest TTFT of 123ms which is extremely quick if you take into consideration network latency.”

    Cerebrium-measured lowest TTFT.

  18. cerebrium.ai
    “As for throughput use cases, it seems SGLang was a clear winner achieving a throughput of 460 tokens per second on a batch size of 64.”

    Cerebrium-measured peak throughput.


Related resources

See all
A single microchip rendered as a square outline with pins on all four sides, containing two short vertical bars inside representing a compact, low-bit numeric value.
4-Bit LLM Inference on H100: Throughput, VRAM & Cost
A monitor or GPU-card outline containing a grid of differently sized rectangular tiles packed together like windows, representing several separate models sharing the space of one processor.
Multiple Models on One GPU: Pack or Dedicate?
A single settings-gear icon combined with a plus sign at its centre, symbolising a configurable, extensible API endpoint.
OpenAI-Compatible Endpoints for Open-Source LLMs