Managing LLM Model Versions: Staging to Production

Connor Blier
Founding GTM

Three stacked ellipse-topped cylinders, like a layered database stack, representing separate versions of a model held one above another.

Managing LLM Model Versions: Staging to Production

Managing model versions across staging and production for LLM inference means treating each version as an immutable, fully-bundled artifact in a registry, pointing mutable aliases at it, and promoting staging to production through shadow and canary stages gated on quality signals, with one-command alias rollback.

Managing model versions across staging and production for LLM inference is less about tracking weights and more about controlling a release boundary. A served LLM behaves differently if you change its tokenizer, its generation config, its serving code, or its inference engine, so the unit you promote has to be the whole bundle, not a checkpoint. This guide is an operator's playbook: build each version as an immutable artifact, promote it through shadow then canary stages gated on quality signals, and keep rollback down to repointing an alias.

Why a model version is more than the weights

The most common LLM-serving mistake is versioning only the weights. As the LLM Serving guide puts it, "The single most common LLM-serving mistake is versioning only the weights. A served model is a bundle. Change any component and outputs can move. Pin all of it or you have not pinned anything." That bundle includes model config, tokenizer, generation config, serving and adapter code, the inference engine version, the quantization recipe, and runtime dependencies.

For an AI application the boundary is even wider. A production release is a release manifest - model, prompt, retrieval, and runtime configuration tested and promoted together. Stack Overflow's engineering blog frames it directly: "For an AI application, a model version is only part of the answer. Inputs, preprocessing, prompts, retrieval, tool contracts, and serving settings can change behavior independently." Pin the manifest, not the file. The exact engine flags you pin matter too; our vLLM server parameter notes show how much throughput shifts when those settings change.

Model registry: immutable identity vs mutable aliases

A model registry is the operational backbone here. Snowflake defines it as "a system for managing identifiable versions of machine learning models and the metadata and relationships needed to operate them through their lifecycle." Each entry "typically includes the model artifact, evaluation metrics, input and output signature, training lineage, ownership, aliases and information about its lifecycle or deployment state."

The critical design choice is separating two layers. Good versioning keeps immutable identity apart from mutable pointers: "Immutable identity - a content hash or an append-only version number that never moves... Mutable pointers (aliases/stages) - human-friendly names like @champion, @production, staging that point at an immutable version and can be repointed during a promotion or rollback." You audit against the hash; you route against the alias. MLflow's registry, for example, "automatically tracks versions of each model, allowing teams to compare iterations, roll back to previous states, and manage multiple versions in parallel (e.g., staging vs. production)."

Promoting staging to production safely

Do not cut traffic over in one step. Start in shadow mode: "duplicate production requests to both the current model (which serves users) and the candidate model (which doesn't). Log both outputs, compare them, and make a promotion decision based on what you observe." You can even front-run deployment by replaying history - running "shadow mode agents on historical production requests before you deploy anything" gives a fast read on regression areas before you touch production infrastructure.

Once shadow mode is clean, move to canary. Route a small share - "start at 1%, sometimes as low as 0.1% for high-stakes applications" - then "gradually increase the canary's traffic share: 1% → 5% → 20% → 50% → 100%" while metrics stay in bounds. The mechanics of wiring this on a GPU endpoint are covered in depth in our canary rollout and rollback walkthrough.

The gate matters as much as the ramp. For LLMs the signal set is broader than errors and latency: a new version "can be technically flawless, with zero errors and lower latency, yet produce subtly worse reasoning, shift tone, or hallucinate more frequently." So monitor error and 5xx counts, P99 latency and GPU OOM events, automated quality scores for hallucination and coherence, LLM-as-a-Judge results, and business signals like task completion and user satisfaction. Pin those gates to the latency budget you already run - see our notes on goodput and latency SLOs.

Rollback with stable, candidate, and previous aliases

Rollback is where immutability pays off. The foundation is that "every model artifact is immutable, content-addressed, and tagged before it touches production," with the registry maintaining three production aliases: stable (currently serving, proven), candidate (under canary or shadow), and previous. Because the bundle is pinned, recovery is one operation - repoint stable at the previous version - with no rebuild and no ambiguity about what you reverted to. The step-by-step rollback flow for a live endpoint is in the same canary and rollback guide.

Runtime model and prompt management

If changing a prompt or swapping a model requires a redeploy, every one of the steps above becomes slow and risky. Externalize model selection and prompts from code: "The ability to update prompts, swap models, or adjust inference parameters without shipping new code is the foundation on which everything else depends." Serve behind a stable contract so the alias swap is invisible to callers - an OpenAI-compatible endpoint lets you move versions underneath without touching client code.

Infrastructure considerations

Shadow and canary both mean running two versions at once, so you need headroom for parallel deployments. In our testing, Cerebrium cold starts land in the 2-4 second range, which keeps a freshly-promoted candidate from adding a long warm-up tax when canary traffic first hits it. For steady dual-version serving you want predictable headroom - our guide to reserved GPU capacity for production LLM inference covers sizing that baseline. Where versions differ only by adapter, you can often avoid doubling hardware by co-locating multiple models on one GPU.

Fine-tuned variants are a version-management problem of their own: each adapter is a distinct entry in the registry with its own bundle and aliases. Our guide to serving fine-tuned LLMs on serverless GPUs walks through keeping those straight. The cost of carrying parallel versions is real but bounded; for reference, our serverless GPU cost breakdown prices a 2xH100 configuration at $0.11604 per minute.

Reference staging-to-production runbook

  1. Build the full bundle (weights, tokenizer, generation config, serving code, engine version, quantization recipe, dependencies) and register it with a content hash.

  2. Replay last week's production traffic through the candidate; have a judge compare against current outputs.

  3. Shadow live traffic to the candidate; serve nothing, log everything, compare.

  4. Canary 1% → 5% → 20% → 50% → 100%, holding at each step until the full signal set stays in bounds.

  5. Promote by repointing the stable alias; keep the prior version tagged previous.

  6. Rollback instantly by repointing stable back to previous if any gate fails.

The short version: pin the whole bundle, route with aliases, gate every promotion on quality signals, and make rollback a one-command alias swap. Memory-snapshot cold starts, covered in our GPU cold start deep dive, keep that swap fast enough to treat as routine.

Frequently asked questions

What exactly should a single LLM model version include?
The whole served bundle, not just weights: model config, tokenizer, generation config, serving and adapter code, inference engine version, quantization recipe, and runtime dependencies. For AI applications, the release boundary also covers the prompt, retrieval, and runtime configuration promoted together as one manifest.
How do aliases differ from version numbers?
A version number or content hash is an immutable identity that never moves - you record it in logs and audit trails. Aliases like stable, candidate, and previous are mutable pointers you repoint at an immutable version during a promotion or rollback. Audit against the hash; route against the alias.
What traffic steps should a canary use for LLMs?
Start small - 1%, or as low as 0.1% for high-stakes apps - then increase 1% to 5% to 20% to 50% to 100%, holding at each step only while error rate, P99 latency, GPU OOM, hallucination and coherence scores, LLM-as-a-Judge results, and business metrics stay within bounds.
How fast can rollback be?
Because each version is an immutable, content-addressed bundle, rollback is a single alias repoint from stable back to previous - no rebuild. The registry maintains stable, candidate, and previous aliases so the version you revert to is never ambiguous.

Get started with Cerebrium

Deploy AI models on serverless GPUs in minutes, with no infrastructure to manage. Start for free and pay only for the compute you use.

Sign up free

Sources

  1. fahimfaisal.info
    “The single most common LLM-serving mistake is versioning only the weights. A served model is a bundle. Change any component and outputs can move. Pin all of it or you have not pinned anything.”

    Supports the section on why a version is the full bundle, not just weights.

  2. fahimfaisal.info
    “Good model versioning gives you both layers, and keeps them separate: 1. Immutable identity — a content hash or an append-only version number that never moves. This is what you record in logs, evals, and audit trails. 2. Mutable pointers (aliases/stages) — human-friendly names like @champion, @production, staging that point at an immutable version and can be repointed during a promotion or rollback.”

    Supports immutable identity vs mutable alias separation.

  3. stackoverflow.blog
    “For an AI application, a model version is only part of the answer. Inputs, preprocessing, prompts, retrieval, tool contracts, and serving settings can change behavior independently. Reliable AI infrastructure needs a release boundary around the components that must work together.”

    Supports the release manifest concept.

  4. snowflake.com
    “A model registry is a system for managing identifiable versions of machine learning models and the metadata and relationships needed to operate them through their lifecycle.”

    Defines the model registry.

  5. snowflake.com
    “Each version typically includes the model artifact, evaluation metrics, input and output signature, training lineage, ownership, aliases and information about its lifecycle or deployment state.”

    Lists what a registry entry contains.

  6. mlflow.org
    “The registry automatically tracks versions of each model, allowing teams to compare iterations, roll back to previous states, and manage multiple versions in parallel (e.g., staging vs. production).”

    Supports parallel staging vs production version management.

  7. tianpan.co
    “Shadow mode is the lowest-risk starting point for any significant LLM change. The idea is simple: duplicate production requests to both the current model (which serves users) and the candidate model (which doesn't). Log both outputs, compare them, and make a promotion decision based on what you observe.”

    Supports the shadow mode step.

  8. tianpan.co
    “One pattern that works well is running shadow mode agents on historical production requests before you deploy anything. Replay last week's traffic through the candidate model and have a judge compare outputs against what the current model produced. This gives you a fast read on regression areas before you even touch production infrastructure.”

    Supports replaying historical traffic before deployment.

  9. tianpan.co
    “route a small percentage of traffic — start at 1%, sometimes as low as 0.1% for high-stakes applications — to the candidate while the rest stays on the baseline. Monitor both cohorts on all metrics. If metrics stay within acceptable bounds, gradually increase the canary's traffic share: 1% → 5% → 20% → 50% → 100%.”

    Supports the canary ramp stages.

  10. mlflow.org
    “A new model version can be technically flawless, with zero errors and lower latency, yet produce subtly worse reasoning, shift tone, or hallucinate more frequently.”

    Supports the broader quality signal set for AI canaries.

  11. valuestreamai.com
    “The foundation of every rollback plan is that every model artifact is immutable, content-addressed, and tagged before it touches production.”

    Supports immutable snapshot foundation of rollback.

  12. launchdarkly.com
    “The ability to update prompts, swap models, or adjust inference parameters without shipping new code is the foundation on which everything else depends. If changing a prompt requires a deployment, then testing variations, gradual rollout, and rollback become slow and risky by default.”

    Supports runtime model and prompt management section.

  13. cerebrium.ai
    “**Serverless CPU/GPU Inference**: With cold start times of 2-4 seconds, its the most performant serverless platform on the market.”

    First-hand cold start figure used in infrastructure section.

  14. cerebrium.ai
    “| Cerebrium | $0.07368 | $0.01572 | 0.02664 | $0.11604 |”

    First-hand per-minute 2xH100 cost cited for parallel-version cost reference.


Related resources

See all
A stylised speaker grille cut into a rounded rectangle, with two curved sound waves radiating from its right edge, suggesting audio being produced quickly from a block of text.
OpenAI TTS Alternatives: Open-Source by TTFA
A set of concentric rings forming a stylised target or node icon, with four small triangular tick marks pointing inward from the top, bottom, left, and right, suggesting multiple interchangeable endpoints converging on one core.
Replicate Alternatives for Production Image & LLM Serving
A diagram of five small server-rack boxes arranged in a cross pattern \u2014 one in the center and four at the corners \u2014 connected to each other by thin lines, resembling a distributed computing cluster.
Multi-Node LLM Inference on Serverless Clusters