OpenAI-Compatible Endpoints for Open-Source LLMs

Connor Blier
Founding GTM

A single settings-gear icon combined with a plus sign at its centre, symbolising a configurable, extensible API endpoint.

OpenAI-Compatible Endpoints for Open-Source LLMs

An OpenAI-compatible endpoint for open-source LLMs on serverless GPU lets you keep your existing OpenAI SDK code and swap only the base_url to a self-hosted model. Serving frameworks like vLLM expose the same /v1 schema, while serverless GPU adds scale-to-zero, autoscaling, and pay-per-use billing.

The fastest way to move off a hosted model without rewriting your application is to put your open-source model behind an OpenAI-compatible endpoint. You keep the OpenAI SDK, keep your prompts and tool-calling code, and change one line: the base_url. This guide covers what that compatibility actually buys you, what serverless GPU changes underneath it, the frameworks that expose the schema, and the cost math.

TL;DR: the drop-in base_url swap

An OpenAI-compatible server replicates OpenAI's /v1/chat/completions and /v1/embeddings routes, request/response schema, and auth conventions. Point your existing client at https://your-endpoint/v1, set the model name your server advertises, and existing code runs unchanged. As Spheron notes, vLLM's --served-model-name flag "creates an OpenAI-compatible endpoint your existing SDK code hits without changes."

What "OpenAI-compatible" means, and why it matters

OpenAI never formally standardized its API, but it became the de facto interface for LLMs. An OpenAI-compatible API is "any API that replicates common OpenAI interface, request/response schema, and authentication conventions." That matters for three reasons: it is a drop-in replacement for the hosted API, it enables seamless migration between providers or self-hosted deployments, and it gives consistent integration with the tools and frameworks already built around the OpenAI schema. In practice, "any modern open-source LLM can be served behind an OpenAI-compatible API." No lock-in, and every LangChain, LlamaIndex, or agent framework that speaks OpenAI speaks to your model too.

What serverless GPU changes

Self-hosting used to mean keeping a GPU box running 24/7. Serverless GPU changes the economics: workers scale to zero when idle and scale back up on demand, so you pay per use. RunPod, for example, lets you "set active workers when startup latency matters, or allow flex workers to scale down when traffic is intermittent" - the same tradeoff every serverless GPU platform manages between cost and cold-start latency.

The cost of scale-to-zero is the cold start. In our testing on Cerebrium, loading a TensorRT-LLM Llama 3 8B engine into GPU memory takes "roughly ~10-15s" on each cold start. Keeping a small pool of warm capacity hides that from users when latency is critical. For a deeper treatment of cold starts, see our write-up on serverless GPU cold starts.

Endpoints also come in two shapes. Queue-based endpoints suit asynchronous jobs, batch work, and retryable requests; load-balanced endpoints suit direct HTTP access, streaming, and custom FastAPI or Flask services. Chat completions for interactive apps generally want the load-balanced, streaming pattern.

Frameworks that expose the schema

vLLM ships an OpenAI-compatible server out of the box - the most common way to stand up a /v1 endpoint for a Hugging Face model. Ollama is the lightweight local option and also exposes an OpenAI-compatible surface, useful for development before you deploy. Our DeepSeek-R1 guide walks the full vLLM path end to end on serverless GPU.

Which framework wins depends on your metric. In our benchmarks of Llama 3.1 70B FP8 on a single H100 - 256 input tokens, batch size 1, measured as total roundtrip to the Cerebrium platform - vLLM delivered the lowest time-to-first-token at 123ms, while SGLang was the throughput winner at 460 tokens per second on a batch size of 64. Different frameworks, different sweet spots.

Performance techniques worth knowing

Three techniques do most of the heavy lifting behind these numbers:

  • PagedAttention manages the KV cache in non-contiguous pages, cutting memory waste so you fit more concurrent sequences on one GPU.

  • Continuous batching dynamically schedules requests so the GPU does not stall waiting for the slowest request in a fixed batch - Hugging Face describes it as "processing multiple conversations in parallel and swapping them out when they are done."

  • KV cache reuse keeps already-computed attention state so decoding does not recompute the whole context each token.

Together they are why a single H100 can serve high concurrency at low latency. We push these further in running Llama 3 8B with TensorRT-LLM, where the tutorial achieves "~1700 output tokens per second (FP8) on a single Nvidia A10 instance."

Supported models and licensing

Because the endpoint is model-agnostic, your choice is driven by license and quality, not the API. Popular open-source families that serve cleanly behind an OpenAI-compatible endpoint include Llama, Mistral, Qwen, DeepSeek, Phi, and Gemma. Check each model's license before commercial use - permissive terms differ across families - and see our notes on serving fine-tuned LLMs when you own the checkpoint.

The cost argument

The reason teams do this is usually the bill. Self-hosting Llama 3.3 70B in FP8 on a single H100 works out to roughly $1.67 per million output tokens at $2.40/hr, versus $10 per million for GPT-4o output - a large gap that widens with volume. For a fuller model of throughput-driven economics, see LLM inference cost at scale, and for hardware tradeoffs, H100 vs H200 throughput.

How to choose

Start from your traffic shape. Spiky or low-volume workloads favor scale-to-zero and tolerate a cold start; latency-critical apps warrant warm capacity. Pick vLLM for lowest TTFT, SGLang for peak batch throughput, and validate on your own prompts - vendor specs rarely match your roundtrip numbers. The migration itself is cheap: swap the base_url, keep the SDK, and you own the model. When you outgrow a single provider, that same compatibility is what makes the next move painless. If you are comparing platforms, our serverless provider overview and Modal alternatives guide are good next reads.

Frequently asked questions

What does OpenAI-compatible actually mean?
It means an API that replicates OpenAI's interface, request/response schema, and auth conventions - routes like /v1/chat/completions and /v1/embeddings. OpenAI never formally standardized it, but it became the de facto LLM interface, so compatible servers act as drop-in replacements for your existing OpenAI SDK code.
How do I switch my code to a self-hosted model?
Change the base_url to your endpoint's /v1 URL and set the model name your server advertises. vLLM's --served-model-name flag creates an OpenAI-compatible endpoint your existing SDK code hits without changes.
What is the cold-start cost of scale-to-zero?
It depends on model size and framework. In our testing, loading a TensorRT-LLM Llama 3 8B engine into GPU memory took roughly 10-15 seconds per cold start. Keeping warm workers hides this when latency matters.
Which framework gives the lowest latency?
In our benchmarks of Llama 3.1 70B FP8 on one H100, vLLM had the lowest time-to-first-token at 123ms, while SGLang led on throughput at 460 tokens/sec at batch size 64. Choose based on whether you optimize latency or throughput.
Is self-hosting actually cheaper than OpenAI?
At volume, yes. Serving Llama 3.3 70B in FP8 on a single H100 costs about $1.67 per million output tokens versus $10 per million for GPT-4o output, based on published figures.

Get started with Cerebrium

Deploy AI models on serverless GPUs in minutes, with no infrastructure to manage. Start for free and pay only for the compute you use.

Sign up free

Sources

  1. handbook.modular.com
    “An OpenAI-compatible API is any API that replicates common OpenAI interface, request/response schema, and authentication conventions. While OpenAI didn't formally define this as an industry standard, their API has become the de facto interface for LLMs.”

    Definition of OpenAI-compatible API.

  2. handbook.modular.com
    “Drop-in replacement: Swap out OpenAI's hosted API for your own self-hosted or open-source model, often without changing application code. Seamless migration: Move between providers or self-hosted deployments with minimal disruption. Consistent integration: Maintain compatibility with tools and frameworks that rely on the OpenAI API schema (e.g., `chat/completions`, `embeddings` endpoints).”

    Why compatibility matters.

  3. handbook.modular.com
    “Any modern open-source LLM can be served behind an OpenAI-compatible API, such as”

    Model-agnostic serving.

  4. spheron.network
    “vLLM's `--served-model-name` flag creates an OpenAI-compatible endpoint your existing SDK code hits without changes.”

    The base_url swap pattern.

  5. spheron.network
    “At $2.40/hr for the H100 SXM5 on Spheron (as of March 2026), you're paying about ~$1.67 per million output tokens. OpenAI charges $10/1M for GPT-4o output.”

    Cost comparison.

  6. runpod.io
    “Set active workers when startup latency matters, or allow flex workers to scale down when traffic is intermittent.”

    Scale-to-zero vs warm capacity.

  7. runpod.io
    “RunPod also supports two endpoint patterns. Use a queue-based endpoint for asynchronous jobs, batch work, and requests that benefit from retries. Use a load-balanced endpoint for direct HTTP access, streaming, or a custom FastAPI or Flask service.”

    Queue vs load-balanced endpoints.

  8. huggingface.co
    “One of the most impactful optimizations is **continuous batching**, which attempts to maximize performance by processing multiple conversations in parallel and swapping them out when they are done.”

    Continuous batching.

  9. cerebrium.ai
    “This code will run on every cold start and takes roughly ~10-15s to load the model into GPU memory.”

    Cold-start model load time.

  10. cerebrium.ai
    “In this tutorial we will achieve ~1700 output tokens per second (FP8)on a single Nvidia A10 instance”

    Throughput achieved in tutorial.

  11. cerebrium.ai
    “The following benchmark was done with 256 input tokens on 1xH100 of batch size 1. This was the total roundtrip time of a request made to the Cerebrium platform.”

    Benchmark methodology.

  12. cerebrium.ai
    “As for throughput use cases, it seems SGLang was a clear winner achieving a throughput of 460 tokens per second on a batch size of 64.”

    Peak throughput figure.

  13. cerebrium.ai
    “From our tests, it seems that vLLM is the best framework for use cases looking for the lowest TTFT of 123ms which is extremely quick if you take into consideration network latency.”

    Lowest TTFT figure.


Related resources

See all
A monitor or GPU-card outline containing a grid of differently sized rectangular tiles packed together like windows, representing several separate models sharing the space of one processor.
Multiple Models on One GPU: Pack or Dedicate?
A hexagonal frame like a coin or badge with a plus sign at its centre, symbolizing a per-image cost unit.
FLUX Image API Cost Per Image: 2026 Provider Guide
A single-colour bar chart icon: five vertical bars of increasing height rising left to right along a baseline, with an upward diagonal arrow above the tallest bars pointing to the upper right, symbolizing rising inference throughput.
MI300X vs H200 LLM Inference Throughput