OpenAI TTS Alternatives: Open-Source by TTFA
Connor Blier
Founding GTM
OpenAI TTS Alternatives: Open-Source by TTFA
OpenAI is shutting down its TTS API models on January 6, 2027, so you need a migration plan now. Rank open-source alternatives by time-to-first-audio (TTFA), real-time factor, VRAM, and license. Self-hosting an open model on GPU infrastructure consistently beats an external API on latency.
OpenAI confirmed by email that its tts-1, tts-hd, and gpt-4o-mini-tts models will be shut off on January 6, 2027. If any production feature depends on those endpoints, you have a hard deadline to migrate. This guide is for the Lead ML Engineer who has to pick a replacement and defend the choice. The good news: the open-source text-to-speech ecosystem is now mature enough that self-hosting can be faster than the API you are leaving, and the single metric that matters most for interactive voice is time-to-first-audio (TTFA).
Why the migration is urgent
A deprecation date is not a suggestion. Once tts-1, tts-hd, and gpt-4o-mini-tts are retired, every request to them fails. Migrating a voice feature is not a drop-in key swap: you re-validate audio quality, re-tune latency budgets, re-run load tests, and re-certify any compliance posture. Starting now gives you time to benchmark candidates properly rather than scrambling in Q4 2026.
How to rank: the TTFA evaluation framework
For anything conversational, users perceive the gap between the end of their speech and the first sound of the reply. That gap is dominated by time-to-first-audio - the time from sending text to receiving the first playable chunk of audio. Rank candidates on four axes:
Time-to-first-audio (TTFA): the latency your users actually feel. For streaming endpoints this is closely tracked by time-to-first-byte (TTFB).
Real-time factor (RTF): how fast audio is generated relative to its playback duration. Below 1.0 means the model generates faster than real time, which you need to sustain a stream without stutter.
VRAM footprint: determines which GPU you need and how many concurrent sessions fit on it, which drives cost per minute.
License: confirm the model and its weights permit commercial use before you build on them.
We have published our own numbers on several of these models. From our testing, a streaming Orpheus TTS endpoint has a TTFB of ~100ms, which is in the range you want for low-latency voice.
The open-source TTS landscape
The open-source TTS field has grown into a genuine alternative to proprietary APIs, spanning high-fidelity models like Orpheus and Sesame CSM. Orpheus is a state-of-the-art open system from Canopy Labs built on a Llama-3B backbone - the full architecture and deployment walkthrough is in our Orpheus deployment guide, and for the most natural-sounding option we cover deploying Sesame CSM as an API. The point is that you are not short of credible replacements; you are short of a disciplined way to choose between them.
Why latency: cascaded vs direct architectures
TTS rarely runs alone. In a voice agent it sits at the end of a pipeline, and the architecture of that pipeline sets your latency ceiling. Most voice agents are cascaded: "speech-to-text transcribes the other side, an LLM writes a reply, and text-to-speech reads it out." Each hop adds delay - as one writeup of direct audio models puts it, "every stage you skip is latency you don't pay and a failure mode you don't have."
That framing matters even if you keep a cascaded stack, because it tells you where to cut. The biggest savings come from co-locating the stages and self-hosting them. We measured this directly: the TTFB from the Deepgram API is ~250ms whereas hosting it on Cerebrium is ~110ms - a 140ms saving. On the LLM side, OpenAI 4o-mini's time-to-first-token varies from 700ms to 1.5s, while a self-hosted llama-3 model consistently hits ~300ms. For a self-hosted open vision model, our realtime commentator build saw response times 50% faster than Claude or GPT-4o, without rate limits. Put the whole pipeline on one platform and the gains compound: our Ultravox plus Cartesia pipeline achieved end-to-end first-time-to-audio in just 600ms.
Benchmark it yourself
Vendor numbers are measured under ideal conditions. Reproduce them against your own traffic:
Measure TTFA, not total synthesis time. Instrument the moment the first audio chunk arrives, from a client in the region your users live in - network latency is often the hidden tax.
Test under concurrency. A model that is fast for one request can collapse under load. Measure how many concurrent voice sessions fit per GPU before TTFA degrades.
Account for cold starts. Serverless GPU scaling can add a startup penalty to the first request after idle; see cold starts for voice AI.
Convert to cost. Fold VRAM and concurrency into a cost-per-minute figure so you can compare self-hosting against the API you are replacing on the same terms.
Decision guidance by use case
Real-time conversational agents: prioritise the lowest TTFA and a streaming endpoint. A self-hosted open model on dedicated GPU infrastructure gives you both the latency and the freedom from rate limits.
Batch narration (audiobooks, voiceover): TTFA matters less; weight quality, voice range, and throughput instead.
Multilingual or custom voices: confirm language coverage and licensing for voice cloning before committing.
If you want to replicate the OpenAI developer experience, you can front an open model with an OpenAI-compatible endpoint so your application code barely changes.
Migration playbook
Inventory every call to
tts-1,tts-hd, andgpt-4o-mini-ttsand the latency budget each one lives under.Shortlist two or three open models (Orpheus and Sesame CSM are strong starting points) and benchmark TTFA, RTF, and VRAM yourself.
Deploy the winner on a GPU platform; wire it into a Pipecat voice pipeline if you are building a full agent.
Run a canary and keep a rollback path open until the shutdown date.
The deadline is fixed, but the outcome need not be a downgrade. In our benchmarks, self-hosting open-source TTS and its neighbouring pipeline stages beat the proprietary APIs on the latency metric your users actually feel.
Frequently asked questions
- When exactly does OpenAI's TTS API shut down?
- OpenAI notified customers by email that the tts-1, tts-hd, and gpt-4o-mini-tts models will be shut off on January 6, 2027. After that date requests to those models will fail, so plan your migration well in advance.
- What is time-to-first-audio (TTFA) and why does it matter most?
- TTFA is the time from sending text to receiving the first playable chunk of audio. For conversational voice it is the delay users actually perceive, so it is the single most important metric when ranking TTS alternatives. For streaming endpoints it is closely tracked by time-to-first-byte.
- Is open-source TTS actually faster than OpenAI's API?
- It can be. In our testing a streaming Orpheus TTS endpoint reached a TTFB of about 100ms, and self-hosting neighbouring pipeline stages cut latency substantially - for example Deepgram STT dropped from ~250ms on its API to ~110ms when hosted on Cerebrium.
- Should I use a cascaded or direct audio architecture?
- Cascaded stacks chain speech-to-text, an LLM, and text-to-speech, and each hop adds latency. Direct audio-to-audio models remove those steps. Either way, the largest savings come from self-hosting and co-locating the stages; we achieved 600ms end-to-end first-time-to-audio by putting an Ultravox plus Cartesia pipeline on one platform.
Get started with Cerebrium
Deploy AI models on serverless GPUs in minutes, with no infrastructure to manage. Start for free and pay only for the compute you use.
Sources
- community.openai.com
“We received an email saying these models will be shut off on January 6, 2027: - `tts-1` - `tts-hd` - `gpt-4o-mini-tts`”
Establishes the hard migration deadline for OpenAI's TTS models.
- skeptrune.com
“Most voice agents are cascaded: speech-to-text transcribes the other side, an LLM writes a reply, and text-to-speech reads it out. call4me isn't. There is exactly one model on the audio path, GPT-Live, and it hears audio and speaks audio directly.”
Defines the cascaded voice architecture that TTS normally sits inside.
- skeptrune.com
“There's no speech-to-text step, no text-to-speech step, and no resampling. Every stage you skip is latency you don't pay and a failure mode you don't have.”
Explains why removing pipeline stages reduces latency.
- cerebrium.ai
“From our testing this endpoints has a TTFB (Time-to-first-byte) of ~100ms which is perfect for low latency voice applications.”
Cerebrium's own TTFB benchmark for the Orpheus TTS streaming endpoint.
- cerebrium.ai
“Typically, the TTFB (time-to-first-byte) from the Deepgram API is ~250ms whereas the TTFB hosting it locally on Cerebrium is ~110ms which is a 140ms saving.”
First-hand latency comparison of hosted API vs self-hosted STT.
- cerebrium.ai
“The TTFT (time-to-first-token) of OpenAI 4o-mini varies from 700ms - 1.5s whereas if you deploy a llama-3-8b model or llama-3-70b, you can consistently achieve ~300ms.”
First-hand LLM latency comparison supporting the self-hosting argument.
- cerebrium.ai
“Using this pipeline, we are able to achieve an end-to-end latency (First time to audio, in just 600 ms).”
Cerebrium-measured end-to-end TTFA for a co-located voice pipeline.
- cerebrium.ai
“The reason we are using this open-source model over Anthropic's Claude or OpenAI GPT4o is that the response times we get with our model are 50% faster and much more reliable since we aren't being rate limited or susceptible to other user requests.”
First-hand evidence that self-hosted open models beat proprietary APIs on response time and reliability.