Last updated 2026-09-21 · updated weekly by an automated research job · run #8.
The verdicts
This topic judges "best" along 4 axes, each with its own standing verdict and history.
Voice cloning
Judged on: Best fidelity to a specific target voice from a SHORT reference sample (roughly 3-60 seconds, zero-shot or light fine-tune) — timbre, accent and prosody carried over, judged on published speaker-similarity numbers or reproducible side-by-side samples, not vendor adjectives.
The current state of the art for Voice cloning, as of 2026-08-18 (confidence vendor-claim):
Mistral Voxtral TTS currently leads on zero-shot voice cloning based on a reported 68.4% win rate against ElevenLabs Flash v2.5 in human evaluations and a speaker similarity score of 0.628 on SEED-TTS. However, this verdict rests on vendor benchmarks and a single secondary report; independent arena rankings show Voxtral significantly lower on overall quality.
Contenders:
- Mistral Voxtral TTS — Reported highest speaker similarity score (0.628) and win rate in zero-shot cloning eval.
- ElevenLabs — The commercial reference point most other models are compared against for cloning.
- Chatterbox — Open-weight model with zero-shot cloning under MIT license, runs on consumer GPU.
Expressivity
Judged on: Widest and most controllable emotional range — laughter, sighs, breath and other nonverbal vocalisations, plus explicit control over delivery. Inline tag/markup control counts double here: our own stack drives performance with tags like [laugh], [sigh], [gasp], [pause], so a model that exposes that surface beats one whose emotion is prompt-luck.
The current state of the art for Expressivity, as of 2026-08-18 (confidence vendor-claim):
No single model has decisive evidence for the widest controllable emotional range. ElevenLabs, Fish Audio S2, and Bark TTS all document inline tags for non-verbal vocalizations like laughter and sighs, but Bark's are non-deterministic. Hume Octave is positioned around emotional delivery. The facet lacks a clear leader with reproducible control surface evidence.
Contenders:
- ElevenLabs — Documents audio tags ([laughs], [sighs]) for emotional TTS in its v3 model.
- Fish Audio S2 — Documents emotion control tags like (laughing) and (sighing) in its developer guide.
- Inworld Realtime TTS-2 — Documents non-verbal cues (laughter, sighing, coughing, yawning) as bracket controls.
- Hume Octave — Proprietary model positioned specifically around emotional delivery and instruction-following.
Local-runnable
Judged on: Best model that actually runs on hardware we own — the 48GB tier (1-2x RTX 3090) or an Apple Silicon Mac — with open weights and a license that permits use. A cloud-only model can never win this facet no matter how good it sounds.
The current state of the art for Local-runnable, as of 2026-09-14 (confidence vendor-claim):
Breeze TTS 2 emerges as a top contender for local deployment, reported to run on a 12GB GPU with sub-40ms time-to-first-audio and leading the open-weight category on the Artificial Analysis TTS leaderboard. However, its restrictive commercial license is a significant caveat. Kokoro remains the efficiency champion for CPU/low-resource deployment.
What triggered or contributed to this call:
- 2026-09-14 — Breeze TTS 2 reported as #1 open-weight model on Artificial Analysis TTS leaderboard. (
update,vendor-claim) - 2026-09-14 — Breeze TTS 2 detailed specs and VRAM requirement reported. (
update,vendor-claim)
Contenders:
- Breeze TTS 2 — Reported #1 open-weight model on Artificial Analysis leaderboard, runs on 12GB GPU, sub-40ms latency.
- Kokoro — 82M parameters, runs on CPU, high MOS, Apache 2.0 license.
- Chatterbox — Runs on consumer GPU, MIT license, zero-shot cloning, multilingual support.
Dethronement history — Local-runnable
Every verdict this facet has ever retired, and the receipts for retiring it.
Dethroned 2026-09-14 — held since 2026-08-18
Kokoro is the efficiency champion for local deployment, with 82M parameters running comfortably on CPU and achieving high MOS. Chatterbox also runs on a consumer GPU under MIT license. The 'best' depends on the hardware constraint: Kokoro for CPU/very low-resource, Chatterbox for GPU with cloning.
How it was beaten, and how we know the new one is better: Breeze TTS 2 is reported to outperform previous open-weight models on the primary public benchmark (Artificial Analysis TTS leaderboard) while maintaining hardware requirements within the 48GB tier (12GB GPU). This provides a new benchmark leader with competitive latency.
Latency
Judged on: Best for real-time conversation: lowest time-to-first-audio and a streaming, faster-than-realtime factor, measured on a stated setup. Quality still has to be usable — the fastest robot does not win.
The current state of the art for Latency, as of 2026-08-24 (confidence secondary):
Cartesia Sonic 3.6 now leads both Artificial Analysis Speech Arenas, reinforcing its position as the top contender for low-latency, real-time conversational TTS. Vendor claims cite
90ms for full-quality and ~40ms for streaming, though independent measurements suggest higher real-world latencies (166-190ms).
What triggered or contributed to this call:
- 2026-08-24 — Cartesia ships Sonic 3.6, leading both Artificial Analysis Speech Arenas. (
launch,secondary)
Contenders:
- Cartesia Sonic — Leads Artificial Analysis Provider Voice Arena; marketed on very low time-to-first-audio.
- Inworld Realtime TTS-2 — Ranked #2 on Artificial Analysis Provider Voice Arena, positioned for real-time applications.
- Kokoro — Achieves very low real-time factor (RTF 0.03 on GPU) for fast synthesis.
Dethronement history — Latency
Every verdict this facet has ever retired, and the receipts for retiring it.
Dethroned 2026-08-24 — held since 2026-08-18
Cartesia Sonic is the most cited contender for low-latency, real-time conversational TTS, with vendor claims of
90ms for full-quality and ~40ms for streaming models, and it leads some real-time arena rankings. However, independent real-world measurements suggest higher latencies (166-190ms).
How it was beaten, and how we know the new one is better: Cartesia Sonic 3.6's launch and reported leadership on the Artificial Analysis leaderboards provide updated evidence of its low-latency performance, though the claim remains vendor-adjacent.
What this is, and how it works
This page is generated, not written. A scheduled job runs weekly on a machine in a homelab. Each run it:
- Searches the open web for both the products already tracked here and for category-level terms designed to turn up ones we've never heard of.
- Feeds those results to a language model along with everything already on this page, and asks it what is genuinely new. Finding nothing is an acceptable answer, and most runs should find little.
- Writes the result into a versioned JSON dataset (
data/live-research/tts-models.json) and commits it. This page is re-rendered from that file.
So: a machine wrote the prose here. Every changelog entry carries at least one source link, and every item and entry carries a confidence label:
primary-source— vendor documentation, a published paper, or a regulator.secondary— reputable press.vendor-claim— marketing or an unreplicated vendor statement.unverified— a single low-quality source, usually auto-discovered.
Any label weaker than primary-source is printed beside the finding, so an unflagged row is a primary source. A tracker that hides its own uncertainty is worse than no tracker.
What this is not. Not medical advice. Not a review site — nothing here has been tested in our hands. No affiliate relationships, no sponsored placements, nothing bought. Vendor claims are attributed to the vendor rather than restated as findings. Where a price, a date, or a regulatory status is unknown, it is left blank instead of guessed.
What's new
11 developments in the last 30 days, newest first.
| Date | Subject | Finding |
|---|---|---|
| Sep 21 | Breeze TTS 2, TTS Arena Artificial Analysis on X |
Artificial Analysis confirms Breeze TTS 2 as leading open-weight model on Provider Voices Arena. (secondary) |
| Sep 21 | Inworld Realtime TTS-2 inworld.ai |
Inworld documents non-verbal cue support for laughter, breathing, sighing, coughing, and yawning. |
| Sep 21 | NVIDIA Magpie TTS docs.nvidia.com |
NVIDIA Magpie TTS Zeroshot model exposes zero_shot_quality parameter for speed/fidelity trade-off. |
| Sep 14 | Breeze TTS 2 MindStudio blog |
Breeze TTS 2 reported as #1 open-weight model on Artificial Analysis TTS leaderboard. (vendor-claim) |
| Sep 14 | Breeze TTS 2 MindStudio blog |
Breeze TTS 2 detailed specs and VRAM requirement reported. (vendor-claim) |
| Sep 14 | Breeze TTS 2 Wavespeed.ai review |
Breeze TTS 2 weights carry a restrictive license for commercial use. (secondary) |
| Sep 14 | Cartesia Sonic Cartesia Docs changelog |
Cartesia Sonic 3.5 is now generally available. |
| Sep 14 | regulatory BuilderAI blog |
Piper TTS successor repository license changed from MIT to GPL-3.0. (secondary) |
| Sep 7 | NVIDIA Magpie TTS GitHub - mudler/magpie-tts.cpp |
C++/ggml port of NVIDIA Magpie TTS released, enabling local deployment. (unverified) |
| Aug 1 | Cartesia Sonic Invideo.io blog |
Cartesia Sonic 3.5 real-world latency measured at ~166-190ms, vendor claims sub-90ms. (secondary) |
| Mar 1 | launch Spheron Network blog |
Hume AI releases TADA (Text-Acoustic Dual Alignment) model for zero-hallucination long-form synthesis. (secondary) |
The last 30 days
The biggest shift this month is the emergence of a new open-weight leader. According to Artificial Analysis, Breeze TTS 2 now leads open-weight models on the Provider Voices Speech Arena with an Elo of 1,215, surpassing Fish Audio S2 Pro. It’s reported to run on a 12GB GPU, targeting sub-40ms latency. However, its model weights carry a restrictive commercial license, which is a significant caveat for self-hosted production use.
On the proprietary side, Cartesia Sonic 3.5 is now generally available, but real-world latency measurements (~166-190ms) are notably higher than vendor claims of sub-90ms. Inworld Realtime TTS-2 has documented its support for non-verbal cues like laughter and sighing, a key feature for expressive systems. The open-weight NVIDIA Magpie TTS saw a C++/ggml port released, enabling local deployment, and its documentation now details a speed/fidelity trade-off parameter.
Otherwise, the landscape is quiet. The Piper TTS successor moved to a GPL-3.0 license, affecting commercial use. A new open-source benchmark, ClonEval, was published for evaluating voice cloning. For a practical system needing control and local deployment, Breeze TTS 2’s performance is compelling, but its license makes it a non-starter for many commercial projects, leaving Fish Audio S2 Pro or Orpheus as more permissive, if less top-ranked, options.
Written 2026-09-21 from the changelog below, not from a fresh search.
Drawn from:
- Artificial Analysis confirms Breeze TTS 2 as leading open-weight model on Provider Voices Arena. — 2026-09-21, confidence
secondary - Inworld documents non-verbal cue support for laughter, breathing, sighing, coughing, and yawning. — 2026-09-21, confidence
primary-source - NVIDIA Magpie TTS Zeroshot model exposes zero_shot_quality parameter for speed/fidelity trade-off. — 2026-09-21, confidence
primary-source - Breeze TTS 2 reported as #1 open-weight model on Artificial Analysis TTS leaderboard. — 2026-09-14, confidence
vendor-claim - Breeze TTS 2 detailed specs and VRAM requirement reported. — 2026-09-14, confidence
vendor-claim - Breeze TTS 2 weights carry a restrictive license for commercial use. — 2026-09-14, confidence
secondary - Cartesia Sonic 3.5 is now generally available. — 2026-09-14, confidence
primary-source - Piper TTS successor repository license changed from MIT to GPL-3.0. — 2026-09-14, confidence
secondary - Cartesia Sonic 3.5 real-world latency measured at ~166-190ms, vendor claims sub-90ms. — 2026-09-07, confidence
secondary - Hume AI releases TADA (Text-Acoustic Dual Alignment) model for zero-hallucination long-form synthesis. — 2026-09-07, confidence
secondary - C++/ggml port of NVIDIA Magpie TTS released, enabling local deployment. — 2026-09-07, confidence
unverified
The last year
The last year’s biggest shift is the open-weight tier finally landing punches on the proprietary leaders, though with major caveats. Mistral’s Voxtral TTS, a 4B parameter model, claims to beat ElevenLabs Flash v2.5 in zero-shot cloning human evaluations and offers a lower cost per character. It’s reported to run on a single 16GB GPU. This is a vendor claim, but if verified, it directly challenges the long-standing quality monopoly of cloud providers like ElevenLabs for voice cloning. However, for expressivity and control—the ability to reliably inject a [laugh] or [sigh] on cue—the proprietary cloud services still dominate. Cartesia’s Sonic 3.5 and 3.6 lead the Artificial Analysis Controlled Voice Arena, and ElevenLabs documents explicit audio tags for its v3 model. In the open-weight space, models like Fish Audio S2, Orpheus, and Bark support similar tags, but Bark’s are notoriously non-deterministic, often requiring multiple generations to get a usable take. For a live character system, that unreliability is a deal-breaker.
On the latency front, the definition of “real-time” solidified around sub-200ms time-to-first-byte for conversational agents. Cartesia’s Sonic has been the headline contender here, with its 3.6 iteration reportedly leading both standard and real-time leaderboards. However, a real-world measurement from September 2026 put Sonic 3.5’s latency at 166-190ms, notably higher than the vendor’s sub-90ms model-latency claim. This gap between controlled benchmarks and real-world network conditions is the critical detail for anyone building a live system. Inworld’s Realtime TTS-2 also claimed the top spot on the real-time arena earlier in the year, highlighting a competitive, high-stakes race among cloud providers where independent verification of latency is essential.
For local deployment, the landscape expanded beyond the small, low-quality models of previous years. New contenders like Neuphonic NeuTTS Air (748M parameters, runs on CPU) and a C++/ggml port of NVIDIA’s Magpie TTS now offer viable paths for on-device inference without a datacenter GPU. Hardware fit is finally being addressed: Orpheus recommends 6-8GB VRAM, Chatterbox runs on a consumer GPU, and Hume’s open-weight TADA model needs about 2.5GB VRAM. The arrival of a Qwen3-TTS port optimized for Apple Silicon via MLX is a significant step for the Mac ecosystem. These options are still generally behind the cloud leaders in expressivity and cloning fidelity, but they close the gap enough for many practical applications where privacy, cost, or network independence are priorities.
What stalled is clear, unified benchmarking for the axes that matter most in production. While the TTS Arena and Artificial Analysis provide valuable Elo scores, they don’t systematically measure the reliability of expressive tag execution or real-world latency under load. The new ClonEval benchmark for voice cloning is a welcome step toward rigorous evaluation. Legally, the regulatory environment is hardening, with Tennessee’s ELVIS Act setting a precedent for prohibiting unauthorized commercial voice cloning, a risk factor for any cloning service.
The verdict for a live character system today is fragmented. For best-in-class voice cloning and expressivity with a reliable control surface, a cloud provider like Cartesia or ElevenLabs remains the safest bet, but you must budget for API costs and verify their latency claims in your own stack. For a locally hosted option where cloning is key, Voxtral TTS is the most promising claim, but it’s unverified. For pure, low-latency conversational speech where every millisecond counts, the cloud latency race between Cartesia and Inworld is where to look, but assume real-world numbers will be slower than the marketing. If you must run on a laptop or consumer GPU, the new generation of open-weight models like Chatterbox, Orpheus, or the Magpie port are finally capable, but be prepared to sacrifice some naturalness and fine-grained control.
Written 2026-09-07 from the changelog below, not from a fresh search.
| Month | Entries |
|---|---|
| September 2026 | 11 |
| August 2026 | 34 |
The ones that mattered:
- 2026-09-21 — Artificial Analysis confirms Breeze TTS 2 as leading open-weight model on Provider Voices Arena. (
study,secondary) - 2026-09-14 — Breeze TTS 2 weights carry a restrictive license for commercial use. (
negative,secondary) - 2026-09-14 — Piper TTS successor repository license changed from MIT to GPL-3.0. (
regulatory,secondary) - 2026-09-07 — Hume AI releases TADA (Text-Acoustic Dual Alignment) model for zero-hallucination long-form synthesis. (
launch,secondary) - 2026-08-31 — ClonEval, an open voice cloning benchmark, is published. (
study,primary-source) - 2026-08-24 — Cartesia ships Sonic 3.6, leading both Artificial Analysis Speech Arenas. (
launch,secondary) - 2026-08-18 — Fish Speech V1.5 scores ELO 1339 in TTS Arena evaluation. (
study,secondary) - 2026-08-18 — Cartesia Sonic 3.5 leads Controlled Voice Arena with ELO 1122. (
study,secondary)
All time
Today’s map starts with a regulatory shift: Tennessee’s ELVIS Act now prohibits unauthorized commercial AI voice cloning, creating legal risk for any unlicensed cloning service. On the technical front, the open-weight tier has a new heavyweight contender. Mistral’s Voxtral TTS, a 4B parameter model released in March, claims to match or beat ElevenLabs’ quality at roughly half the cost per character; one report states it won 68.4% of zero-shot cloning evaluations against ElevenLabs Flash v2.5. If these claims hold, Voxtral could disrupt the proprietary cloning market, but it remains unverified and its hardware requirements aren’t specified—likely placing it beyond a single consumer GPU.
For expressivity and inline vocal effects, control surfaces are improving but remain messy. ElevenLabs now documents audio tags like [laugh] and [sigh] for its v3 model. Cartesia Sonic 3.5, now generally available, also supports inline generation of laughter and sighs, alongside refined prosody. In the open tier, Fish Audio S2 offers emotion control tags like (laughing), while Bark TTS supports similar tags but applies them non-deterministically, requiring multiple takes for production—making it unfit for reliable live systems. Hume Octave remains a cloud option marketed purely for emotional delivery.
On conversational latency, the real-time TTS arena has a new leader. Inworld’s Realtime TTS-2, launched in May 2026, reportedly leads the Artificial Analysis Realtime TTS Arena, adding natural-language steering. Cartesia Sonic continues to be marketed on very low time-to-first-audio. The benchmark definition for this category is streaming delivery with sub-200ms Time-to-First-Byte.
For local deployment, the practical open-weights options are Chatterbox and Kokoro. Resemble AI’s Chatterbox Multilingual v3, released under MIT license, now covers 25 languages with zero-shot cloning and intensity control, and runs on a consumer GPU. Kokoro remains the choice for CPU-viable or low-resource deployment, but is not a contender for cloning or expressivity.
The TTS Arena leaderboard, a blind comparison, currently ranks SpeechifyAI Simba 3.2 first, followed by Alibaba Qwen-Audio-3.0-TTS-Plus and Google Gemini 3.1 Flash TTS—none of which are in our tracked item list, highlighting a gap in our coverage for general quality leaders.
Pricing context: ElevenLabs Flash/Turbo sits at $50 per million characters, Google Cloud TTS Standard at $4, and Fish Audio S2 Pro at $15. This makes Voxtral’s claimed half-cost proposition financially significant if proven.
The verdicts per axis: for voice cloning, ElevenLabs is the commercial reference but faces a potentially superior unverified challenger in Voxtral. For expressivity, Hume Octave and Cartesia Sonic are cloud contenders with growing tag support; no open-weight model offers deterministic, production-ready tag control. For local deployment, Chatterbox (consumer GPU) and Kokoro (low-resource) lead. For conversational latency, Inworld Realtime TTS-2 is the new benchmark claimant, with Cartesia Sonic as a competing vendor claim.
Written 2026-08-17 from the changelog below, not from a fresh search.
How the field breaks down, by what we're actually tracking:
- open-weights — 15 items: Kokoro, Chatterbox, Mistral Voxtral TTS, Fish Audio S2, Bark TTS, Sesame CSM, Neuphonic NeuTTS Air, Fish Speech V1.5, CosyVoice2-0.5B, IndexTTS-2, Orpheus, NVIDIA Magpie TTS, CtrlSpeech, Breeze TTS 2, Hume TADA.
- proprietary — 4 items: ElevenLabs, Cartesia Sonic, Hume Octave, Inworld Realtime TTS-2.
- benchmark — 1 item: TTS Arena.
- research — 1 item: Microsoft EmoCtrl-TTS.
- runtime — 1 item: RealtimeTTS.
Gone quiet or discontinued. Kept here rather than deleted, because a tracker that removes its failures is a product page:
- Microsoft EmoCtrl-TTS — Research project for emotion-controllable zero-shot TTS with non-verbal vocalizations; no product plans. Last activity 2026-08-18.
Tracked products
Available now. Price and regulatory status are blank where we have not read them on a primary source — they are never inferred.
| Product | Category | Price | Regulatory | Confidence | Last activity | Links |
|---|---|---|---|---|---|---|
| ElevenLabs | proprietary | unknown | not established | vendor-claim |
2026-08-17 | Site |
| Cartesia Sonic | proprietary | unknown | not established | secondary |
2026-09-14 | Site |
| Hume Octave | proprietary | unknown | not established | vendor-claim |
2026-08-17 | Site |
| Kokoro | open-weights | unknown | not established | unverified |
2026-08-17 | Weights |
| Chatterbox | open-weights | unknown | not established | unverified |
2026-08-17 | Repo · chatterbox-multilingual-tts Model by Resemble.AI | NVIDIA NIM |
| TTS Arena | benchmark | unknown | not established | unverified |
2026-09-21 | Arena |
Upcoming
Announced, no date.
Nothing in this bucket right now.
Coming Soon
Announced with a date, or an open pre-order.
Nothing in this bucket right now.
What We're Watching
Exists, unproven, or newly discovered. This is where auto-discovered items land.
Mistral Voxtral TTS
4B parameter open-weight TTS model from Mistral, claimed to match ElevenLabs quality at lower cost, with strong zero-shot cloning scores.
First seen 2026-08-17 · confidence unverified · open-weights
Why it's here: Voxtral TTS scores ELO 1077 and runs on a single GPU with 16GB VRAM. (2026-08-18) — A guide reports Voxtral TTS has an ELO score of 1,077 and is straightforward to run on a single GPU with 16GB VRAM, served through vLLM.
Sources: Mistral Voxtral TTS: Open-Weight Voice AI for Developers
Fish Audio S2
Open-weight TTS model with emotion control tags like (laughing) and (sighing); weights released under a research license requiring separate license for commercial use.
First seen 2026-08-17 · confidence unverified · open-weights
Why it's here: No changelog entry explains this status yet.
Sources: Emotion Control - Fish Audio · Best Local TTS Models 2026: 8 Open-Source Voices Tested
Bark TTS
Open-weight TTS model supporting inline nonverbal tags like [sighs] and [laughs], but tag application is non-deterministic; runs on consumer GPU (e.g., RTX 3090).
First seen 2026-08-17 · confidence unverified · open-weights
Why it's here: No changelog entry explains this status yet.
Sources: Run Bark AI Locally 2026: Setup on Windows, Mac & Linux
Inworld Realtime TTS-2
Proprietary real-time TTS model launched May 2026, leading the Artificial Analysis Realtime TTS Arena, with natural-language steering across 8 expressive dimensions.
First seen 2026-08-17 · confidence unverified · proprietary
Why it's here: Inworld documents non-verbal cue support for laughter, breathing, sighing, coughing, and yawning. (2026-09-21) — Inworld's documentation states its TTS-2 model supports non-verbal cues written as controls for laughter, breathing, throat clearing, sighing, coughing, and yawning.
Sources: Best voice AI / TTS APIs for real-time voice agents (2026 benchmarks)
Sesame CSM
Open-source TTS model released in early 2026, mentioned as a new contender.
First seen 2026-08-18 · confidence unverified · open-weights
Why it's here: Sesame CSM emerges as a new open-source TTS contender in 2026. (2026-08-18) — A roundup article mentions Sesame CSM as a new open-source TTS model released in the first half of 2026.
Sources: Mention in Befreed.ai article
Neuphonic NeuTTS Air
748M-parameter open-source on-device TTS speech LLM, runs on CPU, supports 3-second cloning.
First seen 2026-08-18 · confidence unverified · open-weights
Why it's here: Neuphonic offers NeuTTS Air, an open-source on-device TTS model. (2026-08-18) — Neuphonic provides a 748M-parameter speech LLM that runs in real-time on mid-tier CPUs, supports instant voice cloning from 3-second samples.
Sources: Speechmatics article
Fish Speech V1.5
Open-source TTS model with DualAR architecture, scoring ELO 1339 on TTS Arena; hardware requirements not specified in source.
First seen 2026-08-18 · confidence unverified · open-weights
Why it's here: No changelog entry explains this status yet.
Sources: SiliconFlow article
CosyVoice2-0.5B
Open-weight TTS model cited as a top pick for 2026; hardware requirements not specified in source.
First seen 2026-08-18 · confidence unverified · open-weights
Why it's here: CosyVoice2-0.5B and IndexTTS-2 mentioned as top open-source TTS models for 2026. (2026-08-24) — A roundup article lists CosyVoice2-0.5B and IndexTTS-2 among its top three open-source TTS picks for 2026, citing innovation, performance, and multilingual support.
Sources: SiliconFlow article
IndexTTS-2
Open-weight TTS model cited as a top pick for 2026; hardware requirements not specified in source.
First seen 2026-08-18 · confidence unverified · open-weights
Why it's here: CosyVoice2-0.5B and IndexTTS-2 mentioned as top open-source TTS models for 2026. (2026-08-24) — A roundup article lists CosyVoice2-0.5B and IndexTTS-2 among its top three open-source TTS picks for 2026, citing innovation, performance, and multilingual support.
Sources: SiliconFlow article
Orpheus
Open-weight TTS model supporting emotional tags like
First seen 2026-08-18 · confidence unverified · open-weights
Why it's here: No changelog entry explains this status yet.
Sources: LocalClaw guide
NVIDIA Magpie TTS
Open-weight TTS model from NVIDIA for building low-latency multilingual voice agents; hardware requirements unspecified.
First seen 2026-08-24 · confidence unverified · open-weights
Why it's here: NVIDIA Magpie TTS Zeroshot model exposes zero_shot_quality parameter for speed/fidelity trade-off. (2026-09-21) — NVIDIA's documentation for Magpie TTS Zeroshot model describes a --zero_shot_quality parameter (range 1-40) that controls the trade-off between synthesis speed and voice similarity.
Sources: Announcement Blog
CtrlSpeech
Research framework for coarse-to-fine controllable expressive speech synthesis with explicit prosodic control.
First seen 2026-08-24 · confidence unverified · open-weights
Why it's here: No changelog entry explains this status yet.
Sources: arXiv Paper
RealtimeTTS
A Python library for real-time, streaming text-to-speech conversion, supporting multiple engines including QwenEngine for low-latency conversational speech.
First seen 2026-08-31 · confidence unverified · runtime
Why it's here: No changelog entry explains this status yet.
Sources: GitHub Repository
Hume TADA
Text-Acoustic Dual Alignment model from Hume AI for zero-hallucination long-form synthesis; 1B parameter version needs ~2.5GB VRAM.
First seen 2026-09-07 · confidence unverified · open-weights
Why it's here: No changelog entry explains this status yet.
Sources: Spheron Network blog
Promising
Early-stage, but the evidence or the approach is genuinely interesting. The only editorial bucket on this page — an item only lands here with a reason recorded in the changelog.
Breeze TTS 2
Open-weight TTS model scoring 1215 Elo on Artificial Analysis Speech Arena, claimed to run on a 12GB GPU.
First seen 2026-09-07 · confidence unverified · open-weights
Why it's here: Artificial Analysis confirms Breeze TTS 2 as leading open-weight model on Provider Voices Arena. (2026-09-21) — Artificial Analysis announced Breeze TTS 2 leads open-weight models on the Provider Voices Speech Arena with an Elo of 1,215, surpassing Fish Audio S2 Pro by 90 points and ranking #6 overall. In the Controlled Voices Arena, it ranks #3 among open-weight models with an Elo of 1,002.
Sources: Announcement reference
Open questions
Publishing what we don't know is the point. These are things the job is actively watching for; when one gets answered it becomes a changelog entry and moves down here to the answered list.
- Which model gives the best speaker similarity from a reference sample under 30 seconds, and what is the measured similarity number rather than the marketing claim? — open since 2026-08-17.
- Which models accept inline nonverbal/delivery tags (laughter, sighs, breaths, pauses) as an actual documented control surface, and what is each one's tag vocabulary? — open since 2026-08-17.
- What is the best-sounding open-weight TTS model that runs on a single 24GB GPU or the 48GB tier, and what does it need — VRAM, runtime, quantisation? — open since 2026-08-17.
- What time-to-first-audio and realtime factor are actually achievable locally versus via the fastest cloud API, measured on a stated setup? — open since 2026-08-17.
- Which of the leading open-weight TTS models carry licenses that permit commercial use and voice cloning, and which are non-commercial or consent-gated? — open since 2026-08-17.
Full changelog
Everything, newest first, grouped by the month we found it. Long by design — it's the receipts.
September 2026
| Date | Subject | Finding |
|---|---|---|
| Sep 21 | Breeze TTS 2, TTS Arena Artificial Analysis on X |
Artificial Analysis confirms Breeze TTS 2 as leading open-weight model on Provider Voices Arena. (secondary) |
| Sep 21 | Inworld Realtime TTS-2 inworld.ai |
Inworld documents non-verbal cue support for laughter, breathing, sighing, coughing, and yawning. |
| Sep 21 | NVIDIA Magpie TTS docs.nvidia.com |
NVIDIA Magpie TTS Zeroshot model exposes zero_shot_quality parameter for speed/fidelity trade-off. |
| Sep 14 | Breeze TTS 2 MindStudio blog |
Breeze TTS 2 reported as #1 open-weight model on Artificial Analysis TTS leaderboard. (vendor-claim) |
| Sep 14 | Breeze TTS 2 MindStudio blog |
Breeze TTS 2 detailed specs and VRAM requirement reported. (vendor-claim) |
| Sep 14 | Breeze TTS 2 Wavespeed.ai review |
Breeze TTS 2 weights carry a restrictive license for commercial use. (secondary) |
| Sep 14 | Cartesia Sonic Cartesia Docs changelog |
Cartesia Sonic 3.5 is now generally available. |
| Sep 14 | regulatory BuilderAI blog |
Piper TTS successor repository license changed from MIT to GPL-3.0. (secondary) |
| Sep 7 | NVIDIA Magpie TTS GitHub - mudler/magpie-tts.cpp |
C++/ggml port of NVIDIA Magpie TTS released, enabling local deployment. (unverified) |
| Aug 1 | Cartesia Sonic Invideo.io blog |
Cartesia Sonic 3.5 real-world latency measured at ~166-190ms, vendor claims sub-90ms. (secondary) |
| Mar 1 | launch Spheron Network blog |
Hume AI releases TADA (Text-Acoustic Dual Alignment) model for zero-hallucination long-form synthesis. (secondary) |
August 2026
| Date | Subject | Finding |
|---|---|---|
| Aug 31 | voice-cloning ClonEval Paper on arXiv |
ClonEval, an open voice cloning benchmark, is published. |
| Aug 31 | local-runnable GitHub - qwen3-tts-apple-silicon |
Open-source Qwen3-TTS port optimized for Apple Silicon via MLX framework. (unverified) |
| Aug 31 | latency GitHub - RealtimeTTS |
RealtimeTTS Python library enables low-latency, streaming TTS. (unverified) |
| Aug 24 | CosyVoice2-0.5B, IndexTTS-2 SiliconFlow |
CosyVoice2-0.5B and IndexTTS-2 mentioned as top open-source TTS models for 2026. (secondary) |
| Aug 18 | Cartesia Sonic MarkTechPost |
Cartesia ships Sonic 3.6, leading both Artificial Analysis Speech Arenas. (secondary) |
| Aug 18 | study SiliconFlow article |
Fish Speech V1.5 scores ELO 1339 in TTS Arena evaluation. (secondary) |
| Aug 18 | discovery SiliconFlow article |
CosyVoice2-0.5B and IndexTTS-2 mentioned as top open-source TTS models for 2026. (secondary) |
| Aug 18 | discovery LocalClaw guide |
Orpheus TTS model supports emotional tags like |
| Aug 18 | Mistral Voxtral TTS Pinggy.io guide |
Voxtral TTS scores ELO 1077 and runs on a single GPU with 16GB VRAM. (secondary) |
| Aug 18 | Cartesia Sonic, TTS Arena Artificial Analysis on X |
Cartesia Sonic 3.5 leads Controlled Voice Arena with ELO 1122. (secondary) |
| Aug 18 | Cartesia Sonic youtube.com |
Cartesia deprecates Sonic 2, Sonic Turbo, and October 2025 Sonic 3. (vendor-claim) |
| Aug 18 | Sesame CSM befreed.ai |
Sesame CSM emerges as a new open-source TTS contender in 2026. (secondary) |
| Aug 18 | Neuphonic NeuTTS Air speechmatics.com |
Neuphonic offers NeuTTS Air, an open-source on-device TTS model. (secondary) |
| Aug 18 | update github.com |
Feature request for inline emotion tag support in Qwen3-TTS. (secondary) |
| Aug 18 | Microsoft EmoCtrl-TTS EmoCtrl-TTS - Microsoft Research |
Microsoft Research publishes EmoCtrl-TTS, an emotion-controllable zero-shot TTS. |
| Aug 18 | TTS Arena artificialanalysis.ai |
Artificial Analysis TTS leaderboard shows Cartesia Sonic 3.6 leading. (secondary) |
| Aug 18 | TTS Arena voicearena.com |
Voice Arena leaderboard shows Gemini 3.1 Flash TTS leading. (secondary) |
| Aug 18 | price costgoat.com |
OpenAI TTS API pricing detailed: Standard $15/1M chars, HD $30/1M chars. (secondary) |
| Aug 18 | price cloud.google.com |
Google Cloud TTS pricing lists Gemini 3.1 Flash TTS at $1 per 1M input tokens. |
| Aug 18 | regulatory fredlaw.com |
Federal court dismisses trademark/copyright claims in Lovo voice cloning lawsuit. (secondary) |
| Aug 17 | ElevenLabs elevenlabs.io |
ElevenLabs documents audio tags for emotional TTS, including [laughs], [sighs], and [clears throat]. (vendor-claim) |
| Aug 17 | expressivity Emotion Control - Fish Audio |
Fish Audio documents emotion control tags including (laughing), (sighing), and (groaning). (vendor-claim) |
| Aug 17 | expressivity localaimaster.com |
Bark TTS supports inline tags like [sighs] and [laughs] but application is non-deterministic, requiring multiple takes for production. (secondary) |
| Aug 17 | TTS Arena offlinetts.com |
SpeechifyAI Simba 3.2 leads the TTS Arena leaderboard with an Elo of 1,229, followed by Alibaba Qwen-Audio-3.0-TTS-Plus. (secondary) |
| Aug 17 | latency camb.ai |
Real-time TTS defined as generating and delivering speech via streaming with sub-200ms Time-to-First-Byte for conversational applications. (secondary) |
| Aug 17 | discovery aipricing.guru |
Pricing comparison shows ElevenLabs Flash/Turbo at $50 per 1M characters, Google Cloud TTS Standard at $4, and Fish Audio S2 Pro at $15. (secondary) |
| Aug 17 | regulatory rock.law |
Tennessee's ELVIS Act prohibits unauthorized AI voice cloning for commercial purposes, providing civil remedies. (secondary) |
| Aug 10 | discovery Hugging Face Blog |
NVIDIA announces Magpie TTS, an open-weight model for low-latency multilingual voice agents. |
| Aug 8 | expressivity arXiv |
CtrlSpeech, a coarse-to-fine controllable expressive speech synthesis framework, is proposed. |
| Jun 10 | Chatterbox localaimaster.com |
Resemble AI releases Chatterbox Multilingual v3, covering 25 languages while remaining MIT-licensed. (vendor-claim) |
| May 5 | voice-cloning marktechpost.com |
Voxtral TTS reportedly beats ElevenLabs Flash v2.5 in zero-shot voice cloning human evaluations and speaker similarity scores. (secondary) |
| May 4 | Cartesia Sonic Sonic 3.5 - Cartesia Docs |
Cartesia Sonic 3.5 is generally available, adding refined prosody, wider emotional range, and real-time laughter. (vendor-claim) |
| May 1 | TTS Arena inworld.ai |
Inworld's Realtime TTS-2 research preview leads the Artificial Analysis Realtime TTS Arena as of May 2026. (vendor-claim) |
| Mar 26 | discovery zenvanriel.com |
Mistral releases Voxtral TTS, a 4B parameter open-weight model challenging ElevenLabs on quality and cost. (vendor-claim) |
Dreamlab Live Research uses autonomous systems to track the state of the art in fields of interest. New trackers appear as the interests do.
Generated from a versioned dataset in a git repository: every state this page has ever been in is a commit, and a bad run is revertible. The dataset is the source; this page is build output and is not itself edited.
- Dataset:
data/live-research/tts-models.json(schema version 2) - Runs completed: 8 · cadence: weekly
- Items tracked: 22 · changelog entries: 45
- Distinct sources seen: 216
- Last run recorded: 2026-09-21T05:52:18Z (status:
ok)