Last updated 2026-09-21 · updated weekly by an automated research job · run #7.
The verdict
The current state of the art, as of 2026-08-18 (confidence vendor-claim):
The current best open-weight TTS model for conversational latency appears to be NVIDIA Magpie TTS, announced for low-latency multilingual voice agents. CosyVoice2-0.5B is also cited for ultra-low latency streaming. Both are newly discovered; hardware fit for 48GB is unconfirmed.
What triggered or contributed to this call:
- 2026-08-18 — NVIDIA releases Magpie TTS, an open-weight model for building low-latency multilingual voice agents. (
discovery,vendor-claim) - 2026-08-18 — CosyVoice2-0.5B is cited as an ultra-low latency streaming TTS model. (
discovery,secondary)
Contenders:
- NVIDIA Magpie TTS — Vendor announcement for low-latency multilingual voice agents with full deployment control.
- CosyVoice2-0.5B — Cited as an ultra-low latency streaming TTS model for real-time applications.
- Chatterbox-Turbo — Claims <150ms streaming latency on consumer GPU.
What this is, and how it works
This page is generated, not written. A scheduled job runs weekly on a machine in a homelab. Each run it:
- Searches the open web for both the products already tracked here and for category-level terms designed to turn up ones we've never heard of.
- Feeds those results to a language model along with everything already on this page, and asks it what is genuinely new. Finding nothing is an acceptable answer, and most runs should find little.
- Writes the result into a versioned JSON dataset (
data/live-research/llm-frontier.json) and commits it. This page is re-rendered from that file.
So: a machine wrote the prose here. Every changelog entry carries at least one source link, and every item and entry carries a confidence label:
primary-source— vendor documentation, a published paper, or a regulator.secondary— reputable press.vendor-claim— marketing or an unreplicated vendor statement.unverified— a single low-quality source, usually auto-discovered.
Any label weaker than primary-source is printed beside the finding, so an unflagged row is a primary source. A tracker that hides its own uncertainty is worse than no tracker.
What this is not. Not medical advice. Not a review site — nothing here has been tested in our hands. No affiliate relationships, no sponsored placements, nothing bought. Vendor claims are attributed to the vendor rather than restated as findings. Where a price, a date, or a regulatory status is unknown, it is left blank instead of guessed.
What's new
16 developments in the last 30 days, newest first.
| Date | Subject | Finding |
|---|---|---|
| Sep 21 | CosyVoice2-0.5B SiliconFlow |
CosyVoice2-0.5B cited as an ultra-low latency streaming TTS model for real-time applications. (secondary) |
| Sep 21 | update NVIDIA Blog |
NVIDIA collaborates with llama.cpp and vLLM communities, claiming up to 1.9x higher throughput on RTX 5090. (vendor-claim) |
| Sep 21 | negative Adobe Help |
Adobe lists deprecation dates for Gemini, OpenAI Sora 2, Runway models in 2026. |
| Sep 14 | update Shattered.io |
NVIDIA announces local AI optimizations for 24GB+ RTX GPUs, claiming up to 1.9x speedup with vLLM and llama.cpp. (vendor-claim) |
| Sep 14 | update THE FUTURE 3D |
Khronos Group announces KHR_gaussian_splatting extension for glTF 2.0 in February 2026, enabling cross-platform GS support. (secondary) |
| Sep 10 | discovery Robotics Business News |
Robbyant open-sources LingBot-World, a high-fidelity world model for embodied AI and real-time simulation. (vendor-claim) |
| Aug 22 | OpenVLA Roboskin.ai |
OpenVLA 7B model card details released, trained on 970,000 real-world robot trajectories from Open X-Embodiment mixture. (secondary) |
| Aug 14 | GLM-5 emergent.sh |
GLM 5.3 launched, built on GLM 5.2 base with post-training gains. (vendor-claim) |
| Jul 23 | negative TrueFoundry Blog |
OpenAI scheduled to shut down 15 listed model entries on July 23, 2026. (secondary) |
| Jun 12 | Kimi K2.6 marktechpost.com |
Moonshot AI releases Kimi K2.7-Code, a coding-focused model. (vendor-claim) |
| May 28 | negative help.openai.com |
OpenAI announces retirement of o3 and GPT-4.5 models from ChatGPT. |
| May 6 | Grok 4.20 (x-ai/grok-4.20) whatllm.org |
Grok 4.3 released in May 2026, representing iterative improvements over Grok 4.20. (secondary) |
| Apr 24 | launch FutureAGI Substack |
DeepSeek releases V4-Pro and V4-Flash, MIT-licensed frontier models with significant cost-performance claims. (secondary) |
| Apr 14 | negative Claude Platform Docs |
Anthropic notifies developers of Claude Sonnet 4 and Claude Opus 4 model retirement on June 15, 2026. |
| Mar 18 | negative TensorOps Blog |
The deprecation of OpenAI's GPT-4.1 forces applications to move to reasoning or open source. (secondary) |
| Mar 13 | discovery Yahoo Finance |
ACE ROBOTICS open-sources Kairos 3.0-4B, a native world model for embodied intelligence. (vendor-claim) |
The last 30 days
The last month saw no new models that change the local deployment picture for the 48GB tier. The most significant updates are vendor claims for existing large open-weight models: Z.ai launched GLM 5.3, reporting (unverified) leading scores on terminal and cybersecurity benchmarks, and Moonshot AI released Kimi K2.7-Code, claiming a 21.8% improvement on its own coding benchmark. For the LiteLLM roster, these are incremental updates to already-tracked 240GB-class models; nothing here prompts a config change for the owned hardware.
The negative news is more concrete and forces action elsewhere. OpenAI, Anthropic, and Adobe have all published deprecation schedules for major models in 2026, including GPT-4.1, Claude Sonnet/Opus 4, and various Gemini variants. This systematic removal of older cloud APIs is a hard push toward either newer reasoning models or open-weights, validating the tracker's focus on self-hostable options.
On the research front, several world model and 3D items saw vendor-claim announcements, including Robbyant's LingBot-World and ACE ROBOTICS's Kairos 3.0-4B, but all are unverified and lack local fit confirmation. Similarly, DeepSeek V4-Pro and V4-Flash were noted in a secondary source as significant MIT-licensed frontier models from April, but they are not yet integrated into the tracked item list with concrete specs.
For a reader maintaining a local proxy, the takeaway is stability: no new models fit the 48GB box, and the roster from August remains current. The pressing work is migrating any lingering dependencies off the cloud models now scheduled for shutdown.
Written 2026-09-21 from the changelog below, not from a fresh search.
Drawn from:
- DeepSeek releases V4-Pro and V4-Flash, MIT-licensed frontier models with significant cost-performance claims. — 2026-09-21, confidence
secondary - CosyVoice2-0.5B cited as an ultra-low latency streaming TTS model for real-time applications. — 2026-09-21, confidence
secondary - ACE ROBOTICS open-sources Kairos 3.0-4B, a native world model for embodied intelligence. — 2026-09-21, confidence
vendor-claim - NVIDIA collaborates with llama.cpp and vLLM communities, claiming up to 1.9x higher throughput on RTX 5090. — 2026-09-21, confidence
vendor-claim - The deprecation of OpenAI's GPT-4.1 forces applications to move to reasoning or open source. — 2026-09-21, confidence
secondary - OpenAI scheduled to shut down 15 listed model entries on July 23, 2026. — 2026-09-21, confidence
secondary - Anthropic notifies developers of Claude Sonnet 4 and Claude Opus 4 model retirement on June 15, 2026. — 2026-09-21, confidence
primary-source - Adobe lists deprecation dates for Gemini, OpenAI Sora 2, Runway models in 2026. — 2026-09-21, confidence
primary-source - Robbyant open-sources LingBot-World, a high-fidelity world model for embodied AI and real-time simulation. — 2026-09-14, confidence
vendor-claim - NVIDIA announces local AI optimizations for 24GB+ RTX GPUs, claiming up to 1.9x speedup with vLLM and llama.cpp. — 2026-09-14, confidence
vendor-claim - Khronos Group announces KHR_gaussian_splatting extension for glTF 2.0 in February 2026, enabling cross-platform GS support. — 2026-09-14, confidence
secondary - OpenVLA 7B model card details released, trained on 970,000 real-world robot trajectories from Open X-Embodiment mixture. — 2026-09-14, confidence
secondary - GLM 5.3 launched, built on GLM 5.2 base with post-training gains. — 2026-09-07, confidence
vendor-claim - Moonshot AI releases Kimi K2.7-Code, a coding-focused model. — 2026-09-07, confidence
vendor-claim - Grok 4.3 released in May 2026, representing iterative improvements over Grok 4.20. — 2026-09-07, confidence
secondary - OpenAI announces retirement of o3 and GPT-4.5 models from ChatGPT. — 2026-09-07, confidence
primary-source
The last year
The last twelve months saw the open-weight frontier consolidate around massive MoE models with 1M+ context windows, but the more practical shift was the maturation of smaller, specialized models that actually fit on owned hardware. MiniMax M3 (428B total, 23B active) landed as the cheapest 1M-context multimodal option at $0.24/$0.96 per M tokens, but it’s a 240GB-class model—aspirational, not deployable. Similarly, GLM-5 saw a vendor-claimed upgrade to version 5.3 with standout benchmarks on Terminal-Bench and CyberGym, but it remains in the same 240GB tier. For the 48GB box you actually own, the meaningful entries are the local-inference fits: gpt-oss-20b at ~16GB remains the best-matched local agentic model, and Qwen3-Coder-30B-A3B provides a purpose-built, 48GB-friendly coding specialist. The Qwen family solidified its position as the community's base model for practical deployment.
In speech and 3D, a slew of unverified claims and discoveries point to a crowded open-weight landscape, but nothing has moved from 'watching' to proven, deployed status. For TTS, claims surfaced about Chatterbox-Turbo (<150ms latency), NVIDIA Magpie TTS, and CosyVoice2-0.5B, all citing ultra-low latency. For STT, Cohere Transcribe and NVIDIA Canary-Qwen 2.5B were noted as leaderboard toppers. In 3D, Hixal3D, Hunyuan3D 2.1, and TRELLIS.2 were all mentioned as available open-weight image-to-3D models, with a secondary source stating the field is now mature with at least eight such models. The confidence for all these is unverified or vendor-claim; none have been integrated or performance-verified for the tracker's purposes.
The most concrete negative findings are retirements and quiet lineages. OpenAI retired its o3 and GPT-4.5 models, and GitHub fully shut down its Models platform. The SmolVLM2 lineage appears frozen since April 2025 with no successor, leaving Qwen3-VL-8B-Instruct as the step-up local vision option. MiniMax M2.7 was flagged by the community as 'open weights, not open source' due to its license, a reminder that 'open-weight' does not guarantee open-source freedom. On the inference side, a secondary source noted that llama.cpp can use roughly 35% less VRAM than vLLM for a 13B model, a useful data point for squeezing models into the 48GB tier.
For proprietary models, the movement was iterative. xAI released Grok 4.3 with lower pricing than 4.20, but no benchmark step change was noted. Kimi K2.7-Code was released as a coding-focused variant, but its reported 21.8% improvement on Kimi's own benchmark is a vendor claim. These updates don't change the calculus: they remain cloud-only, closed-weight options.
The arc of the year is clear: the frontier keeps scaling to sizes that require a small data center (240GB+), but the action for actual deployment is in the 48GB tier and in specialized splits. The roster was already pruned on 2026-08-03, retiring Kimi K2.6 in favor of cheaper, more specialized models like Qwen3-Coder. No new model in the last year has displaced the established local fits—gpt-oss-20b for general agentic work, Qwen3-Coder-30B for coding, and Qwen3-VL-8B for vision. The tracker's LiteLLM config does not need an update based on this period's changes.
Written 2026-09-07 from the changelog below, not from a fresh search.
| Month | Entries |
|---|---|
| September 2026 | 16 |
| August 2026 | 20 |
The ones that mattered:
- 2026-09-21 — DeepSeek releases V4-Pro and V4-Flash, MIT-licensed frontier models with significant cost-performance claims. (
launch,secondary) - 2026-09-21 — The deprecation of OpenAI's GPT-4.1 forces applications to move to reasoning or open source. (
negative,secondary) - 2026-09-21 — OpenAI scheduled to shut down 15 listed model entries on July 23, 2026. (
negative,secondary) - 2026-09-21 — Anthropic notifies developers of Claude Sonnet 4 and Claude Opus 4 model retirement on June 15, 2026. (
negative,primary-source) - 2026-09-21 — Adobe lists deprecation dates for Gemini, OpenAI Sora 2, Runway models in 2026. (
negative,primary-source) - 2026-09-07 — OpenAI announces retirement of o3 and GPT-4.5 models from ChatGPT. (
negative,primary-source) - 2026-08-18 — MiniMax M2.7 flagged as 'open weights, not open source' due to its license. (
negative,secondary) - 2026-08-17 — Kimi K2.5 model deprecated, replaced by K2.6. (
negative,secondary)
All time
The field is now decisively split by hardware-fit and modality. For the 48GB tier — the box that actually exists — the local agentic stack is stable and cheap. gpt-oss-20b remains the anchor: a 21B MoE with native 4-bit inference, tool-calling, and a 131K context, costing ~$0.03/$0.13 per M tokens hosted. It fits in ~16GB, making it the default for any task where cloud latency or cost is unacceptable. For coding, Qwen3-Coder-30B-A3B-Instruct is the direct upgrade, purpose-built for agentic tool-use with a 262K context, fitting in the same tier. For vision, Qwen3-VL-8B-Instruct is the step-change over the dormant SmolVLM2, offering 262K context and decent quality, while Gemma 4 31B IT is the cheapest hosted vision option at $0.09/$0.34 per M tokens, though it’s tight on VRAM at long context. The takeaway: local multimodal reasoning and coding are solved for the owned hardware, at near-cloud pricing.
The open-weight frontier has consolidated around massive MoEs that require a 240GB-class cluster. MiniMax-M3 (428B/23B active) is now the cheapest 1M-context multimodal option at $0.24/$0.96 per M tokens, and it reportedly beat GPT-5.5 on SWE-Bench Pro. GLM-5 (up to 744B/40B active) claims a higher agentic ceiling but costs 2.5x more. For pure coding, Qwen3-Coder (480B/35B active) is cheaper and more specialised than Kimi K2.6, which was retired from the local roster for this reason. DeepSeek-V3.2 remains the estate's default cheap cloud fallback at $0.21/$0.31 per M tokens. The pattern is clear: open-weight labs are competing on price-per-token at the 1M-context, multimodal scale, but all require hardware beyond immediate reach.
What does not work, or is gone: Kimi K2.5 is deprecated. The GitHub Models platform (playground, API) was fully retired on July 30, 2026, removing a once-convenient endpoint. The SmolVLM lineage appears frozen since April 2025, with no successor found. For speech, claims are thin: Chatterbox-Turbo (vendor-claim) promises <150ms TTS latency on consumer GPUs, and Cohere Transcribe (unverified) reportedly topped an ASR leaderboard in March 2026, but neither has been integrated or verified locally. In 3D, Hixal3D (unverified) is a claimed open-weight image-to-3D model from Tencent, but its local requirements are unknown.
The biggest change this window is the arrival of MiniMax-M3, which sets a new price floor for 1M-context multimodal inference and makes the older MiniMax-M2 nearly obsolete. For someone deciding where to spend, the local tier is settled and cost-effective; the frontier tier is a price war among giants you cannot host; and the speech/3D frontiers are still just claims and research announcements.
Written 2026-08-17 from the changelog below, not from a fresh search.
How the field breaks down, by what we're actually tracking:
- open-weights — 11 items: MiniMax-M3, GLM-5, Qwen3-Coder (480B-A35B), Kimi K2.6, DeepSeek-V3.2, MiniMax-M2, Gemma 4 31B IT, GLM-4.6V, Ideogram 4.0, DeepSeek-V4-Pro, DeepSeek-V4-Flash.
- local-inference — 7 items: gpt-oss-20b, gpt-oss-120b, Qwen3-Coder-30B-A3B-Instruct, Qwen3-30B-A3B-Instruct-2507, Qwen3-VL-8B-Instruct, Moondream 3 (preview), SmolVLM2-2.2B-Instruct.
- speech — 6 items: Chatterbox-Turbo, Cohere Transcribe, NVIDIA Magpie TTS, CosyVoice2-0.5B, NVIDIA Canary-Qwen 2.5B, NVIDIA Parakeet-TDT-0.6B-v3.
- world-model — 6 items: OpenVLA, Octo, pi0, Psi0, LingBot-World, ACE ROBOTICS Kairos 3.0-4B.
- 3d — 3 items: Hixal3D, Hunyuan3D 2.1, TRELLIS.2.
- proprietary — 1 item: Grok 4.20 (x-ai/grok-4.20).
Gone quiet or discontinued. Kept here rather than deleted, because a tracker that removes its failures is a product page:
- Qwen3-30B-A3B-Instruct-2507 — 30.5B MoE / 3.3B active, 262K context, $0.04815/$0.1931 per M tokens. Cheap and fast in the cloud and a genuine local fit at ~15-19GB Q4 — one of the few models good on both sides of the line. Local: fits 48GB. Last activity 2026-08-17.
- SmolVLM2-2.2B-Instruct — 2.2B vision-language, the incumbent local vision route. No SmolVLM3 or successor found as of 2026-08-03 — the lineage looks frozen since Apr 2025. Tracked so a revival is noticed. Local: fits 48GB. Last activity 2026-08-17.
Tracked products
Available now. Price and regulatory status are blank where we have not read them on a primary source — they are never inferred.
| Product | Category | Price | Regulatory | Confidence | Last activity | Links |
|---|---|---|---|---|---|---|
| Grok 4.20 (x-ai/grok-4.20) | proprietary | unknown | not established | primary-source |
2026-09-07 | OpenRouter |
| MiniMax-M3 | open-weights | unknown | not established | primary-source |
2026-08-17 | OpenRouter · Hugging Face |
| GLM-5 | open-weights | unknown | not established | primary-source |
2026-09-07 | OpenRouter · Hugging Face |
| Qwen3-Coder (480B-A35B) | open-weights | unknown | not established | primary-source |
2026-08-17 | OpenRouter |
| Kimi K2.6 | open-weights | unknown | not established | secondary |
2026-09-07 | OpenRouter |
| DeepSeek-V3.2 | open-weights | unknown | not established | primary-source |
2026-08-17 | OpenRouter |
| MiniMax-M2 | open-weights | unknown | not established | primary-source |
2026-08-18 | OpenRouter |
| gpt-oss-20b | local-inference | unknown | not established | primary-source |
2026-08-17 | OpenRouter |
| gpt-oss-120b | local-inference | unknown | not established | primary-source |
2026-08-17 | OpenRouter |
| Qwen3-Coder-30B-A3B-Instruct | local-inference | unknown | not established | primary-source |
2026-08-17 | OpenRouter |
| Qwen3-VL-8B-Instruct | local-inference | unknown | not established | primary-source |
2026-08-17 | OpenRouter |
| Gemma 4 31B IT | open-weights | unknown | not established | primary-source |
2026-08-17 | OpenRouter |
| GLM-4.6V | open-weights | unknown | not established | primary-source |
2026-08-17 | OpenRouter |
Upcoming
Announced, no date.
Nothing in this bucket right now.
Coming Soon
Announced with a date, or an open pre-order.
Nothing in this bucket right now.
What We're Watching
Exists, unproven, or newly discovered. This is where auto-discovered items land.
Moondream 3 (preview)
9B total / 2B active MoE with a SigLIP encoder, 32K context — plausibly near-CPU-viable given the active parameter count. The OpenRouter model page 404'd on 2026-08-03, so it may be self-host/HF-only. Local: fits 48GB.
First seen 2026-08-17 · confidence unverified · local-inference
Why it's here: No changelog entry explains this status yet.
Sources: Hugging Face
Chatterbox-Turbo
350M-parameter open-source TTS model from Resemble AI, claiming <150ms streaming latency on consumer GPUs. Local: fits 48GB (claimed).
First seen 2026-08-17 · confidence unverified · speech
Why it's here: No changelog entry explains this status yet.
Sources: BentoML blog article
Cohere Transcribe
2B parameter Apache 2.0 speech-to-text model that reportedly topped the Hugging Face Open ASR Leaderboard in March 2026 with 5.42% WER. Local: unconfirmed.
First seen 2026-08-17 · confidence unverified · speech
Why it's here: No changelog entry explains this status yet.
Sources: MarkTechPost article on open ASR models
Hixal3D
Open-weight image-to-3D model from Tencent Arc Research Lab, accepted to SIGGRAPH 2026. Local: unconfirmed.
First seen 2026-08-17 · confidence unverified · 3d
Why it's here: No changelog entry explains this status yet.
Sources: YouTube video report
NVIDIA Magpie TTS
Open-weight text-to-speech model from NVIDIA for building low-latency multilingual voice agents. Local: unconfirmed.
First seen 2026-08-18 · confidence unverified · speech
Why it's here: No changelog entry explains this status yet.
Sources: Hugging Face Blog Post
CosyVoice2-0.5B
Open-source streaming TTS model cited for ultra-low latency in real-time applications. Local: unconfirmed.
First seen 2026-08-18 · confidence unverified · speech
Why it's here: CosyVoice2-0.5B cited as an ultra-low latency streaming TTS model for real-time applications. (2026-09-21) — A 2026 guide lists CosyVoice2-0.5B as an open-source TTS model offering ultra-low latency streaming, positioning it for real-time applications.
Sources: SiliconFlow Article
NVIDIA Canary-Qwen 2.5B
Speech-to-text model leading the Hugging Face Open ASR Leaderboard with a 5.63% average WER. Local: unconfirmed.
First seen 2026-08-18 · confidence unverified · speech
Why it's here: No changelog entry explains this status yet.
Sources: Gladia Blog
NVIDIA Parakeet-TDT-0.6B-v3
Multilingual speech-to-text model cited as the strongest open-source self-host STT option. Local: unconfirmed.
First seen 2026-08-18 · confidence unverified · speech
Why it's here: No changelog entry explains this status yet.
Sources: Coval Blog
OpenVLA
Open-weight vision-language-action model for robotics. Local: unconfirmed.
First seen 2026-08-18 · confidence unverified · world-model
Why it's here: OpenVLA 7B model card details released, trained on 970,000 real-world robot trajectories from Open X-Embodiment mixture. (2026-09-14) — A model card for OpenVLA 7B shows it was trained on 970,000 real-world robot trajectories from the Open X-Embodiment mixture, including DROID data.
Sources: RoboticsCenter.ai Guide
Octo
Open-weight robot foundation model. Local: unconfirmed.
First seen 2026-08-18 · confidence unverified · world-model
Why it's here: No changelog entry explains this status yet.
Sources: RoboticsCenter.ai Guide
pi0
Open foundation model for robotics (partial release). Local: unconfirmed.
First seen 2026-08-18 · confidence unverified · world-model
Why it's here: No changelog entry explains this status yet.
Sources: RoboticsCenter.ai Guide
Psi0
Open foundation model for universal humanoid loco-manipulation, presented in February 2026. Local: unconfirmed.
First seen 2026-08-18 · confidence unverified · world-model
Why it's here: No changelog entry explains this status yet.
Sources: YouTube - Stanford Robotics Seminar
Hunyuan3D 2.1
Tencent's open-source image-to-3D model with a two-stage pipeline for shape generation and texture painting. Local: unconfirmed.
First seen 2026-08-18 · confidence unverified · 3d
Why it's here: No changelog entry explains this status yet.
Sources: Trellis2.app Blog
Ideogram 4.0
9.3B-parameter open-weight text-to-image model with native 2K output, multilingual in-image text, and JSON prompting; non-commercial license. Local: unconfirmed.
First seen 2026-09-07 · confidence unverified · open-weights
Why it's here: No changelog entry explains this status yet.
Sources: 3DAI Studio blog post
TRELLIS.2
4B-parameter MIT-licensed image-to-3D model using sparse voxel representation (O-Voxel), exports GLB with textures, OBJ, PLY, radiance fields, and Gaussian splats. Local: unconfirmed.
First seen 2026-09-07 · confidence unverified · 3d
Why it's here: No changelog entry explains this status yet.
Sources: Cinevva guide on AI 3D model generators
LingBot-World
Open-weight high-fidelity world model from Robbyant for embodied AI and real-time simulation, generating interactive video from a single image. Local: unconfirmed.
First seen 2026-09-14 · confidence unverified · world-model
Why it's here: No changelog entry explains this status yet.
Sources: Robotics Business News
DeepSeek-V4-Pro
MIT-licensed frontier model released April 2026, cited for significant cost-performance advantage. Local: unconfirmed.
First seen 2026-09-21 · confidence unverified · open-weights
Why it's here: No changelog entry explains this status yet.
Sources: Release mention
DeepSeek-V4-Flash
MIT-licensed frontier model released April 2026, cited for significant cost-performance advantage. Local: unconfirmed.
First seen 2026-09-21 · confidence unverified · open-weights
Why it's here: No changelog entry explains this status yet.
Sources: Release mention
ACE ROBOTICS Kairos 3.0-4B
Open-source native world model for embodied intelligence, announced March 2026. Local: unconfirmed.
First seen 2026-09-21 · confidence unverified · world-model
Why it's here: No changelog entry explains this status yet.
Sources: Press release
Promising
Early-stage, but the evidence or the approach is genuinely interesting. The only editorial bucket on this page — an item only lands here with a reason recorded in the changelog.
Nothing in this bucket right now.
Open questions
Publishing what we don't know is the point. These are things the job is actively watching for; when one gets answered it becomes a changelog entry and moves down here to the answered list.
- What is the current best open-weight TTS model that runs at conversational latency inside 48GB, and how does it compare to hosted alternatives on naturalness and streaming? — open since 2026-08-17.
- What is the current best open-weight streaming STT model for 48GB, and has anything displaced the Whisper lineage on word error rate at real-time factor? — open since 2026-08-17.
- Has any interactive world model or embodied prediction model shipped downloadable weights that run on consumer hardware, or is the whole category still API-only demos? — open since 2026-08-17.
- What is the current state of the art for image/video to Gaussian splatting — is there a model that produces a usable splat from casual capture without a photogrammetry pipeline? — open since 2026-08-17.
- Can gpt-oss-120b be made to run usefully in 48GB with offload or a tighter quant, or is the 80GB sizing a hard floor? — open since 2026-08-17.
- Has any quantisation or serving advance since 2026-08 moved a previously-240GB-class model into the 48GB tier? — open since 2026-08-17.
Full changelog
Everything, newest first, grouped by the month we found it. Long by design — it's the receipts.
September 2026
| Date | Subject | Finding |
|---|---|---|
| Sep 21 | CosyVoice2-0.5B SiliconFlow |
CosyVoice2-0.5B cited as an ultra-low latency streaming TTS model for real-time applications. (secondary) |
| Sep 21 | update NVIDIA Blog |
NVIDIA collaborates with llama.cpp and vLLM communities, claiming up to 1.9x higher throughput on RTX 5090. (vendor-claim) |
| Sep 21 | negative Adobe Help |
Adobe lists deprecation dates for Gemini, OpenAI Sora 2, Runway models in 2026. |
| Sep 14 | update Shattered.io |
NVIDIA announces local AI optimizations for 24GB+ RTX GPUs, claiming up to 1.9x speedup with vLLM and llama.cpp. (vendor-claim) |
| Sep 14 | update THE FUTURE 3D |
Khronos Group announces KHR_gaussian_splatting extension for glTF 2.0 in February 2026, enabling cross-platform GS support. (secondary) |
| Sep 10 | discovery Robotics Business News |
Robbyant open-sources LingBot-World, a high-fidelity world model for embodied AI and real-time simulation. (vendor-claim) |
| Aug 22 | OpenVLA Roboskin.ai |
OpenVLA 7B model card details released, trained on 970,000 real-world robot trajectories from Open X-Embodiment mixture. (secondary) |
| Aug 14 | GLM-5 emergent.sh |
GLM 5.3 launched, built on GLM 5.2 base with post-training gains. (vendor-claim) |
| Jul 23 | negative TrueFoundry Blog |
OpenAI scheduled to shut down 15 listed model entries on July 23, 2026. (secondary) |
| Jun 12 | Kimi K2.6 marktechpost.com |
Moonshot AI releases Kimi K2.7-Code, a coding-focused model. (vendor-claim) |
| May 28 | negative help.openai.com |
OpenAI announces retirement of o3 and GPT-4.5 models from ChatGPT. |
| May 6 | Grok 4.20 (x-ai/grok-4.20) whatllm.org |
Grok 4.3 released in May 2026, representing iterative improvements over Grok 4.20. (secondary) |
| Apr 24 | launch FutureAGI Substack |
DeepSeek releases V4-Pro and V4-Flash, MIT-licensed frontier models with significant cost-performance claims. (secondary) |
| Apr 14 | negative Claude Platform Docs |
Anthropic notifies developers of Claude Sonnet 4 and Claude Opus 4 model retirement on June 15, 2026. |
| Mar 18 | negative TensorOps Blog |
The deprecation of OpenAI's GPT-4.1 forces applications to move to reasoning or open source. (secondary) |
| Mar 13 | discovery Yahoo Finance |
ACE ROBOTICS open-sources Kairos 3.0-4B, a native world model for embodied intelligence. (vendor-claim) |
August 2026
| Date | Subject | Finding |
|---|---|---|
| Aug 18 | discovery SiliconFlow Article |
CosyVoice2-0.5B is cited as an ultra-low latency streaming TTS model. (secondary) |
| Aug 18 | discovery Gladia Blog |
NVIDIA Canary-Qwen 2.5B holds #1 position on Hugging Face Open ASR Leaderboard with 5.63% WER. (secondary) |
| Aug 18 | discovery Coval Blog |
NVIDIA Parakeet-TDT-0.6B-v3 cited as strongest open-source self-host STT option. (secondary) |
| Aug 18 | discovery RoboticsCenter.ai Guide |
OpenVLA, Octo, and pi0 are listed as available open-weight vision-language-action models. (secondary) |
| Aug 18 | discovery BuilderAI Tools Blog |
Single image to 3D is mature, with at least eight open-weight models available as of mid-2026. (secondary) |
| Aug 18 | discovery Trellis2.app Blog |
Hunyuan3D 2.1 is Tencent's latest open-source image-to-3D model with a two-stage pipeline. (secondary) |
| Aug 18 | discovery GigaGPU Article |
llama.cpp uses roughly 35% less VRAM than vLLM for a 13B model at equivalent quality on an RTX 5090. (secondary) |
| Aug 17 | discovery marktechpost.com |
Cohere's Transcribe model topped Open ASR Leaderboard in March 2026. (secondary) |
| Aug 17 | discovery bentoml.com |
Chatterbox-Turbo TTS model claims low-latency streaming on consumer GPU. (vendor-claim) |
| Aug 17 | discovery YouTube video on Pixal3D/Hixal3D |
Tencent Arc Research Lab announced Hixal3D, an open-weight image-to-3D model. (vendor-claim) |
| Aug 14 | update Hugging Face Blog |
Hugging Face publishes 'State of Open Models: Summer 2026 Observations' report. |
| Aug 11 | discovery Hugging Face Blog |
NVIDIA releases Magpie TTS, an open-weight model for building low-latency multilingual voice agents. (vendor-claim) |
| Jul 12 | Kimi K2.6 tinker-docs.thinkingmachines.ai |
Kimi K2.5 model deprecated, replaced by K2.6. (secondary) |
| Jul 1 | negative GitHub Changelog |
GitHub Models platform fully retired on July 30, 2026. |
| Jun 1 | MiniMax-M3 Medium article on MiniMax M3 |
MiniMax M3 released with 1M context, native multimodal input, and 59.0% on SWE-Bench Pro. (secondary) |
| Apr 30 | Grok 4.20 (x-ai/grok-4.20) artificialanalysis.ai |
xAI launched Grok 4.3 with lower pricing than Grok 4.20. (secondary) |
| Apr 20 | Kimi K2.6 discretestack.com |
Kimi K2.6 released with improved agentic coding benchmarks. (secondary) |
| Apr 12 | MiniMax-M2 Insights.marvin-42.com |
MiniMax M2.7 flagged as 'open weights, not open source' due to its license. (secondary) |
| Mar 27 | discovery EurekAlert news release |
RoVid-X, a 4M clip dataset for embodied world models, published. (secondary) |
| Feb 20 | discovery youtube.com |
Psi0: An open foundation model for universal humanoid loco-manipulation presented. (vendor-claim) |
Dreamlab Live Research uses autonomous systems to track the state of the art in fields of interest. New trackers appear as the interests do.
Generated from a versioned dataset in a git repository: every state this page has ever been in is a commit, and a bad run is revertible. The dataset is the source; this page is build output and is not itself edited.
- Dataset:
data/live-research/llm-frontier.json(schema version 2) - Runs completed: 7 · cadence: weekly
- Items tracked: 34 · changelog entries: 36
- Distinct sources seen: 281
- Last run recorded: 2026-09-21T05:45:00Z (status:
ok)