Deep|LLM: Open-Weight Models’ Low-Cost Catch-Up Is Unlikely to Last

A succession of strong open-weight models were released this summer. Zhipu launched GLM-5.2 on June 22, followed by GLM-5.3 on August 14, a release focused solely on scaling post-training. On Zhipu’s own evaluations, GLM-5.3 matched or outperformed Fable 5 and GPT-5.6 Sol on several agentic coding tasks. Moonshot AI launched Kimi K3 on July 17. It scored 57 on the Artificial Analysis Intelligence Index, behind only Fable 5 and GPT-5.6 Sol. DeepSeek’s August 13 update to V4 Pro scored 87.9 on Terminal-Bench 2.1, nearly matching Fable 5’s 88 at roughly one-fiftieth of its cost per task. DeepSeek followed with V4.1 Flash in September. Xiaomi released MiMo-V2.6 with open weights on September 22; its Pro variant now leads the open-weight rankings on the Artificial Analysis Intelligence Index.

Many in the industry now put the gap between open-weight and top proprietary models at just a few months. Some argue that distillation is not the main reason open-weight models have kept pace. We agree it is not the only reason, but it has been one of the key ones. What worries investors is whether open-weight developers can keep matching top proprietary models at a fraction of the R&D cost, and if so, how long the proprietary labs’ lead can last. That worry is feeding into AI and compute stocks, as it did after DeepSeek R1 in January 2025.

We do not think open-weight developers can keep catching up this cheaply, and we expect the recent narrowing of the gap to prove temporary, as it did after R1. Distilling capabilities from top proprietary models is getting harder. The latest models from Anthropic, OpenAI and Google no longer expose their full chain of thought as plaintext, and answers and summaries alone provide weaker training signals. Open-weight developers will need to invest more in their own reinforcement learning, synthetic data generation and training environments. All require compute. Over the next few model generations, we expect the leading open-weight models to fall roughly six to twelve months behind top proprietary models again. The same thing happened in 2025: R1 briefly approached o1, then o3, Claude 4 and GPT-5 pulled ahead.

1. Full reasoning now leaves the API only in encrypted form

Anthropic, OpenAI and Google now take a similar approach to hidden reasoning. Their latest models expose a summary of the chain of thought, or no summary at all. Their APIs can also return encrypted reasoning blocks for reuse in later turns: Anthropic uses a signature field, OpenAI uses encrypted_content, and Google uses thought signatures. Developers pass these blocks back unchanged in subsequent requests. The client cannot decrypt them.

Access has tightened in stages. OpenAI has withheld raw reasoning traces since o1 launched in September 2024. Early Gemini reasoning models still exposed them in early 2025, when researchers at Stanford and other institutions reproduced reasoning behavior by fine-tuning on just 1,000 Gemini traces. In May 2025, Gemini 2.5 and Claude 4 switched to summaries. Opus 4.7, released in April 2026, omits even the summary by default. Fable 5.1, released on September 1, does not expose raw reasoning under any setting. Anthropic’s latest models also refuse requests to reveal their hidden reasoning in the answer. Since September 24, Anthropic has charged for refusals classified as reasoning extraction at the model’s standard rates.

Researchers have nevertheless recovered hidden reasoning by exploiting how APIs handle encrypted blocks. An August paper (arxiv.org/abs/2608.09867) showed that, among the compatible models tested, these blocks could be replayed across sessions and between models from the same provider. The researchers passed a stronger model’s encrypted reasoning to a weaker model, then prompted the latter to transcribe the underlying text. This exploited cross-model compatibility without breaking the encryption itself. Using Claude Haiku 4.5 to extract reasoning generated by Claude Opus 4.8 cost an estimated US$720 per 10,000 traces. The example below comes from the study.

Providers subsequently patched the vulnerability, and Anthropic introduced two restrictions with Fable 5.1. Its encrypted thinking blocks can be reused only by Fable 5.1 or newer models; older models discard them. The blocks are also bound to the full preceding conversation, so changing that context invalidates the signature. This context check is enabled by default for API accounts created on or after August 31. Encrypted reasoning is still returned, but for the latest protected models, such as Fable 5.1, there is no publicly known way to extract the original hidden reasoning. The remaining route is reconstruction: generating approximate traces from the question, answer and summary. The next section looks at how well that works.

2. Summaries are a poor substitute for full reasoning traces

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论