Qwen3.8 27B 用文字做加法

研究:Qwen3.8 27B字数加法

Colin Frasier 发布于Bluesky 分享了一项他在两年前开展的实验,该实验使用 GPT-4o 测试其在处理日益庞大的数字时,“计算总和但以文字形式返回答案”的能力。

以下是他就此结果所展示的图表:

Heatmap chart of accuracy on an addition prompt, colored from dark green (high) through yellow to dark red (low). Title: "What is {a} + {b}? Please write your answer in words. Do not include any other text or information, just the answer in words." Subtitle: 30 randomly selected pairs for each digit combination (n = 30 * 13 * 13 = 5070). X axis: Number of digits in a, 1 to 13. Y axis: Number of digits in b, 1 to 13. Legend: Accuracy, 1.00, 0.75, 0.50, 0.25, 0.00. Values by row, listed for a =...

我确信 GPT-4o 并未作弊使用计算器,尤其是考虑到它在许多计算中都出现了错误;因此,我决定在本地硬件(DGX Spark)上重新开展这一实验,以便在完全受控的环境中进一步探究相关影响。

我将他的图片粘贴到一个 Codex Remote 会话(GPT-6 Astra)中,并让其使用 Qwen3.8-27B-Q4_K_M.gguf 运行相同的实验。以下是禁用推理功能、每种组合运行 30 次的结果:

Heatmap in the same layout as the previous chart, using an orange (low) to white to blue (high) color scale, showing much lower accuracy overall. Title: Addition in words — Qwen3.8 27B Q4_K_M. Subtitle: Reasoning disabled · 30 fixed pairs per ordered digit-length cell (n = 5,070). Overall numeric accuracy: 1,195 / 5,070 (23.57%). X axis: Number of digits in a, 1 to 13. Y axis: Number of digits in b, 1 to 13. Legend: Accuracy, 100%, 75%, 50%, 25%, 0%. Values by row, listed for a = 1 to 13. b =...

随后,我又在启用推理功能的情况下进行了测试。由于每对计算所需时间显著延长,因此我没有再进行每格 30 次的采样,而是仅运行一次——这使得热力图的视觉效果大打折扣,因为每个方格要么为 100%,要么为 0%:

Heatmap in the same layout as the previous charts, almost entirely blue. Title: Addition in words — Qwen3.8 27B — medium reasoning pilot. Subtitle: 1 fixed pair per ordered digit-length cell · easiest first (n = 169). X axis: Number of digits in a, 1 to 13. Y axis: Number of digits in b, 1 to 13. Legend: Accuracy, 1.00, 0.75, 0.50, 0.25, 0.00. Every cell shows 100% except two orange cells showing 0%: a = 2 with b = 8, and a = 12 with b = 9.

在 169 次尝试中,模型正确回答了 167 次;由于这些均为单次生成,我确信再次运行可能会得到不同的结果。

这里提供一个 报告版本 链接,其中包含部分较大规模计算的推理轨迹,其中包括如下文本:

等等,让我再仔细算一遍。

4,299,366,105,622
6,088,794,067,970

我把它们对齐:
4 2 9 9 3 6 6 1 0 5 6 2 2
6 0 8 8 7 9 4 0 6 7 9 7 0

从右至左逐位相加:
第 1 位(个位):2 + 0 = 2
第 2 位(十位):2 + 7 = 9
第 3 位(百位):6 + 9 = 15,写 5,进 1

标签:数学、前往、生成式人工智能、本地大型语言模型、大型语言模型、Qwen、LLM推理、DGX-火花

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论