TL;DR:在被要求直接回答(不使用思维链)时,大语言模型仅能跟随如 K = apple; B = K; D = B; print(D) 这样的引用链短短几行。在早期层中引入一个微小的 rank-8 LoRA,可使冻结的中间层将链条延续得更远。计算能力确实存在;只是它过早停止了。
主要发现
13个基础模型(0.6B–32B)仅能可靠地跟随1.4–3.6行。将深度加倍(OLMo-3 7B → 32B)后,链条长度仍保持在约2.6行。
Qwen3-8B + 在第14层训练的 rank-8 LoRA(65,537个参数,所有基础权重冻结):在24行链条上的精确准确率从15.5%提升至99%。经过更长训练的版本可在单次前向传播中处理长达50行的链条。
机制:LoRA 对每个 token 独立作用,且不传递任何跨 token 的信息。它启动了一个接力过程:程序行通过冻结的第16–22层将链条身份传递下去。在第14–22层切断每行对其父节点的注意力会导致准确率降至随机水平;而在第23–29层进行相同切断则影响甚微。
放置悬崖:使用相同的配方,LoRA 位于第20层时可达20.5行;位于第21层时仅达5.2行。对冻结模型的测量在4个预留模型中的3个中,将该界限定位在预注册容差范围内。
循环模型:在 Ouro-1.4B 中,每轮循环应用 LoRA,经过4次循环后可达60行,经过8次循环后≥160行(双链选择准确率)。
多跳问答:在 MuSiQue(黄金段落)上,分别训练的早期层 LoRA 在三个标准模型上使 EM(精确匹配)分数提升9.4–17.9。
🎬 网站:https://lunamos.github.io/stop-thinking-too-early/
💻 代码:https://github.com/Lunamos/stop-thinking-too-early
欢迎提问!
TL;DR: Asked to answer directly (no CoT), LLMs follow reference chains like K = apple; B = K; D = B; print(D) for only a few lines. A tiny rank-8 LoRA at one early layer lets the frozen middle layers carry the chain much further. The computation is there; it just stops too early.
Key findings
13 base models (0.6B–32B) reliably follow only 1.4–3.6 lines. Doubling depth (OLMo-3 7B → 32B) leaves reach at ~2.6 lines.
Qwen3-8B + a task-trained rank-8 LoRA at layer 14 (65,537 params, all base weights frozen): exact accuracy on 24-line chains goes from 15.5% → 99%. A longer-trained version reaches 50 lines in a single forward pass.
Mechanism: the LoRA acts on each token independently and moves no information between tokens. It starts a relay : program lines pass their chain identity along through frozen layers 16–22. Cutting each line's attention to its parent in layers 14–22 drops accuracy to chance; the same cut in layers 23–29 barely matters.
Placement cliff: same recipe, LoRA at layer 20 → 20.5 lines; at layer 21 → 5.2 lines. A frozen-model measurement located this limit within a preregistered tolerance in 3 of 4 held-out models.
Looped models: in Ouro-1.4B, a LoRA applied every loop reaches 60 lines after 4 loops and ≥160 after 8 (two-chain choice accuracy).
Multi-hop QA: on MuSiQue (gold paragraphs), separately trained early-layer LoRAs add 9.4–17.9 EM across three standard models.
🎬 Website : https://lunamos.github.io/stop-thinking-too-early/
💻 Code: https://github.com/Lunamos/stop-thinking-too-early
Happy to answer questions!
首次收录 · 2026-10-04 · 10.22 分