Stop Thinking Too Early
Preprint2026Mechanistic interpretability · Looped transformers

Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It

Give a pretrained transformer a chain of references in its prompt, K = apple, B = K, D = B, …, print(D), and ask for the answer directly. Every model we tried follows only a few lines and then guesses, whatever its size. The weights can do far more: one rank-8 LoRA at one early layer, under 0.01% of the parameters, lets Qwen3-8B follow 50-line chains in a single pass and lets a looped model follow at least 160.

Watch the relay

Qwen3-8B reads a program with two chains of sixteen lines. A line lights up once its hidden state tells which chain it belongs to. Drag through the 36 layers, frozen and with the LoRA.

after layer 22
Frozen Qwen3-8Bline 4 / 16
print(E)Output: ?
Exact accuracy, 16-line chains15.5%
+ rank-8 LoRA at layer 14line 16 / 16
print(E)Output: apple
Exact accuracy, 16-line chains98.5%
frozenwith the LoRAfurthest readable line after each layer

Measured, not simulated: a linear read-out at each line's pointer, fitted on 450 programs and scored on 150, counts a line as readable at 75% accuracy. Frozen, the relay reaches line 4 at layer 17 and stops. With the LoRA it runs through layers 16 to 22 to the last line, and the question can read it from layer 22.

15.5%→99%exact accuracy on 24-line chains, Qwen3-8B, one pass
50 linesfollowed in one pass by a LoRA trained on longer programs
≥160 linesfollowed by Ouro-1.4B after eight loops
65,537trained parameters, under 0.01% of Qwen3-8B
+9.4 to +17.9MuSiQue exact-match points from an early-layer LoRA
1 · The failure

Every model stops after a few lines

Thirteen base models from 0.6B to 32B parameters, with 16 to 64 layers, follow only 1.4 to 3.6 lines reliably. Depth does not help: OLMo-3-32B has twice the layers of OLMo-3-7B and both reach about 2.6 lines. DeepSeek-V4-Flash (292B MoE) reaches 4.0 lines and is near chance by six.

Looped models, which run their layers again, do not escape it. Ouro-1.4B gains one or two lines from its second loop and almost nothing after that: its reach is 0, 1.6, 2.3, 2.2 and 2.2 lines after one to five loops.

Accuracy against chain length on programs with three chains, as the model's choice among the three root values (chance 1/3). Hover or tap a grey line to see which model it is.
2 · The fix

One rank-8 LoRA at one layer

We add a rank-8 LoRA to the residual stream at the input of one early layer, h ← s·h + BAh, and train only A, B and s. A penalty keeps the model's predictions on ordinary text: WikiText perplexity goes from 10.14 to 10.148. Every weight of the model stays frozen.

Qwen3-8B then answers 24-line chains 99% of the time instead of 15.5%. A LoRA trained on longer programs answers 98% of 40-line and 88% of 48-line chains. In looped models every loop now counts: Ouro-1.4B follows 60 lines after four loops and at least 160 after eight, and re-running Qwen3-8B's layers 14 to 22 with the LoRA lifts 64-line chains from 34% to 92%.

Reach, frozen and with the LoRA
Longest chain followed with at least 80% accuracy, two chains; log scale
Ouro-1.4B: loops pay once the relay is on
Reach against the number of loops at inference; open circles are lower bounds
A tiny change, large gains. Qwen3-8B uses exact accuracy for both bars; Ouro and Huginn use the choice between the two roots. The 24-line evaluations mark every tested length as answered; the LoRA trained on up to 40 lines is evaluated through 160 lines.
3 · Why it works

A relay in the middle layers

By default each line passes on which chain it belongs to for only two or three lines, the question resolves one or two more pointers itself, and a late layer copies the answer. With the LoRA, every line passes its chain on. This relay runs in one short range of middle layers: layers 16 to 22 of Qwen3-8B, and layers 7 to 15 of every loop of Ouro.

  • In each of these layers, attention from a line reaches a little further up its own chain. Frozen heads that read only the parent line take over the longer reads.
  • Cutting each line's attention to its parent line in layers 14 to 22 returns 6-, 8- and 12-line chains to chance (53, 48 and 55%). The same cut after the relay leaves 100, 100 and 98%.
  • Removing the ten heads that read the parent line drops 16-line accuracy from 89% to 53%; ten random heads from the same layers leave 83%.
  • Small looped transformers trained from scratch on these programs learn the same relay.
Schematic of the default and LoRA computation, and the relay front per layer and per loop
The LoRA extends a relay that runs in the same middle layers. (a, b) Qwen3-8B's default and LoRA computation: program lines carry chain identity; by default the question follows the remaining pointers itself, with the LoRA it reads the relay; late layers copy the value. (c, d) The longest prefix of lines whose chain is decoded at 75% accuracy, after each layer of Qwen3-8B and Llama-3.1-8B, and across four loops of Ouro-1.4B.
4 · Where to put it

Early enough to start the relay

The LoRA works only while it can still start the relay. In Qwen3-8B, moving it from layer 20 to layer 21 drops reach from 20.5 to 5.2 lines. A measurement on the frozen model, the cutoff layer, located this limit in advance for three of four held-out models, within a preregistered tolerance.

Place matters far more than form. At layer 14, four different small changes all extend the chain; at layer 26, after the relay's layers, all four stay at chance.

Reach against the layer of the LoRA in nine models, and the frozen cutoff layer against the measured limit
A model-specific layer separates working and failing placements. (a) Reach of independently trained rank-8 LoRAs against their layer relative to depth. (b) The frozen-model cutoff layer against the measured limit in nine models; filled markers are held-out models.
Change to Qwen3-8BParametersLayer 14Layer 26
LoRA on the residual stream65,53796.554.5
Projection LoRA (all seven projections)606,20890.052.5
FLAS-style low-rank flow, 3 steps65,87294.054.0
FLAS flow block, 1 step168M95.556.0
Exact accuracy (%) on 24-line, two-chain programs, 200 programs per cell, every change trained the same way. Answering one of the two root values at random gives 50%.
5 · Beyond programs

Multi-hop questions

On chains of fictional facts, a Qwen3-8B LoRA raises exact match from 41.2% to 97.4%, and six-hop questions, longer than any in training, from 23% to 88%. On MuSiQue, a LoRA trained on the benchmark at the earliest layer tested adds 11.4, 9.4 and 17.9 exact-match points to Qwen3-8B, OLMo-3-7B and Llama-3.1-8B, and 11.7 to Ouro-1.4B. At the latest layer tested it adds little in the standard models.

LoRA at the earliest layer testedLoRA at the latest layer tested
MuSiQue exact-match gain over the frozen model, 900 development questions with their paragraphs (Ouro: 300, four loops), with paired 95% bootstrap intervals. Layers 6 and 30 of Qwen3-8B, 4 and 28 of OLMo-3-7B and Llama-3.1-8B, and 6 and 20 of Ouro's loop body applied in every loop, where a late LoRA still comes before the next loop's middle layers.
Citation

BibTeX

@article{jin2026thinking,
  title   = {Transformers Stop Thinking Too Early, and a Tiny {LoRA} Fixes It},
  author  = {Jin, Zehao and Deng, Ruixuan and Wang, Junran},
  journal = {arXiv preprint},
  year    = {2026}
}