Looped transformers have emerged as a parameter-efficient architecture by reusing a shared stack of layers across multiple input positions during a single forward pass. This reuse treats computation as a dynamic process that can be extended per-token, allowing the model to allocate more processing steps to challenging inputs while keeping parameter count low. Prior work has compared looped and non-looped models at matched parameter counts or per-token FLOPs, but whether looping genuinely improves test-time scaling as output length grows has remained an open question.
The core idea behind test-time scaling is straightforward: as models receive more compute during decoding, accuracy should improve. The accuracy-compute slope quantifies this relationship as the accuracy gain per doubling of test-time decoding FLOPs. A steeper slope means each additional round of computation yields greater predictive gains. The paper investigates this slope for looped transformers and finds that existing approaches often produce steeper slopes than their non-looped counterparts when measured per doubling of FLOPs. However, at matched compute budgets, these same looped models underperform the baseline, revealing a trade-off between slope steepness and absolute accuracy.
To understand why this trade-off exists, the authors analyzed where extra loop iterations actually help. Their analysis showed that many tokens achieve near-optimal predictions after just one or two iterations, while a minority of tokens benefit from deeper looping. Fixed-depth looping, which applies the same number of iterations to every token regardless of difficulty, wastes compute on easy tokens and may deny hard tokens the time they need. This uneven benefit profile suggests that an adaptive approach—one that decides per-token whether additional iteration is worthwhile—could capture the best of both worlds: a steep slope from selective looping and high accuracy from giving hard tokens sufficient compute.
Adaptive Iteration Allocation with TaH2
The paper's central contribution, TaH2, introduces an iteration decider module that works alongside the backbone transformer to determine, for each token, how many loop iterations are needed. This decider is not hand-designed; it is jointly post-trained with the backbone through a supervision signal called lookahead depth supervision. The key insight is that during training, the model can observe online labels indicating whether an additional iteration would improve the current prediction or not. These labels act as a binary signal: if the prediction stabilizes or improves with one more iteration, the label says "continue"; otherwise, it says "stop."
By backpropagating through this decision process, the backbone learns to produce representations that are more informative after each iteration, while the iteration decider learns to recognize tokens that still stand to gain from further computation. The result is a system that automatically allocates extra loop steps only where they matter, rather than applying a fixed depth uniformly. This adaptive allocation improves both the efficiency of test-time scaling—the slope—and the accuracy that can be reached at any given compute budget.
The lookahead depth supervision signal is computed by comparing the model's prediction after k iterations to its prediction after k+1 iterations. If the change falls below a threshold or the accuracy metric improves, the label for that token at that step is set to "sufficient." Otherwise, the label encourages another iteration. This comparison happens online during the forward pass, meaning the model is effectively self-supervising its own compute allocation. Over many such supervised steps, both the backbone and the decider adapt to each other: the backbone produces more discriminative representations, and the decider becomes better at predicting when enough computation has been done.
An important detail is that the supervision does not require ground-truth labels for every intermediate step. Instead, it uses the final ground-truth label from the training dataset to compute a reference accuracy, and then evaluates whether intermediate predictions move toward or away from that reference. This makes the signal tractable even though the looped forward pass does not produce a full set of supervised targets at every step.
AIME Benchmark Results and Slope Analysis
On the AIME benchmark, which consists of challenging mathematics problems, TaH2 delivers striking improvements. The accuracy-compute slope increases from 1.79 for the non-looped baseline to 2.74 for TaH2, a 53% relative gain. This means that at each doubling of test-time decoding FLOPs, TaH2 gains more accuracy than the baseline, and the gap widens as more compute is available.
At matched test-time compute, TaH2 exceeds the non-looped baseline's peak accuracy by about 3.4 points. This is a substantial improvement given the difficulty of the AIME problems and the already strong performance of the base model. The result demonstrates that TaH2 not only scales more efficiently but also achieves higher accuracy when operating within the same compute budget.
The paper also examines how performance changes as the maximum allowed iteration depth increases. Existing looped transformers with fixed depth largely plateau as the depth grows deeper—additional iterations yield diminishing returns across the board. In contrast, TaH2's gain over the non-looped baseline continues to grow: from +2.8 points at a maximum depth of 2, to +3.4 points at depth 4, to +3.9 points at depth 8. This trend suggests that as more iteration capacity becomes available, adaptive methods can extract progressively more value, while fixed-depth methods hit a ceiling.
Practical Implications
For a working developer, TaH2's approach offers a pattern that can be adapted beyond the specific AIME setting. The idea of an iteration decider trained with lookahead depth supervision can be applied to any looped transformer where test-time compute is a resource constraint. The key implementation steps are: add a lightweight decision head on top of the backbone's per-token representations, define a supervision signal that compares consecutive iteration outputs, and jointly fine-tune both components. The result is a model that respects compute budgets while maximizing predictive power where it matters most.
Developers working with long-generation tasks—such as theorem proving, code completion, or extended reasoning benchmarks—may find that adaptive looping reduces the need for manual hyperparameter tuning of iteration counts. Instead of guessing a fixed depth that works "okay" for most inputs, the model learns to allocate its FLOPs dynamically. This can lead to more predictable scaling behavior and better accuracy at lower average compute cost.
The paper's code, released alongside the study, provides a reference implementation that researchers can examine and extend. The authors note that the lookahead depth supervision framework is agnostic to the specific backbone architecture, meaning it can be combined with other post-training enhancements or integrated into existing looped transformer pipelines with modest modifications.
Overall, the work addresses a gap in understanding how looping affects test-time scaling trends and provides a concrete method that improves both the efficiency and the attainable accuracy of scaling. The results on AIME demonstrate that adaptive iteration allocation can outperform both pure non-looped baselines and fixed-depth looped models across a range of compute budgets, and the underlying principle—training a decision mechanism with online labels—may prove useful in other settings where compute must be allocated dynamically per input.