What Makes Recurrence Effective in Looped Language Models?

Hugging Face Daily Papers Papers

Summary

This paper systematically studies when recurrence helps in looped language models (LoopLMs), finding that extra recurrence can improve reasoning beyond the training horizon but degrade knowledge retention, and proposes channel-wise history-state injection with timestep conditioning as a more robust design for variable inference budgets.

Looped language models (LoopLMs) increase computational depth through parameter sharing, offering a path to scale inference computation without adding parameters. However, it remains unclear when additional recurrence is beneficial and how architectural choices affect its effectiveness. Through controlled experiments, we systematically examine (1) when recurrence helps, (2) where it should be applied, and (3) how its conditioning affects performance. Our evaluation covers inference budgets below, within, and beyond the training horizon under knowledge and reasoning tasks. (1) We find that recurrence can improve reasoning beyond the training horizon while degrading knowledge performance, but harder reasoning instances do not consistently benefit more. (2) Performance also depends on how distinct layers and recurrent iterations are allocated, showing that effective depth alone is insufficient to predict behavior. Non-recurrent output layers improve robustness to under-unrolling, while the preferred placement of input and output layers varies with inference budget. (3) Finally, we find that conventional initial-state injection offers limited robustness to varying recurrence depth. We therefore propose history-state injection as an alternative, and show that channel-wise history-state injection combined with timestep conditioning offers a low-cost and more effective design, better preserving knowledge under extended unrolling while improving robustness across inference budgets. Overall, our results clarify when recurrent computation helps, where it fails, and offer practical guidelines for designing LoopLMs across variable inference budgets.
Original Article
View Cached Full Text

Cached at: 10/01/26, 12:19 AM

Paper page - What Makes Recurrence Effective in Looped Language Models?

Source: https://huggingface.co/papers/2609.36636

Abstract

Loopedlanguagemodels(LoopLMs)increasecomputationaldepththroughparametersharing,offeringapathtoscaleinferencecomputationwithoutaddingparameters.However,itremainsunclearwhenadditionalrecurrenceisbeneficialandhowarchitecturalchoicesaffectitseffectiveness.Throughcontrolledexperiments,wesystematicallyexamine(1)whenrecurrencehelps,(2)whereitshouldbeapplied,and(3)howitsconditioningaffectsperformance.Ourevaluationcoversinferencebudgetsbelow,within,andbeyondthetraininghorizonunderknowledgeandreasoningtasks.(1)Wefindthatrecurrencecanimprovereasoningbeyondthetraininghorizonwhiledegradingknowledgeperformance,butharderreasoninginstancesdonotconsistentlybenefitmore.(2)Performancealsodependsonhowdistinctlayersandrecurrentiterationsareallocated,showingthateffectivedepthaloneisinsufficienttopredictbehavior.Non-recurrentoutputlayersimproverobustnesstounder-unrolling,whilethepreferredplacementofinputandoutputlayersvarieswithinferencebudget.(3)Finally,wefindthatconventionalinitial-stateinjectionofferslimitedrobustnesstovaryingrecurrencedepth.Wethereforeproposehistory-stateinjectionasanalternative,andshowthatchannel-wisehistory-stateinjectioncombinedwithtimestepconditioningoffersalow-costandmoreeffectivedesign,betterpreservingknowledgeunderextendedunrollingwhileimprovingrobustnessacrossinferencebudgets.Overall,ourresultsclarifywhenrecurrentcomputationhelps,whereitfails,andofferpracticalguidelinesfordesigningLoopLMsacrossvariableinferencebudgets.

View arXiv pageView PDFProject pageAdd to collection

Get this paper in your agent:

hf papers read 2609\.36636

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.36636 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.36636 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.36636 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

Looped Language Models Improve Compositional Tool Calling

Hugging Face Daily Papers

Looped language models enhance compositional tool calling by leveraging recurrent computation, improving accuracy on multi-step tasks while adaptive inference optimizes the balance between performance and compute cost. The study suggests these models are promising for reliable agentic systems.

Allocating Recurrent Compute in Looped Language Models

arXiv cs.LG

This paper introduces MixerLoop, a method that allocates recurrent compute by selectively looping the mixer component in language models while applying the feed-forward network once, achieving performance improvements with reduced computational costs.

Recursive Language Models

Papers with Code Trending

This paper introduces Recursive Language Models (RLMs), an inference strategy that enables LLMs to process arbitrarily long prompts by treating them as external environments and recursively calling themselves over prompt snippets. RLMs handle inputs two orders of magnitude beyond context windows and outperform base LLMs on long-context tasks with comparable cost.

Looped State-Space Language Models with Adaptive Exit-State Selection

arXiv cs.AI

This paper explores looped (recurrent) state-space language models using Mamba and hybrid Mamba-Transformer backbones, showing they outperform non-looped baselines on reasoning tasks and remain competitive under iso-parameter and iso-FLOPs pretraining, with adaptive exit-state selection improving intermediate-depth performance.