What Makes Recurrence Effective in Looped Language Models?
Summary
This paper systematically studies when recurrence helps in looped language models (LoopLMs), finding that extra recurrence can improve reasoning beyond the training horizon but degrade knowledge retention, and proposes channel-wise history-state injection with timestep conditioning as a more robust design for variable inference budgets.
View Cached Full Text
Cached at: 10/01/26, 12:19 AM
Paper page - What Makes Recurrence Effective in Looped Language Models?
Source: https://huggingface.co/papers/2609.36636
Abstract
Loopedlanguagemodels(LoopLMs)increasecomputationaldepththroughparametersharing,offeringapathtoscaleinferencecomputationwithoutaddingparameters.However,itremainsunclearwhenadditionalrecurrenceisbeneficialandhowarchitecturalchoicesaffectitseffectiveness.Throughcontrolledexperiments,wesystematicallyexamine(1)whenrecurrencehelps,(2)whereitshouldbeapplied,and(3)howitsconditioningaffectsperformance.Ourevaluationcoversinferencebudgetsbelow,within,andbeyondthetraininghorizonunderknowledgeandreasoningtasks.(1)Wefindthatrecurrencecanimprovereasoningbeyondthetraininghorizonwhiledegradingknowledgeperformance,butharderreasoninginstancesdonotconsistentlybenefitmore.(2)Performancealsodependsonhowdistinctlayersandrecurrentiterationsareallocated,showingthateffectivedepthaloneisinsufficienttopredictbehavior.Non-recurrentoutputlayersimproverobustnesstounder-unrolling,whilethepreferredplacementofinputandoutputlayersvarieswithinferencebudget.(3)Finally,wefindthatconventionalinitial-stateinjectionofferslimitedrobustnesstovaryingrecurrencedepth.Wethereforeproposehistory-stateinjectionasanalternative,andshowthatchannel-wisehistory-stateinjectioncombinedwithtimestepconditioningoffersalow-costandmoreeffectivedesign,betterpreservingknowledgeunderextendedunrollingwhileimprovingrobustnessacrossinferencebudgets.Overall,ourresultsclarifywhenrecurrentcomputationhelps,whereitfails,andofferpracticalguidelinesfordesigningLoopLMsacrossvariableinferencebudgets.
View arXiv pageView PDFProject pageAdd to collection
Get this paper in your agent:
hf papers read 2609\.36636
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.36636 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.36636 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.36636 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Looped Language Models Improve Compositional Tool Calling
Looped language models enhance compositional tool calling by leveraging recurrent computation, improving accuracy on multi-step tasks while adaptive inference optimizes the balance between performance and compute cost. The study suggests these models are promising for reliable agentic systems.
RecurTrace: Adaptive Latent Reasoning with Loop-Time Memory
RecurTrace introduces loop-time memory and adaptive halting to improve latent reasoning in language models, achieving higher accuracy on MathQA with optimized compute compared to fixed-loop methods.
Allocating Recurrent Compute in Looped Language Models
This paper introduces MixerLoop, a method that allocates recurrent compute by selectively looping the mixer component in language models while applying the feed-forward network once, achieving performance improvements with reduced computational costs.
Recursive Language Models
This paper introduces Recursive Language Models (RLMs), an inference strategy that enables LLMs to process arbitrarily long prompts by treating them as external environments and recursively calling themselves over prompt snippets. RLMs handle inputs two orders of magnitude beyond context windows and outperform base LLMs on long-context tasks with comparable cost.
Looped State-Space Language Models with Adaptive Exit-State Selection
This paper explores looped (recurrent) state-space language models using Mamba and hybrid Mamba-Transformer backbones, showing they outperform non-looped baselines on reasoning tasks and remain competitive under iso-parameter and iso-FLOPs pretraining, with adaptive exit-state selection improving intermediate-depth performance.