Tag
This paper systematically studies when recurrence helps in looped language models (LoopLMs), finding that extra recurrence can improve reasoning beyond the training horizon but degrade knowledge retention, and proposes channel-wise history-state injection with timestep conditioning as a more robust design for variable inference budgets.
WaveFront Decoding introduces a training-free self-speculative decoding framework for looped language models that reduces latency by concurrently batching drafting and verification, achieving up to 4.81x speedup on Huginn-3.5B.
This paper introduces MixerLoop, a method that allocates recurrent compute by selectively looping the mixer component in language models while applying the feed-forward network once, achieving performance improvements with reduced computational costs.
The paper explores how looped language models, which use iterative latent computation, improve compositional tool calling in agentic systems, showing benefits for multi-step API interactions.
Looped language models enhance compositional tool calling by leveraging recurrent computation, improving accuracy on multi-step tasks while adaptive inference optimizes the balance between performance and compute cost. The study suggests these models are promising for reliable agentic systems.
This paper investigates whether a frozen looped transformer can read its own computation quality (pre-answer prediction reaching AUROC 0.797) and whether external interventions can improve outcomes, finding that no tested frozen intervention produces a validated capability gain, a property termed operational proto-introspection.