Tag
The paper introduces LIFT (Latent Information Feedback Transformer), which propagates deep-layer hidden states across generation steps instead of relying solely on decoded tokens, using teacher-forced pretraining with states derived from an off-the-shelf LM's next-token distribution. Experiments on 135M–1B models show LIFT outperforms standard Transformers on language modeling, reasoning, and procedural tasks under token-matched budgets, with code and models released.
Introduces Loop Scaling Laws, the first scaling law jointly modeling recurrence (looped transformers) and sparsity (MoE) alongside model size and data, showing looped MoEs match models ~2x larger on reasoning benchmarks at matched compute.
A blog post by Günter Klambauer providing an illustrated overview of recurrence in neural networks, tracing its history from Amari-Hopfield Nets to mentions of GPT-6 Astra and categorizing various architectures based on recurrence patterns.
This paper investigates whether the global workspace functionality in transformers persists under recurrence by applying Jacobian lens analysis to looped and recurrent models, finding that recurrence alters workspace accessibility and content transport.
The article shares an intuition that compaction in RNNs represents agentic recurrence while recursion in RLMs represents agentic attention, comparing different context-handling approaches in AI models.
The paper proposes Source-Centered State Evolution (SCSE), a method for looped Transformers that reconciles input conditioning with reference-preserving shared recurrence, improving recurrent quality frontiers across multiple benchmarks.
This paper shows that state-of-the-art vision-language models fail at length generalization in visual reasoning due to 'global shortcuts', and demonstrates that combining local foveated perception with recurrence enables robust out-of-distribution generalization.
Introduces ITNet, a neural architecture based on a learnable integral transform that unifies convolution, attention, and recurrence, achieving strong results across multiple modalities.
Ali Hatamizadeh announces Gated DeltaNet-2, a new linear attention model that outperforms KDA and Mamba-3 at 1.3B scale; @BlinkDL_AI notes its recurrence is nearly identical to RWKV-7's DPLR.
Proposes MMoA, a novel AI-agent framework that incorporates recurrence mechanisms for a memoried mixture-of-agent architecture. The paper introduces a method to improve agent collaboration and memory in multi-agent systems.