recurrence

Tag

Cards List
#recurrence

A new transformer passes its hidden state to the next token instead of recomputing it (25 minute read)

TLDR AI ↗ · 21h ago Cached

The paper introduces LIFT (Latent Information Feedback Transformer), which propagates deep-layer hidden states across generation steps instead of relying solely on decoded tokens, using teacher-forced pretraining with states derived from an off-the-shelf LM's next-token distribution. Experiments on 135M–1B models show LIFT outperforms standard Transformers on language modeling, reasoning, and procedural tasks under token-matched budgets, with code and models released.

0 favorites 0 likes
#recurrence

Scaling Laws for Looped Mixture of Experts

Hugging Face Daily Papers ↗ · yesterday Cached

Introduces Loop Scaling Laws, the first scaling law jointly modeling recurrence (looped transformers) and sparsity (MoE) alongside model size and data, showing looped MoEs match models ~2x larger on reasoning benchmarks at matched compute.

0 favorites 0 likes
#recurrence

@gklambauer: The Illustrated Recurrence: From Amari-Hopfield Nets to GPT-6 Astra First blog. Recurrence has been discussed since GPT…

X AI KOLs Timeline ↗ · 2026-09-21 Cached

A blog post by Günter Klambauer providing an illustrated overview of recurrence in neural networks, tracing its history from Amari-Hopfield Nets to mentions of GPT-6 Astra and categorizing various architectures based on recurrence patterns.

0 favorites 0 likes
#recurrence

Looped Transformers under the Jacobian Lens: Does the Global Workspace Survive Recurrence?

arXiv cs.AI ↗ · 2026-09-03 Cached

This paper investigates whether the global workspace functionality in transformers persists under recurrence by applying Jacobian lens analysis to looped and recurrent models, finding that recurrence alters workspace accessibility and content transport.

0 favorites 0 likes
#recurrence

@lateinteraction: Intuition: Compaction is agentic recurrence (RNNs), whereas recursion (RLMs) is agentic attention. Recurrence maintains…

X AI KOLs Timeline ↗ · 2026-08-15 Cached

The article shares an intuition that compaction in RNNs represents agentic recurrence while recursion in RLMs represents agentic attention, comparing different context-handling approaches in AI models.

0 favorites 0 likes
#recurrence

Looped Transformers with Source-Centered State Evolution

arXiv cs.LG ↗ · 2026-07-31 Cached

The paper proposes Source-Centered State Evolution (SCSE), a method for looped Transformers that reconciles input conditioning with reference-preserving shared recurrence, improving recurrent quality frontiers across multiple benchmarks.

0 favorites 0 likes
#recurrence

On Locality and Length Generalization in Visual Reasoning

Hugging Face Daily Papers ↗ · 2026-07-10 Cached

This paper shows that state-of-the-art vision-language models fail at length generalization in visual reasoning due to 'global shortcuts', and demonstrates that combining local foveated perception with recurrence enables robust out-of-distribution generalization.

0 favorites 0 likes
#recurrence

ITNet: A Learnable Integral Transform That Subsumes Convolution, Attention, and Recurrence

arXiv cs.AI ↗ · 2026-06-20 Cached

Introduces ITNet, a neural architecture based on a learnable integral transform that unifies convolution, attention, and recurrence, achieving strong results across multiple modalities.

0 favorites 0 likes
#recurrence

@BlinkDL_AI: Gated DeltaNet-2 is almost exactly RWKV-7's DPLR recurrence, not acknowledging the elephant in the room

X AI KOLs Following ↗ · 2026-05-22 Cached

Ali Hatamizadeh announces Gated DeltaNet-2, a new linear attention model that outperforms KDA and Mamba-3 at 1.3B scale; @BlinkDL_AI notes its recurrence is nearly identical to RWKV-7's DPLR.

0 favorites 0 likes
#recurrence

MMoA: An AI-Agent framework with recurrence for Memoried Mixure-of-Agent

arXiv cs.CL ↗ · 2026-05-20 Cached

Proposes MMoA, a novel AI-agent framework that incorporates recurrence mechanisms for a memoried mixture-of-agent architecture. The paper introduces a method to improve agent collaboration and memory in multi-agent systems.

0 favorites 0 likes
← Back to home

Submit Feedback