Tag
A debate resurfaces between AI pioneers Geoffrey Hinton and Yann LeCun regarding the efficacy of autoregressive LLMs, with recent advances in reasoning models reigniting the discussion on whether transformers alone suffice for human-like reasoning.
This study finds that activation steering in latent chain-of-thought reasoning is less effective than in explicit CoT, highlighting a transition gap where interventions in latent space fail to transfer to language generation.
The paper proposes NarraLite, an efficient multimodal generative recommendation framework that uses latent narrative reasoning to improve episodic content prediction with better accuracy and efficiency.
Recurrent Looped Transformer (RLT) is a novel architecture combining a causal encoder with a recurrent decoder to achieve latent reasoning with unbounded temporal depth, model-hardware co-design, and model-RL algorithm co-design.
The article argues that relying on chain-of-thought traces for AI safety is ineffective, as they can be manipulated and do not faithfully represent model behavior, instead emphasizing the need to focus on harness control mechanisms.
This paper proposes Prototype-Mediated Process Supervision (PMPS) for latent chain-of-thought reasoning, achieving token compression and accuracy improvements over explicit CoT methods.
A*-Thought-V2 models chain-of-thought reasoning as hidden-state trajectories, using geometric dynamics to compress non-essential steps into latent tokens, improving accuracy and efficiency in LLM reasoning.
RecurTrace introduces loop-time memory and adaptive halting to improve latent reasoning in language models, achieving higher accuracy on MathQA with optimized compute compared to fixed-loop methods.
The article explores latent reasoning as an alternative to chain-of-thought in AI, categorizing five families of approaches and discussing implications for AGI and interpretability.
Chart Pathway's BDH-CQ, a 150M parameter reasoning model, achieves 29.5% on ARC-AGI-1 at a much lower cost per task compared to larger models like GPT-5.6 Luna, showcasing improved cost-accuracy trade-offs.
Proposes Retrieval Grounding Latent Reasoning (RGLR), a latent reasoning framework for dense retrieval that explicitly connects intermediate latent transitions with retrieval improvements, outperforming baselines on reasoning-intensive tasks.
This paper introduces SELR, a unified framework for self-explainable latent reasoning that trains a single model to perform efficient reasoning while generating human-readable explanations, eliminating the need for external decoders.
The paper investigates the interpretability of latent reasoning models, finding that reasoning tokens are often unnecessary but can be decoded to reveal interpretable traces when needed, suggesting these models implement expected solutions.
The article introduces BDH-CQ, a 150M parameter recurrent model that combines in-context learning with latent reasoning, achieving 29.5% on ARC-AGI-1 at a cost of $0.0007 per task, setting a new standard for cost efficiency.
This paper explores surgically retrofitting a pretrained language model (e.g., Qwen2.5-0.5B) with recurrent depth, demonstrating that the resulting model can perform deeper latent reasoning, extrapolate past supervised depth, and outperform dense models fine-tuned to reason in tokens, while also revealing catastrophic interference limits.
Pathway's 150M-parameter BDH-CQ model achieves 29.5% on ARC-AGI-1 at a record-low cost of $0.0007 per task, using recurrent memory and latent reasoning instead of long token chains. The architecture may be the breakthrough Andrew Curran teased, with OpenAI researcher Lukasz Kaiser as an investor and adviser.
ENTLORE is a graph-grounded benchmark framework for enterprise question answering that evaluates latent organizational reasoning, showing that even with gold document sources, many implicit relation questions remain unanswered.
Proposes LatentRM, a reward modeling framework that learns intermediate reasoning traces as discrete latent variables to explicitly maximize downstream scalar reward likelihood, improving preference modeling and policy alignment across in-distribution and OOD tasks.
This paper introduces GradCuit, a method for test-time latent reasoning that inserts optimizable latent states at a selected Transformer layer. It achieves 64.5% average accuracy across five backbones and three reasoning benchmarks, outperforming chain-of-thought prompting and showing improved robustness and interpretability.
This paper introduces J-CoT, a recurrent reasoning framework that uses vocabulary-indexed coefficients (J-thoughts) as intermediate interfaces, enabling improved reasoning performance on mathematical, scientific, coding, and path-reasoning tasks without requiring full verbalization or dense hidden state recurrence.