Principled Thoughts for Latent Recursive LLM Systems

Hugging Face Daily Papers Papers

Summary

The paper presents REST, a novel training objective for latent recursive LLM systems that enhances accuracy by up to 7.5 percentage points across benchmarks by incorporating properties like causality and minimality into differentiable losses.

Large language models can reason in continuous space instead of decoded text, by recurring on their own hidden states or by passing those states between agents, while training supervises only the Cross-Entropy (CE) of the final decoded answer and does not constrain the thought. Theoretical and empirical analyses establish and confirm four failures of CE-only training that lead to a lower probability of the correct answer such as collapsing thoughts across distinct questions and retaining irrelevant information. We introduce REST (REpresentation-Supervised Thoughts), a training objective that turns four properties of a valid thought representation (causality, minimality, separability, and stability) into differentiable losses added to CE. We instantiate it in latent single-agent and multi-agent systems, without architectural changes or added parameters at inference. Across 7 benchmarks spanning mathematics, science, medicine, and code generation, with the same training data, compute, and latent budget, REST increases accuracy over CE-only training across agent settings and model sizes by up to 7.5 percentage points and convergence on a final answer by 30\%. Furthermore, REST thoughts encode more of what is required to achieve the correct answer, and decoding them better recovers the intended output of the agent, which makes latent communication easier to interpret. Project Website: https://fard-lab.github.io/REST
Original Article
View Cached Full Text

Cached at: 09/30/26, 04:19 AM

Paper page - Principled Thoughts for Latent Recursive LLM Systems

Source: https://huggingface.co/papers/2609.36159

Abstract

Largelanguagemodelscanreasonincontinuousspaceinsteadofdecodedtext,byrecurringontheirownhiddenstatesorbypassingthosestatesbetweenagents,whiletrainingsupervisesonlytheCross-Entropy(CE)ofthefinaldecodedansweranddoesnotconstrainthethought.TheoreticalandempiricalanalysesestablishandconfirmfourfailuresofCE-onlytrainingthatleadtoalowerprobabilityofthecorrectanswersuchascollapsingthoughtsacrossdistinctquestionsandretainingirrelevantinformation.WeintroduceREST(REpresentation-SupervisedThoughts),atrainingobjectivethatturnsfourpropertiesofavalidthoughtrepresentation(causality,minimality,separability,andstability)intodifferentiablelossesaddedtoCE.Weinstantiateitinlatentsingle-agentandmulti-agentsystems,withoutarchitecturalchangesoraddedparametersatinference.Across7benchmarksspanningmathematics,science,medicine,andcodegeneration,withthesametrainingdata,compute,andlatentbudget,RESTincreasesaccuracyoverCE-onlytrainingacrossagentsettingsandmodelsizesbyupto7.5percentagepointsandconvergenceonafinalanswerby30\%.Furthermore,RESTthoughtsencodemoreofwhatisrequiredtoachievethecorrectanswer,anddecodingthembetterrecoverstheintendedoutputoftheagent,whichmakeslatentcommunicationeasiertointerpret.ProjectWebsite:https://fard-lab.github.io/REST

View arXiv pageView PDFProject pageGitHub0Add to collection

Get this paper in your agent:

hf papers read 2609\.36159

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.36159 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.36159 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.36159 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

Learning to Refine Hidden States for Reliable LLM Reasoning

arXiv cs.LG

Proposes ReLAR, a reinforcement-guided latent refinement framework that iteratively updates hidden representations in LLMs before decoding, improving reasoning reliability and efficiency compared to chain-of-thought methods.

Formalizing Latent Thoughts: Four Axioms of Thought Representation in LLMs

Hugging Face Daily Papers

Introduces an axiomatic evaluation framework for latent thought representations in LLMs, revealing that current representations fail to satisfy four fundamental functional axioms (Causality, Minimality, Separability, Stability) across 23 reasoning tasks, indicating a structural gap in representation quality.

When Do LLMs Reason? A Dynamical Systems View via Entropy Phase Transitions

arXiv cs.LG

This paper investigates when chain-of-thought reasoning is beneficial for LLMs, showing that early-stage entropy dynamics reliably indicate reasoning utility, and introduces EDRM, a lightweight, training-free framework that adaptively selects inference strategies to achieve significant token savings while maintaining or improving accuracy.

Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents

Hugging Face Daily Papers

This paper introduces 'progress advantage', an implicit advantage function derived from reinforcement learning post-training that enables effective step-level scoring for LLM agents without requiring dedicated reward model training. It outperforms confidence-based baselines and trained reward models across multiple benchmarks and model families.