Principled Thoughts for Latent Recursive LLM Systems
Summary
The paper presents REST, a novel training objective for latent recursive LLM systems that enhances accuracy by up to 7.5 percentage points across benchmarks by incorporating properties like causality and minimality into differentiable losses.
View Cached Full Text
Cached at: 09/30/26, 04:19 AM
Paper page - Principled Thoughts for Latent Recursive LLM Systems
Source: https://huggingface.co/papers/2609.36159
Abstract
Largelanguagemodelscanreasonincontinuousspaceinsteadofdecodedtext,byrecurringontheirownhiddenstatesorbypassingthosestatesbetweenagents,whiletrainingsupervisesonlytheCross-Entropy(CE)ofthefinaldecodedansweranddoesnotconstrainthethought.TheoreticalandempiricalanalysesestablishandconfirmfourfailuresofCE-onlytrainingthatleadtoalowerprobabilityofthecorrectanswersuchascollapsingthoughtsacrossdistinctquestionsandretainingirrelevantinformation.WeintroduceREST(REpresentation-SupervisedThoughts),atrainingobjectivethatturnsfourpropertiesofavalidthoughtrepresentation(causality,minimality,separability,andstability)intodifferentiablelossesaddedtoCE.Weinstantiateitinlatentsingle-agentandmulti-agentsystems,withoutarchitecturalchangesoraddedparametersatinference.Across7benchmarksspanningmathematics,science,medicine,andcodegeneration,withthesametrainingdata,compute,andlatentbudget,RESTincreasesaccuracyoverCE-onlytrainingacrossagentsettingsandmodelsizesbyupto7.5percentagepointsandconvergenceonafinalanswerby30\%.Furthermore,RESTthoughtsencodemoreofwhatisrequiredtoachievethecorrectanswer,anddecodingthembetterrecoverstheintendedoutputoftheagent,whichmakeslatentcommunicationeasiertointerpret.ProjectWebsite:https://fard-lab.github.io/REST
View arXiv pageView PDFProject pageGitHub0Add to collection
Get this paper in your agent:
hf papers read 2609\.36159
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.36159 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.36159 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.36159 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Learning to Refine Hidden States for Reliable LLM Reasoning
Proposes ReLAR, a reinforcement-guided latent refinement framework that iteratively updates hidden representations in LLMs before decoding, improving reasoning reliability and efficiency compared to chain-of-thought methods.
Formalizing Latent Thoughts: Four Axioms of Thought Representation in LLMs
Introduces an axiomatic evaluation framework for latent thought representations in LLMs, revealing that current representations fail to satisfy four fundamental functional axioms (Causality, Minimality, Separability, Stability) across 23 reasoning tasks, indicating a structural gap in representation quality.
When Do LLMs Reason? A Dynamical Systems View via Entropy Phase Transitions
This paper investigates when chain-of-thought reasoning is beneficial for LLMs, showing that early-stage entropy dynamics reliably indicate reasoning utility, and introduces EDRM, a lightweight, training-free framework that adaptively selects inference strategies to achieve significant token savings while maintaining or improving accuracy.
Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates
This paper proposes computation-efficient strategies for latent adversarial training (LAT) of LLMs, using low-rank representation fine-tuning and circuit-guided surrogate models to reduce per-step FLOPs by 48.1% while requiring only 0.0118% trainable parameters.
Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents
This paper introduces 'progress advantage', an implicit advantage function derived from reinforcement learning post-training that enables effective step-level scoring for LLM agents without requiring dedicated reward model training. It outperforms confidence-based baselines and trained reward models across multiple benchmarks and model families.