Adaptive Latent Agentic Reasoning

arXiv cs.CL Papers

Summary

This paper introduces Adaptive Latent Agentic Reasoning (ALAR), a dual-mode framework for LLM agents that uses compact latent reasoning for routine turns and selectively escalates to explicit chain-of-thought for harder decisions, achieving up to 84.6% token reduction while maintaining task accuracy.

arXiv:2606.02871v1 Announce Type: new Abstract: Large reasoning models improve performance by generating extended chain-of-thought (CoT) reasoning, but this behavior becomes inefficient when applied to LLM agents. Current LLM agents often generate verbose textual reasoning at every decision step and allocate reasoning effort nearly uniformly across turns, leading to substantial inefficiency in multi-turn agentic trajectories. We propose Adaptive Latent Agentic Reasoning (ALAR), a dual-mode framework that uses compact latent reasoning for routine turns and selectively escalates to explicit chain-of-thought when deeper deliberation is needed. ALAR learns latent reasoning by using the agent's actions as supervision anchors and is further optimized to use latent reasoning when it is sufficient for task success and reserve explicit CoT for harder decisions. Experiments on agentic search and tool-use benchmarks show that ALAR maintains comparable or better task accuracy while substantially reducing generated tokens by up to 43.6% in search and 84.6% in tool use. These results demonstrate that ALAR improves the accuracy-efficiency trade-off of LLM agents by reducing unnecessary textual reasoning while preserving explicit deliberation for harder decision steps.
Original Article
View Cached Full Text

Cached at: 06/03/26, 09:35 AM

# Adaptive Latent Agentic Reasoning
Source: [https://arxiv.org/html/2606.02871](https://arxiv.org/html/2606.02871)
Dongwon Jung1Peng Shi2Yi Zhang3Junshan Zhang1Muhao Chen1

1University of California, Davis2University of Waterloo3Greenshoe, Inc\. \{dwojung,jazh,muhchen\}@ucdavis\.edupeng\.shi@uwaterloo\.cayi@greenshoe\.ai

###### Abstract

Large reasoning models improve performance by generating extended chain\-of\-thought \(CoT\) reasoning, but this behavior becomes inefficient when applied to LLM agents\. Current LLM agents often generate verbose textual reasoning at every decision step and allocate reasoning effort nearly uniformly across turns, leading to substantial inefficiency in multi\-turn agentic trajectories\. We proposeAdaptive Latent Agentic Reasoning\(ALAR\), a dual\-mode framework that uses compact latent reasoning for routine turns and selectively escalates to explicit chain\-of\-thought when deeper deliberation is needed\.ALARlearns latent reasoning by using the agent’s actions as supervision anchors, and is further optimized to use latent reasoning when it is sufficient for task success and reserve explicit CoT for harder decisions\. Experiments on agentic search and tool\-use benchmarks show thatALARmaintains comparable or better task accuracy while substantially reducing generated tokens by up to 43\.6% in search and 84\.6% in tool use\. These results demonstrate thatALARimproves the accuracy\-efficiency trade\-off of LLM agents by reducing unnecessary textual reasoning while preserving explicit deliberation for harder decision steps\.

Adaptive Latent Agentic Reasoning

Dongwon Jung1Peng Shi2Yi Zhang3Junshan Zhang1Muhao Chen11University of California, Davis2University of Waterloo3Greenshoe, Inc\.\{dwojung,jazh,muhchen\}@ucdavis\.edupeng\.shi@uwaterloo\.cayi@greenshoe\.ai

## 1Introduction

Recent advances in large reasoning models \(LRMs\) have shown that extended chain\-of\-thought \(CoT\) reasoning improves performance on mathematical, logical, and coding tasksJaechet al\.\([2024](https://arxiv.org/html/2606.02871#bib.bib36)\); Guoet al\.\([2025](https://arxiv.org/html/2606.02871#bib.bib1)\)\. In the standard single\-pass setting, reasoning is primarily answer\-directed, where the model deliberates before producing a final response\. By contrast, LLM agents extend this paradigm to interactive environments, where reasoning is interleaved with actions such as retrieval, tool use, and environment interactionYaoet al\.\([2022](https://arxiv.org/html/2606.02871#bib.bib37)\); Shinnet al\.\([2023](https://arxiv.org/html/2606.02871#bib.bib38)\)\. We refer to this per\-turn computation performed at each decision step as agentic reasoning\(Weiet al\.,[2026](https://arxiv.org/html/2606.02871#bib.bib10)\): reasoning used to choose the next action, incorporate observations, and decide when to terminate\.

However, current LLM agents largely inherit the reasoning behavior of single\-pass LRMs\. As a result, they often generate lengthy CoTChenet al\.\([2024](https://arxiv.org/html/2606.02871#bib.bib30)\)even when the next action mainly depends on external observations, and they allocate nearly even reasoning effort across turns despite substantial variations in reasoning demands\. This inefficiency compounds in multi\-turn trajectories, where reasoning tokens from earlier turns accumulate in the growing context\. We therefore ask how to make LLM agents reason more efficiently while preserving the deliberation needed for challenging decision steps\.

![Refer to caption](https://arxiv.org/html/2606.02871v1/x1.png)Figure 1:Traditional LRMs generate verbose CoT at every decision step, introducing significant inefficiency in multi\-turn agentic trajectories\.ALARuses compact latent reasoning by default and falls back to explicit CoT only for turns that require deeper planning\.A natural approach is to apply recent reasoning token compression methods, which reduce verbose CoT through pruning, length budgets, or rewards for shorter correct solutions\(Luoet al\.,[2025](https://arxiv.org/html/2606.02871#bib.bib11); Houet al\.,[2025](https://arxiv.org/html/2606.02871#bib.bib12); Yiet al\.,[2026](https://arxiv.org/html/2606.02871#bib.bib13)\)\. However, these methods still operate within the explicit CoT interface where every turn must produce a textual reasoning trace, and efficiency is obtained only by shortening that trace\. This is limiting in the agentic setting, where many turns do not require even a shortened textual rationale, but only sufficient internal computation to choose the next environment\-coupled action\. Thus, efficient agentic reasoning requires a more structural change beyond compressing explicit CoT\.

A promising alternative is implicit chain\-of\-thought or latent reasoningHaoet al\.\([2024](https://arxiv.org/html/2606.02871#bib.bib6)\); Shenet al\.\([2025b](https://arxiv.org/html/2606.02871#bib.bib7)\), which replaces textual reasoning tokens with a fixed\-length sequence of continuous thoughts in the model’s hidden\-state space\. By avoiding the generation of explicit reasoning tokens, implicit CoT provides a compact form of internal computation\. However, extending implicit CoT to agentic settings introduces two challenges\. First, training the latent reasoning mode is nontrivial because continuous thoughts live in hidden\-state space, so variable\-length textual CoT cannot serve as a direct supervision target\. Moreover, in agents, per\-turn reasoning should support intermediate action selection rather than only final\-answer generation\. Second, the model should not rely on latent reasoning uniformly\. Instead, it must retain the ability to escalate to explicit CoT on turns that genuinely require deeper reasoning, to achieve the desired level of performance\.

To address these challenges, we introduce*Adaptive Latent Agentic Reasoning*\(ALAR\), a reasoning architecture for LLM agents that uses latent reasoning by default and escalates to explicit CoT only when the current turn requires deeper reasoning\.ALARconsists of two components\. First,*Action\-Anchored Self\-Distillation*\(AASD\) trains the latent reasoning mode without directly supervising latent states\. Instead of aligning latent thoughts with textual CoT, AASD replaces each teacher CoT span with a latent block and trains the student to reproduce the teacher’s subsequent action\. Since actions are the points where the agent interacts with the environment, they provide natural anchors for supervision\. Second,*Adaptive Reasoning GRPO*\(AR\-GRPO\) learns adaptive mode selection by rewarding latent reasoning when it preserves task success, while encouraging explicit CoT on turns that require more detailed reasoning\.

![Refer to caption](https://arxiv.org/html/2606.02871v1/x2.png)Figure 2:ALARachieves a better accuracy\-efficiency trade\-off than reasoning token compression baselines across search and tool\-use benchmarks\.We evaluate ALAR on agentic search and tool\-use benchmarks against recent reasoning token compression baselines\. As shown in[Figure˜2](https://arxiv.org/html/2606.02871#S1.F2), ALAR achieves a better accuracy\-efficiency trade\-off by reducing tokens more aggressively while preserving task accuracy\. Our contributions are summarized as follows:

- •We introduceALAR, a dual\-mode framework that combines latent reasoning with adaptive mode selection, allowing LLM agents to use compact latent reasoning when it suffices and escalate to explicit CoT at turns where additional deliberation is needed for action selection\.
- •We propose*Action\-Anchored Self\-Distillation*\(AASD\), a self\-distillation method which trains latent agentic reasoning without latent\-state supervision by replacing teacher CoT spans with latent blocks and supervising the student to reproduce the teacher’s next environment\-facing action\.
- •We propose*AR\-GRPO*, a reinforcement learning method that optimizes per\-turn reasoning\-mode selection by rewarding latent\-mode use when task success is preserved and discouraging unnecessary explicit CoT

![Refer to caption](https://arxiv.org/html/2606.02871v1/x3.png)Figure 3:Overview ofALAR\. At each turn, LRM adaptively chooses latent mode for routine decisions or explicit mode for harder turns\. Action\-Anchored Self\-Distillation trains the latent mode by using the teacher’s actions as anchors\. AR\-GRPO further learns when to use latent reasoning by rewarding it only when task success is preserved\.
## 2Related Work

### 2\.1Latent Reasoning

Recent work has explored latent reasoning as an efficient alternative to explicit CoT\. Early methods train models to internalize or compress textual CoT into continuous hidden states\(Denget al\.,[2024](https://arxiv.org/html/2606.02871#bib.bib27); Haoet al\.,[2024](https://arxiv.org/html/2606.02871#bib.bib6); Shenet al\.,[2025b](https://arxiv.org/html/2606.02871#bib.bib7); Cheng and Van Durme,[2024](https://arxiv.org/html/2606.02871#bib.bib24)\)\. More recent hybrid approaches combine latent and explicit reasoning through switching, gating, or token\-level mixing\(Shiet al\.,[2025](https://arxiv.org/html/2606.02871#bib.bib25); Xuet al\.,[2026](https://arxiv.org/html/2606.02871#bib.bib29); Yueet al\.,[2026](https://arxiv.org/html/2606.02871#bib.bib28); Suet al\.,[2025](https://arxiv.org/html/2606.02871#bib.bib26)\)\. These methods mainly target single\-pass reasoning, where latent computation is used to produce a final answer\. Our setting differs in that latent reasoning is action\-oriented, environment\-coupled, and repeated across turns, making the central challenge not only how to compress reasoning, but also how to allocate reasoning modes throughout a trajectory\.

### 2\.2Reasoning Token Reduction

To mitigate overthinking in LRMsChenet al\.\([2024](https://arxiv.org/html/2606.02871#bib.bib30)\); Suiet al\.\([2025](https://arxiv.org/html/2606.02871#bib.bib35)\), recent work has sought to reduce reasoning cost by shortening explicit CoT traces\. One group of methods uses reinforcement learning or fine\-tuning rewards to favor concise\-but\-correct reasoning and prune redundant thinking steps\(Arora and Zanette,[2026](https://arxiv.org/html/2606.02871#bib.bib33); Luoet al\.,[2025](https://arxiv.org/html/2606.02871#bib.bib11); Houet al\.,[2025](https://arxiv.org/html/2606.02871#bib.bib12); Chenget al\.,[2025](https://arxiv.org/html/2606.02871#bib.bib32)\)\. Another group introduces length control or difficulty\-adaptive budgets, allowing models to adjust reasoning length according to a user\-specified budget, sampled optimal length, or problem difficulty\(Aggarwal and Welleck,[2025](https://arxiv.org/html/2606.02871#bib.bib31); Yiet al\.,[2026](https://arxiv.org/html/2606.02871#bib.bib13); Shenet al\.,[2025a](https://arxiv.org/html/2606.02871#bib.bib34)\)\. While effective, these methods still optimize efficiency within the textual CoT interface\. Our work instead changes the reasoning substrate itself, using latent reasoning to bypass unnecessary textual CoT and enable more aggressive token reduction across multi\-turn trajectories\.

## 3Adaptive Latent Agentic Reasoning

To this end, we proposeAdaptive Latent Agentic Reasoning\(ALAR\), a dual\-mode reasoning framework for efficient LLM agents\. We first formulate the LRM as a multi\-turn agent policy \([Section˜3\.1](https://arxiv.org/html/2606.02871#S3.SS1)\), then introduce two core design components:*Latent Agentic Reasoning*\([Section˜3\.2](https://arxiv.org/html/2606.02871#S3.SS2)\) and*Adaptive Mode Selection*\([Section˜3\.3](https://arxiv.org/html/2606.02871#S3.SS3)\)\. We then present the two\-stage optimization procedure:*Action\-Anchored Self\-Distillation*\([Section˜3\.4](https://arxiv.org/html/2606.02871#S3.SS4)\) learns the latent agentic reasoning, and AR\-GRPO learns adaptive mode selection \([Section˜3\.5](https://arxiv.org/html/2606.02871#S3.SS5)\)\.

### 3\.1LRMs as LLM Agents

We consider a large reasoning model \(LRM\) parameterized byθ\\thetathat produces an explicit chain\-of\-thought \(CoT\) before each output\(Guoet al\.,[2025](https://arxiv.org/html/2606.02871#bib.bib1); Xianget al\.,[2025](https://arxiv.org/html/2606.02871#bib.bib4)\)\. In an agentic setting, this reason\-before\-output pattern is repeated across multiple environment\-coupled decision steps\. Specifically, we treat the LRM as the policy of an LLM agent that interacts with a tool environment over up toTTturns\. Given a queryxx, at each turnttthe agent generates an explicit CoTct∼πθ\(⋅∣st\)c\_\{t\}\\sim\\pi\_\{\\theta\}\(\\cdot\\mid s\_\{t\}\), conditioned on the statests\_\{t\}\(the current context\), then emits an actionat∼πθ\(⋅∣st,ct\)a\_\{t\}\\sim\\pi\_\{\\theta\}\(\\cdot\\mid s\_\{t\},c\_\{t\}\)that is either a tool call or the final response\. Ifata\_\{t\}is a tool call, the environment returns an observationoto\_\{t\}that is appended to the context; otherwise, the episode terminates\. The resulting trajectory isτ=\(x,c1,a1,o1,…,cT,aT\)\\tau=\(x,c\_\{1\},a\_\{1\},o\_\{1\},\\ldots,c\_\{T\},a\_\{T\}\), whereaTa\_\{T\}is the final response\.

### 3\.2Latent Agentic Reasoning

The formulation exposes the main inefficiency we target: explicit CoT is generated at every turn, even when the next action may require only lightweight internal computation\. Latent reasoning has so far been studied primarily in single\-pass reasoning tasks\(Haoet al\.,[2024](https://arxiv.org/html/2606.02871#bib.bib6); Shenet al\.,[2025b](https://arxiv.org/html/2606.02871#bib.bib7)\), where continuous thoughts replace the CoT before producing a final answer\. We adapt this idea to multi\-turn agentic reasoning, where reasoning serves a different role: at each intermediate turn, the agent reasons to select the next action toward a long\-horizon goal rather than to directly produce the final answer\.

Specifically, at each turntt, instead of generating an explicit CoTctc\_\{t\}, the agent produces a fixed\-length sequence ofKKcontinuous thoughtszt=\(zt1,…,ztK\)z\_\{t\}=\(z\_\{t\}^\{1\},\\ldots,z\_\{t\}^\{K\}\)\. Starting from the hidden stateht0h\_\{t\}^\{0\}corresponding to the current statests\_\{t\}, each latent thought is generated autoregressively in hidden\-state space:

ztk=fϕ​\(htk−1\),k=1,…,K,z\_\{t\}^\{k\}=f\_\{\\phi\}\(h\_\{t\}^\{k\-1\}\),\\quad k=1,\\ldots,K,wherefϕf\_\{\\phi\}is a projection layer and eachztkz\_\{t\}^\{k\}is fed back as the input embedding for the next latent position\. After the latent block is produced, the agent samples the next action asat∼πθ\(⋅∣st,zt\)a\_\{t\}\\sim\\pi\_\{\\theta\}\(\\cdot\\mid s\_\{t\},z\_\{t\}\), whereata\_\{t\}is decoded over the vocabularyVVconditioned on the current state and the latent thoughts\. We refer to this process as latent agentic reasoning: the agent performs implicit per\-turn computation through a latent block rather than a discrete CoT\.

### 3\.3Adaptive Mode Selection

Although the latent agentic reasoning is sufficient for routine turns, some decisions require more substantive reasoning than a fixed\-length latent block can accommodate\. We therefore equip the agent with a per\-turn choice between latent and explicit*mode*, with the mode sampled directly from the policy:

mt∼πθ\(⋅∣st\),rt∼πθ\(⋅∣st,mt\),m\_\{t\}\\sim\\pi\_\{\\theta\}\(\\cdot\\mid s\_\{t\}\),\\qquad r\_\{t\}\\sim\\pi\_\{\\theta\}\(\\cdot\\mid s\_\{t\},m\_\{t\}\),where the modemt∈\{<lat\>,<think\>\}m\_\{t\}\\in\\\{\\textsc\{<lat\>\},\\textsc\{<think\>\}\\\}determines the form of the per\-turn reasoning tracertr\_\{t\}, which is the latent blockztz\_\{t\}in the latent mode and an explicit CoTctc\_\{t\}in the explicit mode\. The action is then sampled fromπθ\(⋅∣st,rt\)\\pi\_\{\\theta\}\(\\cdot\\mid s\_\{t\},r\_\{t\}\)\. Lettingrt∈\{zt,ct\}r\_\{t\}\\in\\\{z\_\{t\},c\_\{t\}\\\}denote the reasoning trace of turnttunder its selected mode, the resulting trajectory isτ=\(x,m1,r1,a1,o1,…,mT,rT,aT\)\\tau=\(x,m\_\{1\},r\_\{1\},a\_\{1\},o\_\{1\},\\ldots,m\_\{T\},r\_\{T\},a\_\{T\}\)\. Becausemtm\_\{t\}is sampled from the same policy that generates the rest of the trajectory, mode selection becomes part of the agent’s decision space rather than a choice imposed by an external orchestrator or router\.

### 3\.4Action\-Anchored Self\-Distillation

Training the latent mode raises a supervision challenge\. The projectorfϕf\_\{\\phi\}that produces the continuous thoughtsztz\_\{t\}is newly initialized and has no targets to learn from\. An obvious candidate is the explicit CoTctc\_\{t\}thatztz\_\{t\}replaces, but the two are structurally mismatched:ctc\_\{t\}is a variable\-length sequence of discrete tokens, whereasztz\_\{t\}is a fixed\-length sequence of continuous vectors\. Matching them position\-wise would tiefϕf\_\{\\phi\}to the token\-level decomposition of the teacher’s reasoning instead of letting it discover its own\.

We address this with Action\-Anchored Self\-Distillation \(AASD\): the same base model acts as a teacher in the explicit mode and a student in the latent mode, with the student anchored to the teacher’s*actions*\. Anchoring on actions sidesteps the alignment problem: actions are the points at which both modes contact the environment and at which correctness is defined, sofϕf\_\{\\phi\}is free to discover whatever trajectory through hidden\-state space best producesata\_\{t\}fromsts\_\{t\}, without being told whatztz\_\{t\}should look like\.

Teacher rollouts\.Letπθ\\pi\_\{\\theta\}denote the base LRM, shared between the two modes\. We roll out the explicit mode on a training set in the agentic environment, and from each resulting trajectory we extract the*action trajectory*τa=\(a1,o1,a2,o2,…,aT−1,oT−1,aT\)\\tau\_\{a\}=\(a\_\{1\},o\_\{1\},a\_\{2\},o\_\{2\},\\ldots,a\_\{T\-1\},o\_\{T\-1\},a\_\{T\}\), which retains the teacher’s actions and the corresponding environment observations while dropping its explicit CoTs\.

Student objective\.The student shares the base parametersθ\\thetawith the teacher and operates in the latent mode, with the projectorfϕf\_\{\\phi\}providing the continuous thoughts\. Given an action trajectoryτa\\tau\_\{a\}, we form a student trajectoryτ~=\(x,m1,z1,a1,o1,…,mT,zT,aT\)\\tilde\{\\tau\}=\(x,m\_\{1\},z\_\{1\},a\_\{1\},o\_\{1\},\\ldots,m\_\{T\},z\_\{T\},a\_\{T\}\)by inserting a latent mode tokenmt=<lat\>m\_\{t\}=\\textsc\{<lat\>\}and a latent blockztz\_\{t\}of lengthKKbefore each anchor actionata\_\{t\}in place of the teacher’s CoT\.

We train\(θ,ϕ\)\(\\theta,\\phi\)by maximizing the log\-likelihood of the teacher’s anchor actions under the student trajectory, conditioned on the statests\_\{t\}and the preceding latent blockztz\_\{t\}, produced by the projector chainztk=fϕ​\(htk−1\)z\_\{t\}^\{k\}=f\_\{\\phi\}\(h\_\{t\}^\{k\-1\}\):

ℒAASD=−𝔼τa∑t=1T\[\\displaystyle\\mathcal\{L\}\_\{\\text\{AASD\}\}=\-\\mathbb\{E\}\_\{\\tau\_\{a\}\}\\sum\_\{t=1\}^\{T\}\\big\[log⁡πθ​\(mt∣st\)\\displaystyle\\log\\pi\_\{\\theta\}\(m\_\{t\}\\mid s\_\{t\}\)\+logπθ\(at∣st,mt,zt\)\]\.\\displaystyle\+\\log\\pi\_\{\\theta\}\(a\_\{t\}\\mid s\_\{t\},m\_\{t\},z\_\{t\}\)\\big\]\.The loss is applied to the action tokens and the mode tokensmtm\_\{t\}, so that the model also learns to emit the mode tokens at the start of each turn\. TheKKlatent positions have no discrete token target to compute cross\-entropy against, sinceztz\_\{t\}lives in continuous space rather than over the vocabularyVV, and the environment observationsoto\_\{t\}are masked out to stabilize training\(Jinet al\.,[2025](https://arxiv.org/html/2606.02871#bib.bib5)\)\. The latent blockztz\_\{t\}that the student inserts betweensts\_\{t\}andata\_\{t\}is learned end\-to\-end: the cross\-entropy at each anchor actionata\_\{t\}back\-propagates through the transformer to theKKlatent input positions and from there through the iterative projector chainztk=fϕ​\(htk−1\)z\_\{t\}^\{k\}=f\_\{\\phi\}\(h\_\{t\}^\{k\-1\}\), accumulating gradient contributions across allKKprojector steps\.

### 3\.5AR\-GRPO

After learning the latent mode with AASD, we train adaptive mode selection by first initializing the mode distribution with a brief mode\-warmup SFT and then optimizing the policy with AR\-GRPO\. The goal is to encourage latent reasoning whenever it improves efficiency without sacrificing task success, while preserving the ability to escalate to explicit CoT when needed\.

Mode warmup\.AASD trains every turn with<LAT\>, so the resulting policy has little probability mass on<THINK\>and provides weak exploration for adaptive mode selection\. We therefore begin with a brief mode\-warmup SFT: starting from the AASD checkpoint, we assign each turn in a small subset of teacher trajectories to either<LAT\>or<THINK\>\.<LAT\>turns are trained with the AASD objective, while<THINK\>turns are trained with standard cross\-entropy on the teacher’s CoT\.

Trajectory reward\.After warmup, we optimize the agent over complete trajectories\. For each query, we sample a group ofGGrollouts\{τ\(i\)\}i=1G\\\{\\tau^\{\(i\)\}\\\}\_\{i=1\}^\{G\}fromπθ\\pi\_\{\\theta\}\. LetnLAT​\(τ\)n\_\{\\mathrm\{LAT\}\}\(\\tau\)denote the number of latent reasoning turns andnturn​\(τ\)n\_\{\\mathrm\{turn\}\}\(\\tau\)denote the total number of reasoning turns\. We define the*latent fraction*of a trajectory as

f​\(τ\)=nLAT​\(τ\)nturn​\(τ\)∈\[0,1\],f\(\\tau\)=\\frac\{n\_\{\\mathrm\{LAT\}\}\(\\tau\)\}\{n\_\{\\mathrm\{turn\}\}\(\\tau\)\}\\in\[0,1\],withf​\(τ\)=0f\(\\tau\)=0when no reasoning turn is taken\.

Based on this latent fraction, we define an asymmetric format reward that encourages latent reasoning only when it preserves task success:

rfmt​\(τ\)=\{1\+α​f​\(τ\),EM​\(τ\)=1,−α​f​\(τ\),otherwise,r\_\{\\mathrm\{fmt\}\}\(\\tau\)=\\begin\{cases\}1\+\\alpha f\(\\tau\),&\\mathrm\{EM\}\(\\tau\)=1,\\\\ \-\\alpha f\(\\tau\),&\\mathrm\{otherwise\},\\end\{cases\}whereα\>0\\alpha\>0controls the latent mode bonus\. Intuitively, correct trajectories are rewarded more when they rely more on latent reasoning, while incorrect trajectories are penalized for overusing latent reasoning\.

To avoid early collapse to a single mode mixture, we add a decayed diversity bonus,

rdiv​\(τ\(i\)\)=ds​\|f​\(τ\(i\)\)−f¯G\|,r\_\{\\mathrm\{div\}\}\(\\tau^\{\(i\)\}\)=d\_\{s\}\\left\|f\(\\tau^\{\(i\)\}\)\-\\bar\{f\}\_\{G\}\\right\|,wheref¯G\\bar\{f\}\_\{G\}denote the group mean latent fraction anddsd\_\{s\}cosine\-decays from11to0during training\. This term encourages early exploration of different latent\-explicit mixtures, then fades so that the success\-conditioned format reward dominates\.

Finally, we apply a length\-scaling factorsL​\(τ\)s\_\{L\}\(\\tau\)that remains11within the tolerance lengthLLand down\-weights trajectories with overlong explicit<THINK\>segments\. The final trajectory reward combines the format and diversity terms under this length scaling:

R\(i\)=sL​\(τ\(i\)\)​\(rfmt​\(τ\(i\)\)\+rdiv​\(τ\(i\)\)\),R^\{\(i\)\}=s\_\{L\}\(\\tau^\{\(i\)\}\)\\left\(r\_\{\\mathrm\{fmt\}\}\(\\tau^\{\(i\)\}\)\+r\_\{\\mathrm\{div\}\}\(\\tau^\{\(i\)\}\)\\right\),withR\(i\)=−1R^\{\(i\)\}=\-1for invalid output formats\.

GRPO optimization\.We normalize the trajectory rewards within each rollout group to obtain the advantageA^\(i\)=\(R\(i\)−μ\)/σ\\hat\{A\}^\{\(i\)\}=\(R^\{\(i\)\}\-\\mu\)/\\sigma, whereμ\\muandσ\\sigmaare the mean and standard deviation of\{R\(j\)\}j=1G\\\{R^\{\(j\)\}\\\}\_\{j=1\}^\{G\}\. This advantage is broadcast to all policy\-generated tokens inτ\(i\)\\tau^\{\(i\)\}, andπθ\\pi\_\{\\theta\}is optimized with the standard GRPO clipped objective with a KL penalty to the reference policy\(Shaoet al\.,[2024](https://arxiv.org/html/2606.02871#bib.bib2)\)\.

## 4Experiment Setting

We evaluateALARin two agentic domains,*search*and*tool use*\. We first describe the implementation details on both domains and then illustrate the evaluation setup\.

### 4\.1Implementation

Models and Datasets\.In the search domain, we use the released Search\-R1Jinet al\.\([2025](https://arxiv.org/html/2606.02871#bib.bib5)\)3B and 7B checkpoints as the base LRM, which are RL\-trained on NQKwiatkowskiet al\.\([2019](https://arxiv.org/html/2606.02871#bib.bib20)\)and HotpotQAYanget al\.\([2018](https://arxiv.org/html/2606.02871#bib.bib18)\)\. We roll out each base model in explicit mode on its training pool and keep only successful trajectories using exact\-matching rejection sampling, yielding 86K trajectories for the 7B model and 76K for the 3B model\.

In the tool\-use domain, we use Qwen3\-4B\-ThinkingYanget al\.\([2025a](https://arxiv.org/html/2606.02871#bib.bib39)\), a 4B LRM with native tool\-calling capability\. Given a query and a set of candidate tools in the system prompt, the model emits a multi\-step tool\-calling trajectory in a single assistant turn, interleaving reasoning with JSON function calls\. Teacher trajectories are collected from thegraph\_synsubset of ToolMindYanget al\.\([2025b](https://arxiv.org/html/2606.02871#bib.bib22)\), and we retain only rollouts whose tool calls exactly match the reference calls under AST\-level matching, resulting in 21K teacher rollouts\. For AR\-GRPO, we useG=8G=8,α=0\.3\\alpha=0\.3in both domains and set generous generation length tolerances ofL=400L=400for search andL=1600L=1600for tool use\.

NQHotpotQATriviaQA2WikiMuSiQueBamboogleAvg\.MethodEMTokAE\\mathrm\{AE\}EMTokAE\\mathrm\{AE\}EMTokAE\\mathrm\{AE\}EMTokAE\\mathrm\{AE\}EMTokAE\\mathrm\{AE\}EMTokAE\\mathrm\{AE\}EMTokAE\\mathrm\{AE\}Qwen2\.5\-3BSearch\-R142\.91380\.0037\.41580\.0061\.31430\.0039\.61720\.0014\.61740\.0033\.61410\.0038\.21540\.00ShorterBetter41\.3132\-0\.1436\.9150\-0\.0260\.2137\-0\.0539\.0166\-0\.0415\.3168\+0\.1833\.6129\+0\.0937\.71470\.00ThinkPrune41\.2132\-0\.1537\.01500\.0060\.3137\-0\.0438\.8166\-0\.0715\.3168\+0\.1834\.4130\+0\.1537\.8147\+0\.01O1\-Pruner19\.3149\-2\.8319\.8158\-2\.3541\.7155\-1\.6826\.1168\-1\.685\.2167\-3\.1828\.8137\-0\.6923\.5156\-2\.07ALARStage 141\.474\+0\.2938\.091\+0\.4755\.9113\-0\.2338\.594\+0\.3114\.8105\+0\.4435\.287\+0\.5337\.394\+0\.30Stage 241\.3106\+0\.0538\.0125\+0\.2660\.1115\+0\.1039\.0130\+0\.1715\.3138\+0\.3536\.8106\+0\.5338\.4120\+0\.24Qwen2\.5\-7BSearch\-R149\.12050\.0043\.22500\.0063\.82340\.0040\.12570\.0019\.12490\.0040\.82180\.0042\.72360\.00ShorterBetter46\.6190\-0\.1842\.8232\+0\.0364\.4215\+0\.1141\.3233\+0\.1818\.5231\-0\.0837\.6196\-0\.2941\.9216\-0\.04ThinkPrune46\.8191\-0\.1742\.8236\+0\.0164\.3218\+0\.0942\.0238\+0\.2218\.4234\-0\.1237\.6200\-0\.3142\.0220\-0\.05O1\-Pruner46\.8190\-0\.1642\.8227\+0\.0564\.3213\+0\.1141\.0228\+0\.1818\.8228\+0\.0136\.8192\-0\.3741\.8213\-0\.03ALARStage 146\.3106\+0\.2042\.0112\+0\.4163\.5112\+0\.5038\.8129\+0\.3417\.7115\+0\.1734\.497\-0\.2340\.5112\+0\.23Stage 246\.8123\+0\.1742\.8140\+0\.3964\.3133\+0\.4639\.6144\+0\.3818\.4135\+0\.2738\.4124\+0\.1441\.7133\+0\.30Table 1:Evaluation results on the search domain using Search\-R1 as the base model, at the Qwen2\.5\-3B and Qwen2\.5\-7B scales\.SimpleMultipleParallelPar\.\-Mult\.Avg\.MethodAccTokAE\\mathrm\{AE\}AccTokAE\\mathrm\{AE\}AccTokAE\\mathrm\{AE\}AccTokAE\\mathrm\{AE\}AccTokAE\\mathrm\{AE\}Qwen3\-4B\-ThinkingQwen3\-4B91\.55640\.0092\.05010\.0085\.08910\.0076\.010830\.0086\.17600\.00ShorterBetter93\.0109\+0\.8690\.0104\+0\.6882\.0201\+0\.6075\.5233\+0\.7585\.1162\+0\.72ThinkPrune93\.5169\+0\.7789\.5159\+0\.5589\.5325\+0\.7981\.0349\+0\.8788\.4251\+0\.75O1\-Pruner92\.5100\+0\.8589\.599\+0\.6787\.5195\+0\.8775\.5230\+0\.7586\.2156\+0\.79ALARStage 193\.851\+0\.9890\.050\+0\.7990\.0116\+1\.0581\.5130\+1\.1088\.887\+0\.98Stage 294\.287\+0\.9490\.584\+0\.7589\.5145\+1\.0082\.5152\+1\.1289\.2117\+0\.95Table 2:Evaluation results on the BFCL benchmark using Qwen3\-4B\-Thinking as the base model\.Latent block\.Each latent block consists ofK=4K=4continuous thoughts framed by surface tags<LAT\>…</LAT\>\. The four placeholders are repurposed as content\-free sentinels: at every<LAT\>, the projectorfϕf\_\{\\phi\}writesKKcontinuous embeddings into these positions, and the closing</LAT\>is prefilled after theKKprojections programmatically\. The tags are standard tokens in the vocabulary of the model, so no new special token is added\. The projector is a two\-layer MLP with GELU and a final LayerNorm whose hidden width matches the base model\.

Training\.Both domains follow the same two\-stage training pipeline\. Stage 1 trains the latent mode with AASD on successful teacher trajectories, replacing each teacher reasoning span with<LAT\>followed by a length\-KKlatent block and supervising only the subsequent anchor actions\. For the mode warmup, we first perform a brief SFT from the Stage 1 checkpoint using 20K instances, where each turn is randomly assigned to either latent or explicit thinking with equal probability\. We then optimize it with the AR\-GRPO objective as the Stage 2\.

### 4\.2Evaluation Setup

Benchmarks\.For the search domain, we evaluate on six open\-domain QA benchmarks: NQ, HotpotQA, TriviaQAJoshiet al\.\([2017](https://arxiv.org/html/2606.02871#bib.bib19)\), 2WikiMultiHopQA \(2Wiki;Hoet al\.[2020](https://arxiv.org/html/2606.02871#bib.bib17)\), MuSiQueTrivediet al\.\([2022](https://arxiv.org/html/2606.02871#bib.bib15)\), and BambooglePresset al\.\([2023](https://arxiv.org/html/2606.02871#bib.bib16)\)\. For the tool\-use domain, we evaluate on the AST\-based BFCLPatilet al\.\([2025](https://arxiv.org/html/2606.02871#bib.bib23)\)on all categories: simple, multiple, parallel, and parallel\-multiple\.

Evaluation metrics\.We report task accuracy \(EM\\mathrm\{EM\}\), average number of generated tokens \(Tok\\mathrm\{Tok\}\), and an Accuracy\-Efficiency \(AE\\mathrm\{AE\}\) score followingLuoet al\.\([2025](https://arxiv.org/html/2606.02871#bib.bib11)\)\.EM\\mathrm\{EM\}is exact\-match accuracy for search and AST\-level tool\-call matching for tool use\.Tok\\mathrm\{Tok\}counts the model\-generated tokens, including reasoning traces, mode tags, tool calls, and final answers\.AE\\mathrm\{AE\}summarizes the accuracy\-efficiency trade\-off relative to the corresponding base model\. Specifically, we computeAE=α​ΔTok\+β​\[ΔEM\]\+\+γ​\[ΔEM\]−\\mathrm\{AE\}=\\alpha\\Delta\_\{\\mathrm\{Tok\}\}\+\\beta\[\\Delta\_\{\\mathrm\{EM\}\}\]\_\{\+\}\+\\gamma\[\\Delta\_\{\\mathrm\{EM\}\}\]\_\{\-\}, whereΔTok=\(Tok0−Tok\)/Tok0\\Delta\_\{\\mathrm\{Tok\}\}=\(\\mathrm\{Tok\}\_\{0\}\-\\mathrm\{Tok\}\)/\\mathrm\{Tok\}\_\{0\},ΔEM=\(EM−EM0\)/EM0\\Delta\_\{\\mathrm\{EM\}\}=\(\\mathrm\{EM\}\-\\mathrm\{EM\}\_\{0\}\)/\\mathrm\{EM\}\_\{0\},\[x\]\+=max⁡\(0,x\)\[x\]\_\{\+\}=\\max\(0,x\), and\[x\]−=min⁡\(0,x\)\[x\]\_\{\-\}=\\min\(0,x\)\. Here,EM0\\mathrm\{EM\}\_\{0\}andTok0\\mathrm\{Tok\}\_\{0\}denote the accuracy and length of the corresponding base model\. FollowingLuoet al\.\([2025](https://arxiv.org/html/2606.02871#bib.bib11)\), we set\(α,β,γ\)=\(1,3,5\)\(\\alpha,\\beta,\\gamma\)=\(1,3,5\)to penalize accuracy degradation more strongly than accuracy improvement\.

Baselines\.In each domain, we compareALARagainst three published reasoning\-token compression methods:*O1\-Pruner*\(Luoet al\.,[2025](https://arxiv.org/html/2606.02871#bib.bib11)\), which rewards concise rollouts relative to a reference baseline;*ThinkPrune*\(Houet al\.,[2025](https://arxiv.org/html/2606.02871#bib.bib12)\), which enforces annealed length budgets on thinking spans; and*ShorterBetter*\(Yiet al\.,[2026](https://arxiv.org/html/2606.02871#bib.bib13)\), which encourages rollouts to match the shortest correct reasoning length in each group\.

## 5Experiment Results

### 5\.1Main Results

[Table˜1](https://arxiv.org/html/2606.02871#S4.T1)and[Table˜2](https://arxiv.org/html/2606.02871#S4.T2)report results on the search and tool\-use domains\. Overall,ALARachieves the best accuracy–efficiency trade\-off across both domains: it matches or improves the base model’s EM while substantially reducing generated tokens, yielding the strongest AE Pareto performance\.

ALARachieves strong token reduction while preserving accuracy\.In the search domain,ALARsubstantially reduces generation with little or no accuracy loss\. For 3B, Stage 2 improves average EM from38\.238\.2to38\.438\.4while reducing generated tokens by22\.1%22\.1\\%\. Stage 1 is even more efficient, reducing tokens by39\.0%39\.0\\%with competitive EM\. For 7B, Stage 2 reduces tokens by43\.6%43\.6\\%while maintaining comparable EM, and Stage 1 achieves a52\.5%52\.5\\%token reduction\.

Text\-based reasoning compression has limited headroom for search agents\.The reasoning token reduction baselines provide only modest gains in the search domain\. Since Search\-R1 already produces compact explicit CoT, methods that only shorten textual reasoning reduce average tokens by about44–10%10\\%\. In contrast,ALARchanges the reasoning interface itself by replacing textual reasoning with latent reasoning, enabling much larger reductions without severe performance degradation\.

ALARis especially effective in the tool\-use domain\.In tool use, the base Qwen3\-4B\-Thinking model is much more verbose than Search\-R1\. All compression baselines therefore achieve substantial token reductions, butALARperforms best\. Stage 1 improves average accuracy from86\.186\.1to88\.888\.8while reducing generated tokens by88\.6%88\.6\\%\. Stage 2 further improves accuracy to89\.289\.2while reducing tokens by84\.6%84\.6\\%, achieving the best EM and the strongest AE score among all methods\.

Stage 1 shows the strength of AASD, while Stage 2 improves adaptivity\.Stage 1 is highly competitive despite using the fewest tokens, showing that AASD effectively injects latent reasoning into agentic policies\. Since Stage 1 uses latent reasoning for every turn, its strong performance suggests that many agentic decisions do not require explicit CoT\. Stage 2 generally uses more tokens but improves EM by learning adaptive mode selection through AR\-GRPO, allowing the model to use latent reasoning for easier turns and explicit CoT for harder ones\.

Overall, adaptive latent reasoning outperforms explicit reasoning compression\.These results support two hypotheses behindALAR\. First, agentic reasoning is often unnecessarily verbose: many turns only require enough internal computation to select the next action, not a full explicit CoT trace\. Second, reasoning demand is heterogeneous across turns: while routine turns can be handled with compact latent reasoning, harder turns still benefit from explicit CoT\. Rather than uniformly compressing every textual reasoning trace,ALARlearns when explicit reasoning is necessary and uses latent reasoning otherwise\. This leads to comparable or better EM, much lower token usage, and the best accuracy–efficiency trade\-off across both search and tool\-use domains\.

### 5\.2Analysis of Adaptive Mode Selection

We examine the adaptive mode selection behavior ofALARin the search domain by analyzing per\-turn latent and explicit reasoning choices in the 7B evaluation trajectories \([Table˜3](https://arxiv.org/html/2606.02871#S5.T3)\)\.

Harder benchmarks retain more explicit reasoning\.Although latent reasoning dominates overall,ALARuses explicit reasoning more often on harder benchmarks\. After AR\-GRPO, Bamboogle and 2Wiki have the lowest latent fractions \(59%59\\%and75%75\\%\), while easier single\-hop datasets such as NQ and TriviaQA rely on latent reasoning much more frequently \(89%89\\%and91%91\\%\)\. Compared with the warmed\-up policy, AR\-GRPO increases latent usage when it is sufficient, as in TriviaQA \(69%→91%69\\%\\\!\\to\\\!91\\%\) and NQ \(84%→89%84\\%\\\!\\to\\\!89\\%\), but keeps it similar or lower on harder datasets such as MuSiQue, 2Wiki, and Bamboogle\. This suggests that AR\-GRPO learns a task\-conditioned mode\-selection policy rather than uniformly increasing latent reasoning\.

Per\-turn Latent FractionTotalDatasetT1T2T3T4≥\\geq5WarmupGRPOMuSiQue738587887888832Wiki44758487937575HotpotQA56879294948584NQ72899595948489TriviaQA73959798956991Bamboogle75394964746859Table 3:Per\-turn latent fraction on the search domain over all 7B trajectories\. Turn\-level columns report the latent fraction after AR\-GRPO;*Total*compares the overall latent fraction before and after AR\-GRPO\.Explicit reasoning is concentrated in early planning turns\.Turn 1 consistently uses the most explicit reasoning, while later turns are mostly latent\. This pattern suggests thatALARuses explicit CoT primarily for initial planning, such as decomposing a comparison or compositional question into sub\-goals before issuing the first search action\. After the initial plan is formed, subsequent retrieve\-and\-gather turns require less textual deliberation and can usually be handled through latent reasoning\. This turn\-level behavior supports the design motivation ofALAR: explicit reasoning is most useful when the agent must plan or decompose the task, whereas latent reasoning is sufficient for many routine environment\-interaction steps\.

*Search**Tool\-use*Latent StepsKKEMTokAEEMTokAE1140\.6128\+0\.2187\.4109\+0\.902241\.2131\+0\.2788\.6114\+0\.944441\.7133\+0\.3289\.2117\+0\.958841\.8138\+0\.3189\.1124\+0\.94Table 4:Effect of the number of latent stepsKK\. Search results are averaged over the search benchmarks using the 7B model, and tool\-use results are averaged over the BFCL categories\.
### 5\.3Effect of Latent Step Length

A few latent steps are sufficient for agentic reasoning\.We ablate the number of continuous latent stepsKKused in each latent block in[Table˜4](https://arxiv.org/html/2606.02871#S5.T4)\. Even with a small number of latent steps,ALARachieves strong accuracy\-efficiency trade\-offs in both domains\. IncreasingKKfrom11to44improves EM and AE, showing that the latent mode benefits from a modest amount of internal computation\. However, performance changes little beyondK=4K=4: usingK=8K=8yields negligible accuracy gains while slightly increasing token usage\. This suggests that most agentic turns require only compact latent computation to select the next action or incorporate observations, while turns requiring deeper reasoning can still be handled by the explicit mode\. We therefore useK=4K=4as the default because it is sufficient to capture most of the benefit of latent reasoning\.

## 6Conclusion

We introducedALAR, an adaptive latent reasoning framework for efficient LLM agents\. Instead of merely shortening explicit CoT,ALARchanges the reasoning interface: agents use latent reasoning for routine turns and reserve explicit CoT for harder ones\. It is trained with Action\-Anchored Self\-Distillation, which teaches latent reasoning from successful agent actions, and AR\-GRPO, which learns when latent reasoning is sufficient for task success\. Across search and tool\-use domains,ALARachieves comparable or better accuracy with substantially fewer generated tokens than base models and reasoning compression baselines\. These findings highlight adaptive use of latent and explicit reasoning as a practical path toward more efficient LLM agents\.

## Limitations

This work has several limitations\. First, we focus on agentic tasks that require interaction with an external environment, such as search and tool use\. We do not evaluate on domains such as math and coding, which have been heavily used in reasoning post\-training and often require less environment interaction\. In this paper, we view such settings as closer to single\-pass LLM reasoning, where additional reasoning tokens may directly improve final\-answer accuracy\. By contrast, in agentic settings, more generated reasoning does not necessarily lead to better performance, since many turns mainly require selecting the next environment\-coupled action\.

Second,ALARrelies on successful teacher trajectories for Action\-Anchored Self\-Distillation, so the latent policy may inherit the coverage and biases of the teacher\. Third, we use a fixed latent block lengthKKand a discrete latent/explicit mode choice, leaving more fine\-grained control of latent computation to future work\. Finally, latent reasoning reduces generated tokens but makes part of the agent’s reasoning less interpretable, which may be undesirable when action transparency is important\.

## Ethics Statement

This work follows the ACL Code of Ethics\. We use existing public benchmarks and do not collect new human\-subject data\. The main ethical consideration is that latent reasoning may reduce the transparency of intermediate reasoning compared with explicit CoT\. We therefore recommend monitoring agent actions and outputs, and retaining explicit reasoning or additional logging in high\-stakes settings\.

## References

- L1: controlling how long a reasoning model thinks with reinforcement learning\.arXiv preprint arXiv:2503\.04697\.Cited by:[§2\.2](https://arxiv.org/html/2606.02871#S2.SS2.p1.1)\.
- D\. Arora and A\. Zanette \(2026\)Training language models to reason efficiently\.Advances in Neural Information Processing Systems38,pp\. 60770–60808\.Cited by:[§2\.2](https://arxiv.org/html/2606.02871#S2.SS2.p1.1)\.
- X\. Chen, J\. Xu, T\. Liang, Z\. He, J\. Pang, D\. Yu, L\. Song, Q\. Liu, M\. Zhou, Z\. Zhang,et al\.\(2024\)Do not think that much for 2\+ 3=? on the overthinking of o1\-like llms\.arXiv preprint arXiv:2412\.21187\.Cited by:[§1](https://arxiv.org/html/2606.02871#S1.p2.1),[§2\.2](https://arxiv.org/html/2606.02871#S2.SS2.p1.1)\.
- J\. Cheng and B\. Van Durme \(2024\)Compressed chain of thought: efficient reasoning through dense representations\.arXiv preprint arXiv:2412\.13171\.Cited by:[§2\.1](https://arxiv.org/html/2606.02871#S2.SS1.p1.1)\.
- Z\. Cheng, D\. Chen, M\. Fu, and T\. Zhou \(2025\)Optimizing length compression in large reasoning models\.arXiv preprint arXiv:2506\.14755\.Cited by:[§2\.2](https://arxiv.org/html/2606.02871#S2.SS2.p1.1)\.
- Y\. Deng, Y\. Choi, and S\. Shieber \(2024\)From explicit cot to implicit cot: learning to internalize cot step by step\.arXiv preprint arXiv:2405\.14838\.Cited by:[§2\.1](https://arxiv.org/html/2606.02871#S2.SS1.p1.1)\.
- M\. Douze, A\. Guzhva, C\. Deng, J\. Johnson, G\. Szilvasy, P\. Mazaré, M\. Lomeli, L\. Hosseini, and H\. Jégou \(2024\)The faiss library\.External Links:2401\.08281Cited by:[§A\.2](https://arxiv.org/html/2606.02871#A1.SS2.p1.1)\.
- D\. Guo, D\. Yang, H\. Zhang, J\. Song, P\. Wang, Q\. Zhu, R\. Xu, R\. Zhang, S\. Ma, X\. Bi,et al\.\(2025\)Deepseek\-r1: incentivizing reasoning capability in llms via reinforcement learning\.arXiv preprint arXiv:2501\.12948\.Cited by:[§1](https://arxiv.org/html/2606.02871#S1.p1.1),[§3\.1](https://arxiv.org/html/2606.02871#S3.SS1.p1.11)\.
- S\. Hao, S\. Sukhbaatar, D\. Su, X\. Li, Z\. Hu, J\. Weston, and Y\. Tian \(2024\)Training large language models to reason in a continuous latent space\.arXiv preprint arXiv:2412\.06769\.Cited by:[§1](https://arxiv.org/html/2606.02871#S1.p4.1),[§2\.1](https://arxiv.org/html/2606.02871#S2.SS1.p1.1),[§3\.2](https://arxiv.org/html/2606.02871#S3.SS2.p1.1)\.
- X\. Ho, A\. D\. Nguyen, S\. Sugawara, and A\. Aizawa \(2020\)Constructing a multi\-hop qa dataset for comprehensive evaluation of reasoning steps\.InProceedings of the 28th International Conference on Computational Linguistics,pp\. 6609–6625\.Cited by:[§4\.2](https://arxiv.org/html/2606.02871#S4.SS2.p1.1)\.
- B\. Hou, Y\. Zhang, J\. Ji, Y\. Liu, K\. Qian, J\. Andreas, and S\. Chang \(2025\)Thinkprune: pruning long chain\-of\-thought of llms via reinforcement learning\.arXiv preprint arXiv:2504\.01296\.Cited by:[§1](https://arxiv.org/html/2606.02871#S1.p3.1),[§2\.2](https://arxiv.org/html/2606.02871#S2.SS2.p1.1),[§4\.2](https://arxiv.org/html/2606.02871#S4.SS2.p3.1)\.
- A\. Jaech, A\. Kalai, A\. Lerer, A\. Richardson, A\. El\-Kishky, A\. Low, A\. Helyar, A\. Madry, A\. Beutel, A\. Carney,et al\.\(2024\)Openai o1 system card\.arXiv preprint arXiv:2412\.16720\.Cited by:[§1](https://arxiv.org/html/2606.02871#S1.p1.1)\.
- B\. Jin, H\. Zeng, Z\. Yue, J\. Yoon, S\. Arik, D\. Wang, H\. Zamani, and J\. Han \(2025\)Search\-r1: training llms to reason and leverage search engines with reinforcement learning\.arXiv preprint arXiv:2503\.09516\.Cited by:[§3\.4](https://arxiv.org/html/2606.02871#S3.SS4.p5.16),[§4\.1](https://arxiv.org/html/2606.02871#S4.SS1.p1.1)\.
- M\. Joshi, E\. Choi, D\. S\. Weld, and L\. Zettlemoyer \(2017\)TriviaQA: a large scale distantly supervised challenge dataset for reading comprehension\.InProceedings of the 55th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 1601–1611\.Cited by:[§4\.2](https://arxiv.org/html/2606.02871#S4.SS2.p1.1)\.
- T\. Kwiatkowski, J\. Palomaki, O\. Redfield, M\. Collins, A\. Parikh, C\. Alberti, D\. Epstein, I\. Polosukhin, J\. Devlin, K\. Lee,et al\.\(2019\)Natural questions: a benchmark for question answering research\.Transactions of the Association for Computational Linguistics7,pp\. 452–466\.Cited by:[§4\.1](https://arxiv.org/html/2606.02871#S4.SS1.p1.1)\.
- H\. Luo, L\. Shen, H\. He, Y\. Wang, S\. Liu, W\. Li, N\. Tan, X\. Cao, and D\. Tao \(2025\)O1\-pruner: length\-harmonizing fine\-tuning for o1\-like reasoning pruning\.arXiv preprint arXiv:2501\.12570\.Cited by:[§1](https://arxiv.org/html/2606.02871#S1.p3.1),[§2\.2](https://arxiv.org/html/2606.02871#S2.SS2.p1.1),[§4\.2](https://arxiv.org/html/2606.02871#S4.SS2.p2.14),[§4\.2](https://arxiv.org/html/2606.02871#S4.SS2.p3.1)\.
- S\. G\. Patil, H\. Mao, F\. Yan, C\. C\. Ji, V\. Suresh, I\. Stoica, and J\. E\. Gonzalez \(2025\)The berkeley function calling leaderboard \(bfcl\): from tool use to agentic evaluation of large language models\.InInternational Conference on Machine Learning,pp\. 48371–48392\.Cited by:[§4\.2](https://arxiv.org/html/2606.02871#S4.SS2.p1.1)\.
- O\. Press, M\. Zhang, S\. Min, L\. Schmidt, N\. A\. Smith, and M\. Lewis \(2023\)Measuring and narrowing the compositionality gap in language models\.InFindings of the Association for Computational Linguistics: EMNLP 2023,pp\. 5687–5711\.Cited by:[§4\.2](https://arxiv.org/html/2606.02871#S4.SS2.p1.1)\.
- Z\. Shao, P\. Wang, Q\. Zhu, R\. Xu, J\. Song, X\. Bi, H\. Zhang, M\. Zhang, Y\. Li, Y\. Wu,et al\.\(2024\)Deepseekmath: pushing the limits of mathematical reasoning in open language models\.arXiv preprint arXiv:2402\.03300\.Cited by:[§3\.5](https://arxiv.org/html/2606.02871#S3.SS5.p7.6)\.
- Y\. Shen, J\. Zhang, J\. Huang, S\. Shi, W\. Zhang, J\. Yan, N\. Wang, K\. Wang, Z\. Liu, and S\. Lian \(2025a\)Dast: difficulty\-adaptive slow\-thinking for large reasoning models\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: Industry Track,pp\. 2322–2331\.Cited by:[§2\.2](https://arxiv.org/html/2606.02871#S2.SS2.p1.1)\.
- Z\. Shen, H\. Yan, L\. Zhang, Z\. Hu, Y\. Du, and Y\. He \(2025b\)Codi: compressing chain\-of\-thought into continuous space via self\-distillation\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,pp\. 677–693\.Cited by:[§1](https://arxiv.org/html/2606.02871#S1.p4.1),[§2\.1](https://arxiv.org/html/2606.02871#S2.SS1.p1.1),[§3\.2](https://arxiv.org/html/2606.02871#S3.SS2.p1.1)\.
- D\. Shi, A\. Asi, K\. Li, X\. Yuan, L\. Pan, W\. Lee, and W\. Xiao \(2025\)SwiReasoning: switch\-thinking in latent and explicit for pareto\-superior reasoning llms\.arXiv preprint arXiv:2510\.05069\.Cited by:[§2\.1](https://arxiv.org/html/2606.02871#S2.SS1.p1.1)\.
- N\. Shinn, F\. Cassano, A\. Gopinath, K\. Narasimhan, and S\. Yao \(2023\)Reflexion: language agents with verbal reinforcement learning\.Advances in neural information processing systems36,pp\. 8634–8652\.Cited by:[§1](https://arxiv.org/html/2606.02871#S1.p1.1)\.
- D\. Su, H\. Zhu, Y\. Xu, J\. Jiao, Y\. Tian, and Q\. Zheng \(2025\)Token assorted: mixing latent and text tokens for improved language model reasoning\.arXiv preprint arXiv:2502\.03275\.Cited by:[§2\.1](https://arxiv.org/html/2606.02871#S2.SS1.p1.1)\.
- Y\. Sui, Y\. Chuang, G\. Wang, J\. Zhang, T\. Zhang, J\. Yuan, H\. Liu, A\. Wen, S\. Zhong, N\. Zou,et al\.\(2025\)Stop overthinking: a survey on efficient reasoning for large language models\.arXiv preprint arXiv:2503\.16419\.Cited by:[§2\.2](https://arxiv.org/html/2606.02871#S2.SS2.p1.1)\.
- H\. Trivedi, N\. Balasubramanian, T\. Khot, and A\. Sabharwal \(2022\)MuSiQue: multi\-hop questions via single\-hop question composition\.Transactions of the Association for Computational Linguistics10,pp\. 539–554\.Cited by:[§4\.2](https://arxiv.org/html/2606.02871#S4.SS2.p1.1)\.
- T\. Wei, T\. Li, Z\. Liu, X\. Ning, Z\. Yang, J\. Zou, Z\. Zeng, R\. Qiu, X\. Lin, D\. Fu,et al\.\(2026\)Agentic reasoning for large language models\.arXiv preprint arXiv:2601\.12538\.Cited by:[§1](https://arxiv.org/html/2606.02871#S1.p1.1)\.
- V\. Xiang, C\. Snell, K\. Gandhi, A\. Albalak, A\. Singh, C\. Blagden, D\. Phung, R\. Rafailov, N\. Lile, D\. Mahan,et al\.\(2025\)Towards system 2 reasoning in llms: learning how to think with meta chain\-of\-thought\.arXiv preprint arXiv:2501\.04682\.Cited by:[§3\.1](https://arxiv.org/html/2606.02871#S3.SS1.p1.11)\.
- X\. Xu, T\. Yu, X\. Chen, H\. Wang, J\. McAuley, and S\. Mitra \(2026\)ThinkRouter: efficient reasoning via routing thinking between latent and discrete spaces\.arXiv preprint arXiv:2602\.11683\.Cited by:[§2\.1](https://arxiv.org/html/2606.02871#S2.SS1.p1.1)\.
- A\. Yang, A\. Li, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Gao, C\. Huang, C\. Lv,et al\.\(2025a\)Qwen3 technical report\.arXiv preprint arXiv:2505\.09388\.Cited by:[§4\.1](https://arxiv.org/html/2606.02871#S4.SS1.p2.4)\.
- C\. Yang, R\. Le, Y\. Xing, Z\. An, Z\. Chen, W\. X\. Zhao, Y\. Song, and T\. Zhang \(2025b\)ToolMind technical report: a large\-scale, reasoning\-enhanced tool\-use dataset\.arXiv preprint arXiv:2511\.15718\.Cited by:[§4\.1](https://arxiv.org/html/2606.02871#S4.SS1.p2.4)\.
- Z\. Yang, P\. Qi, S\. Zhang, Y\. Bengio, W\. Cohen, R\. Salakhutdinov, and C\. D\. Manning \(2018\)HotpotQA: a dataset for diverse, explainable multi\-hop question answering\.InProceedings of the 2018 Conference on Empirical Methods in Natural Language Processing,pp\. 2369–2380\.Cited by:[§4\.1](https://arxiv.org/html/2606.02871#S4.SS1.p1.1)\.
- S\. Yao, J\. Zhao, D\. Yu, N\. Du, I\. Shafran, K\. Narasimhan, and Y\. Cao \(2022\)React: synergizing reasoning and acting in language models\.arXiv preprint arXiv:2210\.03629\.Cited by:[§1](https://arxiv.org/html/2606.02871#S1.p1.1)\.
- J\. Yi, J\. Wang, and S\. Li \(2026\)Shorterbetter: guiding reasoning models to find optimal inference length for efficient reasoning\.Advances in Neural Information Processing Systems38,pp\. 39011–39043\.Cited by:[§1](https://arxiv.org/html/2606.02871#S1.p3.1),[§2\.2](https://arxiv.org/html/2606.02871#S2.SS2.p1.1),[§4\.2](https://arxiv.org/html/2606.02871#S4.SS2.p3.1)\.
- Z\. Yue, B\. Jin, H\. Zeng, H\. Zhuang, Z\. Qin, J\. Yoon, L\. Shang, J\. Han, and D\. Wang \(2026\)Hybrid latent reasoning via reinforcement learning\.Advances in Neural Information Processing Systems38,pp\. 5501–5530\.Cited by:[§2\.1](https://arxiv.org/html/2606.02871#S2.SS1.p1.1)\.

## Appendix AImplementation Details

### A\.1Models

#### Search domain\.

#### Tool\-use domain\.

### A\.2Search Environment

Following Search\-R1, the search agent interleaves reasoning with retrieval over the Wikipedia\-18 corpus\. Retrieval uses a FAISSDouzeet al\.\([2024](https://arxiv.org/html/2606.02871#bib.bib14)\)index built on E5\-large\-v2444[https://huggingface\.co/intfloat/e5\-large\-v2](https://huggingface.co/intfloat/e5-large-v2)embeddings and returns the top\-3 documents per query\. Each search trajectory is capped at six turns\. During supervision trajectory construction, we cap each retrieved document at 500 characters and the total context length at 4096 tokens\.

### A\.3Latent Block Implementation

Each latent block containsK=4K=4continuous thoughts framed by surface tags<LAT\>and</LAT\>\. The latent positions are implemented as content\-free sentinel placeholders: when the decoder reaches<LAT\>, the projectorfϕf\_\{\\phi\}writesKKcontinuous embeddings into the subsequent latent positions, and</LAT\>is prefilled after theKKprojections\. The tags are standard vocabulary tokens, so no new special tokens are added\. The projectorfϕf\_\{\\phi\}is a two\-layer MLP with GELU activation and a final LayerNorm, with hidden width matching the base model\.

### A\.4Training Details

For the supervised stages, we train for one epoch with LoRA at rank 16 andα=32\\alpha=32on the attention q/k/v/o projections, while the projectorfϕf\_\{\\phi\}is fully fine\-tuned\. We use AdamW with learning rate1×10−41\\times 10^\{\-4\}, 3% linear warmup followed by a cosine schedule, global batch size 32, bf16 precision, and gradient checkpointing\. For AR\-GRPO, we continue to use LoRA with the same rank and target modules, while optimizing the projectorfϕf\_\{\\phi\}together with the policy\. We use latent bonus coefficientα=0\.3\\alpha=0\.3in both domains, a KL coefficient of10−310^\{\-3\}against the SFT reference policy, and length tolerances ofL=400L=400for search andL=1600L=1600for tool use\.

### A\.5Infrastructure

We use theverlframework with FSDP for the actor and vLLM for rollout generation\. The rollout engine is augmented with a state machine that detects<LAT\>, splices projector outputs into the nextKKsentinel positions, and hot\-reloads the projector between rollout steps\.

### A\.6Prompts

[Table˜5](https://arxiv.org/html/2606.02871#A4.T5)specifies the domain\-specific system prompts that are used for training and inference\.

## Appendix BBaseline Implementation Details

#### Shared setting\.

All explicit\-CoT compression baselines use the same base model, retrieval environment, and evaluation setup asALAR\. Training usesverlwith FSDP for the actor and a co\-located vLLM rollout engine\. We use LoRA with rank 16 andα=32\\alpha=32on the attention q/k/v/o projections, AdamW with learning rate1×10−61\\times 10^\{\-6\}, KL coefficient10−310^\{\-3\}against the Search\-R1 actor as the reference policy, a batch of 12 prompts withG=8G=8rollouts per prompt, and 200 update steps in bf16 with gradient checkpointing\.

#### ShorterBetter\.

ShorterBetter uses a group\-relative length reward that encourages each rollout to approach the shortest correct reasoning length within its group\. We use EM weightα=2\.0\\alpha=2\.0and length\-penalty weightβ=5×10−3\\beta=5\\times 10^\{\-3\}, computed against the shortest correct rollout amongG=8G=8rollouts\.

#### ThinkPrune\.

ThinkPrune applies a staged annealing schedule over the thinking\-token budget\. We use three stages with budgetsT∈\{200,120,70\}T\\in\\\{200,120,70\\\}, corresponding to the p75, p50, and p25 thinking\-length percentiles of the Search\-R1 actor’s training pool\.

#### O1\-Pruner\.

O1\-Pruner uses per\-prompt reference statistics for length\-harmonizing reward computation\. We precompute\(Lref​\(x\),Aref​\(x\)\)\(L\_\{\\mathrm\{ref\}\}\(x\),A\_\{\\mathrm\{ref\}\}\(x\)\)fromK=8K=8rollouts at temperature 0\.7 over a 3,200\-prompt training sub\-pool using the frozen Search\-R1 actor\. The online reward uses EM weightα=2\.0\\alpha=2\.0and clip bounds\(clo,chi\)=\(−4\.0,\+2\.0\)\(c\_\{\\mathrm\{lo\}\},c\_\{\\mathrm\{hi\}\}\)=\(\-4\.0,\+2\.0\)\. Out\-of\-pool prompts fall back to the EM\-only reward during training\.

## Appendix CUse of AI Assistants

We used Claude Code to assist with implementation and experimentation\. We also used ChatGPT to help revise sentences for grammar, clarity, and fluency\.

## Appendix DLicenses

- •Natural Questions \(NQ\): CC BY\-SA 3\.0
- •TriviaQA: Apache\-2\.0
- •HotpotQA: CC BY\-SA 4\.0
- •Wikipedia\-18 Corpus: Apache\-2\.0
- •MuSiQue: CC BY 4\.0
- •2WikiMultiHopQA \(2Wiki\): CC BY\-SA 4\.0
- •Bamboogle: MIT
- •BFCL: CC BY\-NC 4\.0
- •ToolMind: Apache\-2\.0
- •Search\-R1\-Qwen2\.5\-3B: Apache\-2\.0
- •Search\-R1\-Qwen2\.5\-7B: Apache\-2\.0
- •Qwen3\-4B\-Thinking\-2507: Apache\-2\.0
- •E5\-large\-v2: MIT

DomainSystem promptSearchAnswer the given question\. You must conduct reasoning first every time you get new information\. You may choose either mode per turn:<latent\>••••</latent\>— compact internal reasoning\. Emit exactly four bullet placeholder tokens between the tags; each carries one step of internal latent state\. Use by default for routine steps\.<think\> \.\.\. </think\>— explicit textual reasoning\. Use when you need to fuse information from multiple searches, when previous searches were insufficient, or when you need to reflect on your previous reasoning\. After reasoning, if you find you lack some knowledge, you can call a search engine by<search\> query </search\>and it will return the top searched results between<information\>and</information\>\. You can search as many times as you want\. If you find no further external knowledge needed, you can directly provide the answer inside<answer\>and</answer\>, without detailed illustrations\. For example,<answer\> Beijing </answer\>\.Tool\-useYou are a careful tool\-using assistant\. Before each action you must reason\. You may choose either mode per turn:<latent\>••••</latent\>— compact internal reasoning\. Emit exactly four bullet placeholder tokens between the tags; each carries one step of internal latent state\. Use by default for routine steps\.<think\> \.\.\. </think\>— explicit textual reasoning\. Use when you need to chain information from previous tool calls, recover from an unexpected response, or plan a multi\-step sequence\. After reasoning, either issue one or more<tool\_call\>calls, or produce the final natural\-language answer\.Table 5:System prompts used for the search and tool\-use domains\.

Similar Articles

Learning to Refine Hidden States for Reliable LLM Reasoning

arXiv cs.LG

Proposes ReLAR, a reinforcement-guided latent refinement framework that iteratively updates hidden representations in LLMs before decoding, improving reasoning reliability and efficiency compared to chain-of-thought methods.

Agentic Chain-of-Thought Steering for Efficient and Controllable LLM Reasoning

Hugging Face Daily Papers

ACTS (Agentic Chain-of-Thought Steering) formulates LLM reasoning control as a Markov decision process where a controller agent adaptively steers a frozen reasoner during inference using reasoning strategies and steering phrases. The approach achieves comparable accuracy to full-thinking models with significant token savings, enabling controllable accuracy-efficiency trade-offs.

BALAR : A Bayesian Agentic Loop for Active Reasoning

arXiv cs.AI

This paper introduces BALAR, a training-free Bayesian agentic loop algorithm that enables large language models to actively reason and ask clarifying questions in multi-turn interactions. It demonstrates significant performance improvements over baselines on detective, puzzle, and clinical diagnosis benchmarks.

ATLAS: Agentic or Latent Visual Reasoning? One Word is Enough for Both

Hugging Face Daily Papers

ATLAS presents a visual reasoning framework that combines agentic operations and latent representations using functional tokens, enabling efficient training via next-token prediction and reinforcement learning while avoiding intermediate image generation.