X-Tree: Tokenizing Reusable Experience for Efficient Agent Generalization
Summary
X-Tree recovers reusable skill hierarchies directly from agent trajectories (no LLM calls) and integrates them into offline RL, online RLVR, and on-policy self-distillation, improving success rates on WebArena, ScienceWorld, and WebShop by up to 5.8% over standard training recipes at matched data and budget.
View Cached Full Text
Cached at: 10/02/26, 08:31 PM
Paper page - X-Tree: Tokenizing Reusable Experience for Efficient Agent Generalization
Source: https://huggingface.co/papers/2609.32993
Abstract
Multi-stepagentsaretrainedonflatactionstreams:SFTandRLVRweighteverytokenuniformlyandignorethesub-proceduresthatrecuracrosstasks,thehierarchythatletshumansplantop-downfromreusableroutines.Thisstructuresitsunused,andflattraininguseseachscarcetrajectorylessfullythanitscontentallows.Recentagentsdousethatstructure,butonlyasLLM-writtenskillsincontext,neverintheweights,sotheirgainsdonotgeneralizebeyondretrieval.Weinsteadrecoverthishierarchyfromthedataitselfandtrainonit,withnoLLMcalls.Followingtexttokenizers,whichbuildavocabularybycountingalone,wescoreactionspansbyreusabilityandmergecanonicalizedactionsintoareusableeXperiencetree(X-Tree).EachX-Treenodecaptureshowafrequentandsuccess-bearingskilliscomposedfromsub-skills,guidingefficientgeneralization.WeintegrateX-Treeintothreetrainingsettings:offlineRL,witheachnodeasatraininginstance;onlineRLVR,withanadaptiveskillbonus;andon-policyself-distillation,withX-Treeastheself-teacher’sprivilegedcontext.AcrossWebArena,ScienceWorld,andWebShopatthreemodelscales,X-Treeimprovesoverstandardrecipesatmatcheddataandbudgetbyupto4.5%SRonWebArena,5.8%SRonScienceWorldand4.1%successonWebShop.MatchedanalysesattributethegainstotheX-Treestructureandthethreeintegrations.
View arXiv pageView PDFProject pageGitHub1Add to collection
Get this paper in your agent:
hf papers read 2609\.32993
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.32993 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.32993 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.32993 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Tree-of-Experience: A Structured Experience-Management Solution for Self-Evolving Agents under Low-Repetition and Implicit-Reward Environments
This paper introduces FinEvolveBench, a benchmark for financial sentiment prediction, and Tree-of-Experience (ToE), a structured experience-management method for LLM agents in low-repetition tasks with implicit rewards. Experiments show that ToE outperforms general-purpose experience mechanisms in such challenging settings.
RS-Claw: Progressive Active Tool Exploration via Hierarchical Skill Trees for Remote Sensing Agents
RS-Claw proposes an active tool exploration paradigm for remote sensing agents using hierarchical skill trees, enabling on-demand sequential decision-making and achieving up to 86% input token compression while outperforming passive selection baselines on Earth-Bench.
Exploit More, Explore Smarter for Budget-Constrained Agentic Search
This paper introduces ExTS, a tree-search policy for budget-constrained agentic search in LLM agents that improves over standard baselines in tasks like prompt optimization and code generation with an average +5.5% gain.
AREX: Towards a Recursively Self-Improving Agent for Deep Research
AREX introduces a family of recursively self-improving agents for deep research, alternating between an inner research loop and an outer self-improvement loop, trained with long-horizon reinforcement learning. It substantially outperforms comparable-scale baselines on benchmarks like BrowseComp and Humanity's Last Exam.
Process Reward Informed Tree Rollout for Effective Multi-Turn RL
Proposes PaTR, a process-reward-guided adaptive tree rollout framework for multi-turn reinforcement learning in LLM agents. It selectively branches from promising states and prunes dead-end paths, achieving up to +5.0 on SWE-Bench and +9.3 on FrozenLake under the same training budget.