X-Tree: Tokenizing Reusable Experience for Efficient Agent Generalization

Hugging Face Daily Papers Papers

Summary

X-Tree recovers reusable skill hierarchies directly from agent trajectories (no LLM calls) and integrates them into offline RL, online RLVR, and on-policy self-distillation, improving success rates on WebArena, ScienceWorld, and WebShop by up to 5.8% over standard training recipes at matched data and budget.

Multi-step agents are trained on flat action streams: SFT and RLVR weight every token uniformly and ignore the sub-procedures that recur across tasks, the hierarchy that lets humans plan top-down from reusable routines. This structure sits unused, and flat training uses each scarce trajectory less fully than its content allows. Recent agents do use that structure, but only as LLM-written skills in context, never in the weights, so their gains do not generalize beyond retrieval. We instead recover this hierarchy from the data itself and train on it, with no LLM calls. Following text tokenizers, which build a vocabulary by counting alone, we score action spans by reusability and merge canonicalized actions into a reusable eXperience tree (X-Tree). Each X-Tree node captures how a frequent and success-bearing skill is composed from sub-skills, guiding efficient generalization. We integrate X-Tree into three training settings: offline RL, with each node as a training instance; online RLVR, with an adaptive skill bonus; and on-policy self-distillation, with X-Tree as the self-teacher's privileged context. Across WebArena, ScienceWorld, and WebShop at three model scales, X-Tree improves over standard recipes at matched data and budget by up to 4.5% SR on WebArena, 5.8% SR on ScienceWorld and 4.1% success on WebShop. Matched analyses attribute the gains to the X-Tree structure and the three integrations.
Original Article
View Cached Full Text

Cached at: 10/02/26, 08:31 PM

Paper page - X-Tree: Tokenizing Reusable Experience for Efficient Agent Generalization

Source: https://huggingface.co/papers/2609.32993

Abstract

Multi-stepagentsaretrainedonflatactionstreams:SFTandRLVRweighteverytokenuniformlyandignorethesub-proceduresthatrecuracrosstasks,thehierarchythatletshumansplantop-downfromreusableroutines.Thisstructuresitsunused,andflattraininguseseachscarcetrajectorylessfullythanitscontentallows.Recentagentsdousethatstructure,butonlyasLLM-writtenskillsincontext,neverintheweights,sotheirgainsdonotgeneralizebeyondretrieval.Weinsteadrecoverthishierarchyfromthedataitselfandtrainonit,withnoLLMcalls.Followingtexttokenizers,whichbuildavocabularybycountingalone,wescoreactionspansbyreusabilityandmergecanonicalizedactionsintoareusableeXperiencetree(X-Tree).EachX-Treenodecaptureshowafrequentandsuccess-bearingskilliscomposedfromsub-skills,guidingefficientgeneralization.WeintegrateX-Treeintothreetrainingsettings:offlineRL,witheachnodeasatraininginstance;onlineRLVR,withanadaptiveskillbonus;andon-policyself-distillation,withX-Treeastheself-teacher’sprivilegedcontext.AcrossWebArena,ScienceWorld,andWebShopatthreemodelscales,X-Treeimprovesoverstandardrecipesatmatcheddataandbudgetbyupto4.5%SRonWebArena,5.8%SRonScienceWorldand4.1%successonWebShop.MatchedanalysesattributethegainstotheX-Treestructureandthethreeintegrations.

View arXiv pageView PDFProject pageGitHub1Add to collection

Get this paper in your agent:

hf papers read 2609\.32993

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.32993 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.32993 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.32993 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

AREX: Towards a Recursively Self-Improving Agent for Deep Research

Hugging Face Daily Papers

AREX introduces a family of recursively self-improving agents for deep research, alternating between an inner research loop and an outer self-improvement loop, trained with long-horizon reinforcement learning. It substantially outperforms comparable-scale baselines on benchmarks like BrowseComp and Humanity's Last Exam.

Process Reward Informed Tree Rollout for Effective Multi-Turn RL

arXiv cs.CL

Proposes PaTR, a process-reward-guided adaptive tree rollout framework for multi-turn reinforcement learning in LLM agents. It selectively branches from promising states and prunes dead-end paths, achieving up to +5.0 on SWE-Bench and +9.3 on FrozenLake under the same training budget.