Dependency-Aware Trajectory Refinement for Efficient Multi-Turn Agent Fine-Tuning

arXiv cs.CL Papers

Summary

This paper proposes a dependency-aware trajectory refinement method for efficient multi-turn agent fine-tuning, improving accuracy and reducing inference costs.

arXiv:2609.18417v1 Announce Type: new Abstract: Multi-turn agent trajectories often contain redundant rounds (failed tool calls, parallel sub-queries, verification-only steps) that inflate both training and inference cost. We propose viewing each trajectory as a \emph{round-level dependency DAG} that exposes which rounds are globally load-bearing for the final answer, and fine-tune agents on trajectories refined through this DAG. Given an LLM-annotated DAG, these edits are deterministic and interpretable, with optional rephrasing. Models trained on these refined trajectories consistently outperform those trained on the original trajectories at lower inference cost. Specifically, across four multi-modal QA benchmarks, our refinements improve downstream accuracy by up to $1.7$\,pp over vanilla SFT (and $5.7$\,pp over an LLM-deletion baseline) while reducing per-sample inference messages by up to approximately $40\%$ and inference tokens by up to approximately $48\%$, translating to substantial savings in compute and serving cost. Code is available.
Original Article
View Cached Full Text

Cached at: 09/17/26, 09:18 AM

# Dependency-Aware Trajectory Refinement for Efficient Multi-Turn Agent Fine-Tuning
Source: [https://arxiv.org/html/2609.18417](https://arxiv.org/html/2609.18417)
Zhuo ChenAffiliation:School of Information Science and Technology, ShanghaiTech UniversityAffiliation:Shanghai Engineering Research Center of Intelligent Vision and ImagingEmail:[chenzhuo@shanghaitech\.edu\.cn](mailto:)Xinyu WangKewei Tu††thanks:Corresponding authorAffiliation:School of Information Science and Technology, ShanghaiTech UniversityAffiliation:Shanghai Engineering Research Center of Intelligent Vision and Imaging

###### Abstract

Multi\-turn agent trajectories often contain redundant rounds \(failed tool calls, parallel sub\-queries, verification\-only steps\) that inflate both training and inference cost\. We propose viewing each trajectory as a*round\-level dependency DAG*that exposes which rounds are globally load\-bearing for the final answer, and fine\-tune agents on trajectories refined through this DAG\. Given an LLM\-annotated DAG, these edits are deterministic and interpretable, with optional rephrasing\. Models trained on these refined trajectories consistently outperform those trained on the original trajectories at lower inference cost\. Specifically, across four multi\-modal QA benchmarks, our refinements improve downstream accuracy by up to1\.71\.7pp over vanilla SFT \(and5\.75\.7pp over an LLM\-deletion baseline\) while reducing per\-sample inference messages by up to approximately40%40\\%and inference tokens by up to approximately48%48\\%, translating to substantial savings in compute and serving cost\. Code is available[here](https://github.com/Chord-Chen-30/LLaMA-Factory-lite)\.

## 1Introduction

Tool\-augmented multi\-modal agents that interleave reasoning with retrieval calls now define the state of the art for visual question answering\([Yao et al\., 2023](https://arxiv.org/html/2609.18417#bib.bib2);[Schick et al\., 2023](https://arxiv.org/html/2609.18417#bib.bib3);[Qin et al\., 2023](https://arxiv.org/html/2609.18417#bib.bib4);[Deng et al\., 2023](https://arxiv.org/html/2609.18417#bib.bib16);[Xie et al\., 2024](https://arxiv.org/html/2609.18417#bib.bib18)\)\. Training such agents from a pretrained model typically needs high\-quality multi\-turn trajectories that demonstrate how to use tools to solve complex problems\. These trajectories are valuable but not always information\-dense\. While synthesizing these trajectories, the source model inevitably runs failed searches, double\-checks already confirmed facts, or branches into parallel sub\-queries whose results never enter the final answer\. Cost then compounds at both training and inference stages\.

A natural temptation is to compress trajectories before SFT\. We focus on two contrasting operations\. The first intuitive baseline method isround deletionwithout the contextual awareness of a dependency DAG\. An LLM judge marks middle rounds as redundant towards the final answer\. Through experiments, we find it aggressive but lossy, since removing a round whose facts are still cited leaves the student “hallucinating sources”\. In contrast, our proposed method operates on a dependency graph\. We first construct a DAG over rounds, then perform two levels of structural pruning\. The first level removes non\-terminal leaf nodes, and the second level merges independent sibling rounds\.

Figure 1:Accuracy vs\. inference\-cost trade\-off\.L1reaches the highest accuracy\.L2ahalves the inference tokens at vanilla\-SFT accuracy\.L1: Leaf Prunedrop out\-degree\-0 nodes𝒯\\mathcal\{T\}: ReActtrajectory12345dependency12345𝒯′\\mathcal\{T\}^\{\\prime\}: after1245L2: Strict Mergesame parents*and*children𝒯\\mathcal\{T\}: ReActtrajectory12345dependency12345𝒯′\\mathcal\{T\}^\{\\prime\}: after12\+345L3: Relaxed Mergesame parents only, iterated to convergence𝒯\\mathcal\{T\}: ReActtrajectory123456dependency123456after iter 112\+3456𝒯′\\mathcal\{T\}^\{\\prime\}: iter 2\(converged\)12\+34\+56Figure 2:Three levels of refinements on a structural dependency DAG\. For each level,*trajectory*traces the recorded round order \(\),*dependency*overlays the actual data dependencies on the same nodes \(\), and*after*shows the rewritten trajectory\. Round 1 is the user query, the highest\-indexed round the assistant answer\.Once a trajectory is recast as a round\-level dependency DAG, edits become*deterministic, interpretable, and inexpensive to apply*\. We instantiate three such variants of increasing aggressiveness \(leaf prune, strict merge, relaxed merge\), each preserving the load\-bearing dependencies of the original trajectory, optionally followed by LLM rephrasing\. Our results reveal two complementary sweet spots \(Fig\.[1](https://arxiv.org/html/2609.18417#S1.F1)\)\. Leaf prune \(L1\) anchors the high\-accuracy end at\+5\.7\+5\.7pp over the Critical\-Path baseline\. Strict merge with LLM rephrasing \(L2a\) anchors the cost\-efficient end, matching vanilla\-SFT accuracy while using only11\.011\.0K inference tokens per sample, vs\.20\.920\.9K for vanilla SFT\. Together these two points define the best accuracy–cost trade\-off among the SFT systems we compare\. We further note that the most aggressive variant, relaxed merge \(L3\), is sensitive to the one\-tool\-call\-per\-turn protocol used at inference\. A sizeable fraction of samples exhibit extended interaction loops at noticeably higher inference cost\. Structural edits work best when they preserve the train–test interaction pattern expected at inference\.

Contributions\.\(1\) A dependency\-DAG view of agent trajectories that turns trajectory refinement into a transparent, deterministic graph\-editing problem, with three concrete levels of structural edit\. \(2\) Across four multi\-modal QA benchmarks, our refined trajectories train models that surpass vanilla\-SFT accuracy at substantially lower inference cost\. \(3\) An LLM\-rephrasing extension that closes the training/inference format gap from round merging, together with an analysis linking each edit level’s behaviour to that gap\.

## 2Method

### 2\.1Preliminary

We discuss agent trajectories in the ReAct paradigm\([Yao et al\., 2023](https://arxiv.org/html/2609.18417#bib.bib2)\), where an agent interleaves reasoning and acting to solve tasks\. An agent trajectory𝒯\\mathcal\{T\}is a sequence of*rounds*\. Round 1 is the user’s question\. Rounds2\.\.N−12\.\.N\{\-\}1each consist of an assistant turn \(<think\>\+<tool\_call\>\) followed by a user turn \(<tool\_response\>\)\. RoundNNis the final assistant answer \(<think\>\+<answer\>\)\. We refine𝒯\\mathcal\{T\}into a𝒯′\\mathcal\{T\}^\{\\prime\}that preserves the final answer while containing no more rounds than𝒯\\mathcal\{T\}\.

### 2\.2Round\-Level Dependency DAG

For each trajectory𝒯\\mathcal\{T\}we extract a DAGG⁡\(𝒯\)G\(\\mathcal\{T\}\)whose edgesi→ji\\to jrecord*globally load\-bearing*dependencies\. Roundiiproduced a concrete artefact \(number, named entity, URL, intermediate conclusion\) that Roundjjuses to reach the final answer\. Edges encoding mere narrative reference \(“… was unhelpful, let me try …”\) are excluded\. After parsing, any cycles are removed by depth\-first traversal\. We obtain these edges by querying GPT\-5\.4 with the prompt shown in Sec\.[J](https://arxiv.org/html/2609.18417#A10)\.

### 2\.3Three Levels of Structural Edit

We apply three edits of increasing aggressiveness, illustrated in Fig\.[2](https://arxiv.org/html/2609.18417#S1.F2):

#### Leaf Prune \(L1\)\.

Iteratively remove every out\-degree\-zero round except round 1 \(the user query\) and roundNN\(the assistant answer\)\. These nodes correspond to dead\-end tool calls \(failed retrievals, abandoned sub\-queries\) whose outputs feed nothing downstream\.

#### Strict Merge \(L2\)\.

Two rounds with the same upstream context and the same downstream consumer are interchangeable in dataflow\. From the perspective of the final answer, they are parallel computations of the same logical step\.L2fuses these sibling rounds into one\. This collapses redundant parallelism while preserving every dependency in the DAG\.L2is applied on top ofL1\.

#### Relaxed Merge \(L3\)\.

L3relaxesL2’s criterion by requiring only shared parents, regardless of downstream consumers, and applies the rule iteratively\. Because every merge changes the parent set of nodes downstream of it, new sibling pairs can emerge after a pass, so we keep fusing until no candidates remain\. In Fig\.[2](https://arxiv.org/html/2609.18417#S1.F2)\(panelL3\), merging\{2,3\}\\\{2,3\\\}into2\+32\{\+\}3makes rounds44and55share the new parent2\+32\{\+\}3, and they are fused into4\+54\{\+\}5in a second pass\.L3is applied on top ofL1\.

#### Canonical Format\.

All edits keep one outer tag per turn \(<think\>\+<tool\_call\>/<answer\>on the assistant,<tool\_response\>on the user\)\. Merged bodies are concatenated inside the tag\.

### 2\.4LLM\-Rephrased Merge \(L2a, L3a\)

A merged round’s<think\>is a concatenation of several disjoint trains of thought, a pattern absent in original trajectories\.L2aandL3aapply an LLM rewrite to the merged<think\>ofL2andL3that produces a single coherent passage while preserving every sub\-task transition and tool\-call reference \(prompt in Sec\.[J](https://arxiv.org/html/2609.18417#A10)\)\. The<tool\_call\>and<tool\_response\>bodies are kept verbatim\.

## 3Experiments

### 3\.1Setup

#### Training Data\.

We use 6196 multi\-modal, tool\-using trajectories extended from[Geng et al\. \(2025\)](https://arxiv.org/html/2609.18417#bib.bib15)\. Table[1](https://arxiv.org/html/2609.18417#S3.T1)reports message and token reductions for each edit on a*fair\-comparison subset*, 1334 trajectories that at least one edit modifies\.Critical\-Pathis the most aggressive, andL1is the most conservative\.L2andL3fall in between whereL3removes more rounds but fewer tokens thanL2\.

#### Systems Compared\.

\(i\)Zero\-shot: the baseQwen3\-VL\-30B\-A3B\-Thinkingmodel without SFT\([Bai et al\., 2025](https://arxiv.org/html/2609.18417#bib.bib1)\)\. \(ii\)Vanilla SFT: SFT on the unmodified 6196 raw trajectories\. \(iii\)Critical\-Path: an LLM marks each middle round to be “Keep or Remove”, with a grounding check that vetoes deletions that would orphan named entities in the final answer\. \(iv\)Ours: SFT on the same base withL1–L3atrajectory refinement\. See Secs\.[A](https://arxiv.org/html/2609.18417#A1)and[B](https://arxiv.org/html/2609.18417#A2)for full settings\.

#### Benchmarks\.

We evaluate on four multi\-modal, tool\-using QA datasets\. SimpleVQA \(visual factuality, 300;[Cheng et al\., 2025](https://arxiv.org/html/2609.18417#bib.bib10)\), LiveVQA \(recent\-knowledge visual QA, 300;[Fu et al\., 2025](https://arxiv.org/html/2609.18417#bib.bib11)\), HLE \(hard knowledge problems, 330;[Phan et al\., 2026](https://arxiv.org/html/2609.18417#bib.bib9)\), and MMSearch \(multi\-modal search, 171;[Jiang et al\., 2024](https://arxiv.org/html/2609.18417#bib.bib8)\)\. We reportgpt\-5\-nanojudge accuracy along with two efficiency proxies\.

ConfigAvg\.\#msgsMsg\.red\.Avg\.\#toksTok\.red\.*Original*11\.77—3,692—Critical\-Path8\.2929\.46%2,78124\.65%L19\.6318\.18%3,09616\.15%L29\.2221\.61%3,09116\.28%L38\.8724\.67%3,08716\.39%L2a9\.2221\.61%3,08016\.57%L3a8\.8724\.67%3,06716\.93%Table 1:Training\-set message and token reduction\.

### 3\.2Main Results

Table[2](https://arxiv.org/html/2609.18417#S3.T2)reports task accuracy, average message and token counts, and derived cost\-effectiveness ratios across the four benchmarks\. Overall, Leaf Prune \(L1\) achieves the highest accuracy, while the rephrased strict merge \(L2a\) emerges as the most cost\-effective variant\. Three key observations follow from these results:

\(i\) Breaking limits on both accuracy and efficiency\. Our methods successfully decouple the tight coupling between high accuracy and high cost\.L1establishes a new performance ceiling \(50\.04%50\.04\\%Avg\. Acc\.\), outperforming vanilla SFT by\+1\.7\+1\.7pp\. Simultaneously,L2aredefines inference efficiency, achieving the highest accuracy\-per\-token ratio among all SFT variants while maintaining accuracy comparable to the baseline\. Breaking the 4\-benchmark average down,L2ais uniformly the cheapest in messages, and the cheapest in tokens, on every single benchmark among SFT systems \(Sec\.[C](https://arxiv.org/html/2609.18417#A3)\)\.

Zero\-shotVanillaSFTCriticalPathL1L2L3L2aL3a*Taskaccuracy*SimpleVQA66\.6768\.6760\.6770\.6768\.6766\.6767\.3364\.33LiveVQA48\.0051\.0046\.3353\.6751\.3350\.3351\.0051\.00HLE8\.489\.3911\.2110\.3010\.3011\.5211\.528\.79MMSearch55\.5664\.3359\.0665\.5062\.5763\.7462\.5766\.08Avg\. Acc\.44\.6848\.3544\.3250\.0448\.2248\.0748\.1047\.55*Inferenceefficiency*Avg\. \#msgs41\.6723\.9427\.3315\.9022\.2530\.1914\.3016\.63Avg\. \#toks8,977†20,94019,54014,65715,84131,77910,97713,766*Cost\-effectiveness*Acc\. / msg1\.072\.021\.623\.152\.171\.593\.362\.86Acc\. / tok \(×103\\times 10^\{3\}\)4\.98†2\.312\.273\.413\.041\.514\.383\.45Table 2:Main results across four multi\-modal QA benchmarks\.Boldmarks the best in each row\.†Artefactual: Zero\-shot frequently enters stuck loops on HLE \(see Sec\.[D](https://arxiv.org/html/2609.18417#A4)\), deflating its token average\.\(ii\) Flexible accuracy/cost selection under diverse deployment constraints\. When compute budgets are generous,L1serves as the optimal choice, leading performance on three of the four benchmarks\. Conversely, under strict cost or latency constraints,L2aprovides an ideal drop\-in alternative, slashing token overhead by approximately 48% without sacrificing accuracy\.

\(iii\) Boundary exploration and specialized strengths\. Relaxed merging \(L3\) increases token costs on standard tasks but achieves a peak accuracy of11\.52%11\.52\\%on the challenging HLE dataset, showing promise for complex reasoning\. In contrast, theCritical\-Pathbaseline drops sharply, confirming that bluntly deleting intermediate rounds hurts generalization\.

## 4Analysis

In this section, first we analyse why relaxed merge raises stuck rates and how LLM rephrasing mitigates that gap\. Second, we compare against Chain\-of\-Draft as an inference\-time baseline\. Third, we test whether L1’s kept\-node set is stable under alternative annotators\.

Figure 3:*Stuck rate*\(%\), the fraction of samples that hit the 128\-round agent\-loop ceiling\.### 4\.1Trajectory Merging and Rephrasing

When a merge group contains sibling nodes with divergent children, the training turn aggregates multiple tool responses into one<tool\_response\>block\. This format departs from typical multi\-turn interactions and creates a shift that can disrupt the model’s behavior during inference\. As a result, samples hit the 128\-round agent\-loop ceiling22–3×3\\timesas often underL3as underL2across all four benchmarks \(Fig\.[3](https://arxiv.org/html/2609.18417#S4.F3)\)\. Excluding these stuck samples nearly alignsL3’s median message count withL2’s\.L2avoids this distribution shift by restricting merges to interchangeable siblings\.

L2aandL3arewrite the concatenated<think\>bodies into a single coherent reasoning passage\. OnL3, this rewrite both largely closes the stuck\-rate gap \(Fig\.[3](https://arxiv.org/html/2609.18417#S4.F3)\) and substantially cuts per\-sample cost \(45%45\\%messages,57%57\\%tokens\) at comparable accuracy\. The residual gap betweenL3aandL2ais structural, sinceL3fuses siblings whose downstream consumers differ, a pattern that rephrasing alone cannot reshape\.

Figure 4:Chain\-of\-Draft sweep on SimpleVQA versus Zero\-shot, Vanilla SFT,L1, andL2a\.
### 4\.2Comparison with Chain\-of\-Draft

We also test Chain of Draft\([Xu et al\., 2025](https://arxiv.org/html/2609.18417#bib.bib21)\), a prompting\-only method that compresses single\-response chain of thought, as an inference\-time baseline on the unmodified Zero\-shot model with same agent loop and judge\. Sweeping the per\-step budget from55words to50005000tokens on SimpleVQA \(Fig\.[4](https://arxiv.org/html/2609.18417#S4.F4)\), accuracy falls about2222pp below Zero\-shot while tokens rise to66–8×8\\timesthe Zero\-shot average\. CoD targets short single\-turn reasoning; tightening per\-round<think\>in a multi\-turn tool loop makes the agent fire tools before framing the sub\-problem, so it becomes more wasteful\. We therefore omit CoD from Table[2](https://arxiv.org/html/2609.18417#S3.T2)\. Prompting\-only compression for single\-response tasks does not transfer here, which supports editing trajectories before training rather than constraining generation at inference\.

### 4\.3Annotator Agreement

We re\-annotate training trajectories with two alternative annotators and compare the kept\-node set that L1 consumes against the GPT\-5\.4 reference \(Table[3](https://arxiv.org/html/2609.18417#S4.T3)\)\. Kept\-P is the fraction of an alternative’s kept rounds that GPT\-5\.4 also keeps\. Kept\-R is the fraction of GPT\-5\.4’s kept rounds that the alternative retains\. Both recover the same load\-bearing core and differ mainly in how aggressively they prune\.

AnnotatorKept\-PKept\-RKept\-F1claude\-sonnet\-4\-60\.9160\.8190\.865deepseek\-v4\-pro0\.9900\.5750\.727Table 3:Kept\-node agreement vs\. GPT\-5\.4\.
### 4\.4Additional Experiments

Appendix material covers training and inference settings \(Secs\.[A](https://arxiv.org/html/2609.18417#A1)–[B](https://arxiv.org/html/2609.18417#A2)\), per\-benchmark cost breakdowns \(Sec\.[C](https://arxiv.org/html/2609.18417#A3)\), stuck\-rate notes \(Sec\.[D](https://arxiv.org/html/2609.18417#A4)\), merge\-format alignment \(Sec\.[E](https://arxiv.org/html/2609.18417#A5)\), DPO results and setting \(Secs\.[F](https://arxiv.org/html/2609.18417#A6)–[G](https://arxiv.org/html/2609.18417#A7)\), checkpoint selection \(Sec\.[H](https://arxiv.org/html/2609.18417#A8)\), a qualitative DAG example \(Sec\.[I](https://arxiv.org/html/2609.18417#A9)\), and prompt templates \(Sec\.[J](https://arxiv.org/html/2609.18417#A10)\)\.

## 5Related Work

Tool\-using LLM agents that interleave reasoning with search, visit, or system\-level actions are now a mainstream recipe\([Yao et al\., 2023](https://arxiv.org/html/2609.18417#bib.bib2);[Schick et al\., 2023](https://arxiv.org/html/2609.18417#bib.bib3);[Qin et al\., 2023](https://arxiv.org/html/2609.18417#bib.bib4);[Deng et al\., 2023](https://arxiv.org/html/2609.18417#bib.bib16);[Xie et al\., 2024](https://arxiv.org/html/2609.18417#bib.bib18);[Yao et al\., 2024](https://arxiv.org/html/2609.18417#bib.bib19)\), and SFT on multi\-turn trajectories from a stronger LLM is the prevailing fine\-tuning route\([Zeng et al\., 2024](https://arxiv.org/html/2609.18417#bib.bib17)\)\.

Closer in spirit, three lines of work pursue better SFT*data*\. The “less\-is\-more” line picks SFT samples by hand, quality score, or influence estimate, and shows that small selected corpora can match larger noisy ones\([Zhou et al\., 2023](https://arxiv.org/html/2609.18417#bib.bib6);[Cao et al\., 2024](https://arxiv.org/html/2609.18417#bib.bib12);[Liu et al\., 2024](https://arxiv.org/html/2609.18417#bib.bib13)\)\. Reasoning\-step supervision edits or scores individual steps inside one response, either by self\-training new rationales\([Zelikman et al\., 2022](https://arxiv.org/html/2609.18417#bib.bib7)\)or via step\-level process rewards\([Wang et al\., 2024](https://arxiv.org/html/2609.18417#bib.bib14)\)\. Trajectory synthesis generates agentic training data from scratch via teacher\-driven agentic flows\([Mitra et al\., 2024](https://arxiv.org/html/2609.18417#bib.bib20)\)\.

These approaches operate across whole samples, inside a single response, or by generating new data, leaving the internal structure of an agent trajectory untouched\. We recast each trajectory as a round\-level dependency DAG and apply deterministic edits inside it, turning trajectory refinement into a transparent graph\-editing problem\.

## 6Conclusion

We recast each multi\-turn agent trajectory as a round\-level dependency DAG and apply three deterministic structural edits, with an optional LLM rephrasing variant\. Across four multi\-modal QA benchmarks, leaf prune \(L1\) anchors the high\-accuracy end at\+1\.7\+1\.7pp over vanilla SFT, while strict merge with rephrasing \(L2a\) anchors the cost\-efficient end by halving per\-sample inference tokens at vanilla\-SFT accuracy\. The graph\-editing view makes the structural assumptions behind each edit explicit and offers a general handle for refining multi\-turn agent trajectories\.

## Limitations

Our study is scoped to one base model family \(Qwen3\-VL\-30B\-A3B\-Thinking\)\. The proposed edits act on training data rather than model weights, but the combined training and per\-checkpoint evaluation already approaches our compute budget, so transfer across base models and modalities is left to follow\-up work\. Dependency annotation and merge\-think rephrasing each add one annotator pass at data preparation time, and these one\-time costs amortise quickly against the inference\-time savings in Sec\.[3\.2](https://arxiv.org/html/2609.18417#S3.SS2)\.

## Acknowledgments

This work was supported by the Core Facility Platform of Computer Science and Communication, SIST, ShanghaiTech University\.

## References

- Baiet al\.\(2025\)S\. Bai, Y\. Cai, R\. Chen, K\. Chen, X\. Chen, Z\. Cheng, L\. Deng, W\. Ding, C\. Gao, C\. Ge, W\. Ge, Z\. Guo, Q\. Huang, J\. Huang, F\. Huang, B\. Hui, S\. Jiang, Z\. Li, M\. Li, M\. Li, K\. Li, Z\. Lin, J\. Lin, X\. Liu, J\. Liu, C\. Liu, Y\. Liu, D\. Liu, S\. Liu, D\. Lu, R\. Luo, C\. Lv, R\. Men, L\. Meng, X\. Ren, X\. Ren, S\. Song, Y\. Sun, J\. Tang, J\. Tu, J\. Wan, P\. Wang, P\. Wang, Q\. Wang, Y\. Wang, T\. Xie, Y\. Xu, H\. Xu, J\. Xu, Z\. Yang, M\. Yang, J\. Yang, A\. Yang, B\. Yu, F\. Zhang, H\. Zhang, X\. Zhang, B\. Zheng, H\. Zhong, J\. Zhou, F\. Zhou, J\. Zhou, Y\. Zhu, and K\. ZhuQwen3\-vl technical report\.External Links:2511\.21631,[Link](https://arxiv.org/abs/2511.21631)Cited by:[Appendix A](https://arxiv.org/html/2609.18417#A1.p1.1),[§3\.1](https://arxiv.org/html/2609.18417#S3.SS1.SSS0.Px2.p1.1)\.
- Caoet al\.\(2024\)Y\. Cao, Y\. Kang, C\. Wang, and L\. SunInstruction mining: instruction data selection for tuning large language models\.External Links:2307\.06290,[Link](https://arxiv.org/abs/2307.06290)Cited by:[§5](https://arxiv.org/html/2609.18417#S5.p2.1)\.
- Chenget al\.\(2025\)X\. Cheng, W\. Zhang, S\. Zhang, J\. Yang, X\. Guan, X\. Wu, X\. Li, G\. Zhang, J\. Liu, Y\. Mai, Y\. Zeng, Z\. Wen, K\. Jin, B\. Wang, W\. Zhou, Y\. Lu, T\. Li, W\. Huang, and Z\. LiSimpleVQA: multimodal factuality evaluation for multimodal large language models\.External Links:2502\.13059,[Link](https://arxiv.org/abs/2502.13059)Cited by:[§3\.1](https://arxiv.org/html/2609.18417#S3.SS1.SSS0.Px3.p1.1)\.
- Denget al\.\(2023\)X\. Deng, Y\. Gu, B\. Zheng, S\. Chen, S\. Stevens, B\. Wang, H\. Sun, and Y\. SuMind2Web: towards a generalist agent for the web\.External Links:2306\.06070Cited by:[§1](https://arxiv.org/html/2609.18417#S1.p1.1),[§5](https://arxiv.org/html/2609.18417#S5.p1.1)\.
- Fuet al\.\(2025\)M\. Fu, Y\. Peng, D\. Chen, Z\. Zhou, B\. Liu, Y\. Wan, Z\. Zhao, P\. S\. Yu, and R\. KrishnaSeeking and updating with live visual knowledge\.External Links:2504\.05288,[Link](https://arxiv.org/abs/2504.05288)Cited by:[§3\.1](https://arxiv.org/html/2609.18417#S3.SS1.SSS0.Px3.p1.1)\.
- Genget al\.\(2025\)X\. Geng, P\. Xia, Z\. Zhang, X\. Wang, Q\. Wang, R\. Ding, C\. Wang, J\. Wu, Y\. Zhao, K\. Li, Y\. Jiang, P\. Xie, F\. Huang, and J\. ZhouWebWatcher: breaking new frontier of vision\-language deep research agent\.External Links:2508\.05748,[Link](https://arxiv.org/abs/2508.05748)Cited by:[§3\.1](https://arxiv.org/html/2609.18417#S3.SS1.SSS0.Px1.p1.1)\.
- Jianget al\.\(2024\)D\. Jiang, R\. Zhang, Z\. Guo, Y\. Wu, J\. Lei, P\. Qiu, P\. Lu, Z\. Chen, G\. Song, P\. Gao,et al\.MMSearch: benchmarking the potential of large models as multi\-modal search engines\.arXiv preprint arXiv:2409\.12959\.Cited by:[§3\.1](https://arxiv.org/html/2609.18417#S3.SS1.SSS0.Px3.p1.1)\.
- Liuet al\.\(2024\)W\. Liu, W\. Zeng, K\. He, Y\. Jiang, and J\. HeWhat makes good data for alignment? a comprehensive study of automatic data selection in instruction tuning\.External Links:2312\.15685,[Link](https://arxiv.org/abs/2312.15685)Cited by:[§5](https://arxiv.org/html/2609.18417#S5.p2.1)\.
- Mitraet al\.\(2024\)A\. Mitra, L\. D\. Corro, G\. Zheng, S\. Mahajan, D\. Rouhana, A\. Codas, Y\. Lu, W\. Chen, O\. Vrousgos, C\. Rosset, F\. Silva, H\. Khanpour, Y\. Lara, and A\. AwadallahAgentInstruct: toward generative teaching with agentic flows\.External Links:2407\.03502,[Link](https://arxiv.org/abs/2407.03502)Cited by:[§5](https://arxiv.org/html/2609.18417#S5.p2.1)\.
- Phanet al\.\(2026\)L\. Phan, A\. Gatti, Z\. Han, N\. Li, J\. Hu, H\. Zhang, C\. B\. C\. Zhang, M\. Shaaban, J\. Ling, S\. Shi, M\. Choi, A\. Agrawal, A\. Chopra, A\. Khoja, R\. Kim, R\. Ren, J\. Hausenloy, O\. Zhang, M\. Mazeika, D\. Dodonov, T\. Nguyen, J\. Lee, D\. Anderson, M\. Doroshenko, A\. C\. Stokes, M\. Mahmood, O\. Pokutnyi, O\. Iskra, J\. P\. Wang, J\. Levin, M\. Kazakov, F\. Feng, S\. Y\. Feng, H\. Zhao, M\. Yu, V\. Gangal, C\. Zou, Z\. Wang, S\. Popov, R\. Gerbicz, G\. Galgon, J\. Schmitt, W\. Yeadon, Y\. Lee, S\. Sauers, A\. Sanchez, F\. Giska, M\. Roth, S\. Riis, S\. Utpala, N\. Burns, G\. M\. Goshu, M\. M\. Naiya, C\. Agu, Z\. Giboney, A\. Cheatom, F\. Fournier\-Facio, S\. Crowson, L\. Finke, Z\. Cheng, J\. Zampese, R\. G\. Hoerr, M\. Nandor, H\. Park, T\. Gehrunger, J\. Cai, B\. McCarty, A\. C\. Garretson, E\. Taylor, D\. Sileo, Q\. Ren, U\. Qazi, L\. Li, J\. Nam, J\. B\. Wydallis, P\. Arkhipov, J\. W\. L\. Shi, A\. Bacho, C\. G\. Willcocks, H\. Cao, S\. Motwani, E\. de Oliveira Santos, J\. Veith, E\. Vendrow, D\. Cojoc, K\. Zenitani, J\. Robinson, L\. Tang, Y\. Li, J\. Vendrow, N\. W\. Fraga, V\. Kuchkin, A\. P\. Maksimov, P\. Marion, D\. Efremov, J\. Lynch, K\. Liang, A\. Mikov, A\. Gritsevskiy, J\. Guillod, G\. Demir, D\. Martinez, B\. Pageler, K\. Zhou, S\. Soori, O\. Press, H\. Tang, P\. Rissone, S\. R\. Green, L\. Brüssel, M\. Twayana, A\. Dieuleveut, J\. M\. Imperial, A\. Prabhu, J\. Yang, N\. Crispino, A\. Rao, D\. Zvonkine, G\. Loiseau, M\. Kalinin, M\. Lukas, C\. Manolescu, N\. Stambaugh, S\. Mishra, T\. Hogg, C\. Bosio, B\. P\. Coppola, J\. Salazar, J\. Jin, R\. Sayous, S\. Ivanov, P\. Schwaller, S\. Senthilkuma, A\. M\. Bran, A\. Algaba, K\. V\. den Houte, L\. V\. D\. Sypt, B\. Verbeken, D\. Noever, A\. Kopylov, B\. Myklebust, B\. Li, L\. Schut, E\. Zheltonozhskii, Q\. Yuan, D\. Lim, R\. Stanley, T\. Yang, J\. Maar, J\. Wykowski, M\. Oller, A\. Sahu, C\. G\. Ardito, Y\. Hu, A\. G\. K\. Kamdoum, A\. Jin, T\. G\. Vilchis, Y\. Zu, M\. Lackner, J\. Koppel, G\. Sun, D\. S\. Antonenko, S\. Chern, B\. Zhao, P\. Arsene, J\. M\. Cavanagh, D\. Li, J\. Shen, D\. Crisostomi, W\. Zhang, A\. Dehghan, S\. Ivanov, D\. Perrella, N\. Kaparov, A\. Zang, I\. Sucholutsky, A\. Kharlamova, D\. Orel, V\. Poritski, S\. Ben\-David, Z\. Berger, P\. Whitfill, M\. Foster, D\. Munro, L\. Ho, S\. Sivarajan, D\. B\. Hava, A\. Kuchkin, D\. Holmes, A\. Rodriguez\-Romero, F\. Sommerhage, A\. Zhang, R\. Moat, K\. Schneider, Z\. Kazibwe, D\. Clarke, D\. H\. Kim, F\. M\. Dias, S\. Fish, V\. Elser, T\. Kreiman, V\. E\. G\. Vilchis, I\. Klose, U\. Anantheswaran, A\. Zweiger, K\. Rawal, J\. Li, J\. Nguyen, N\. Daans, H\. Heidinger, M\. Radionov, V\. Rozhoň, V\. Ginis, C\. Stump, N\. Cohen, R\. Poświata, J\. Tkadlec, A\. Goldfarb, C\. Wang, P\. Padlewski, S\. Barzowski, K\. Montgomery, R\. Stendall, J\. Tucker\-Foltz, J\. Stade, T\. R\. Rogers, T\. Goertzen, D\. Grabb, A\. Shukla, A\. Givré, J\. A\. Ambay, A\. Sen, M\. F\. Aziz, M\. H\. Inlow, H\. He, L\. Zhang, Y\. Kaddar, I\. Ängquist, Y\. Chen, H\. K\. Wang, K\. Ramakrishnan, E\. Thornley, A\. Terpin, H\. Schoelkopf, E\. Zheng, A\. Carmi, E\. D\. L\. Brown, K\. Zhu, M\. Bartolo, R\. Wheeler, M\. Stehberger, P\. Bradshaw, J\. Heimonen, K\. Sridhar, I\. Akov, J\. Sandlin, Y\. Makarychev, J\. Tam, H\. Hoang, D\. M\. Cunningham, V\. Goryachev, D\. Patramanis, M\. Krause, A\. Redenti, D\. Aldous, J\. Lai, S\. Coleman, J\. Xu, S\. Lee, I\. Magoulas, S\. Zhao, N\. Tang, M\. K\. Cohen, O\. Paradise, J\. H\. Kirchner, M\. Ovchynnikov, J\. O\. Matos, A\. Shenoy, M\. Wang, Y\. Nie, A\. Sztyber\-Betley, P\. Faraboschi, R\. Riblet, J\. Crozier, S\. Halasyamani, S\. Verma, P\. Joshi, E\. Meril, Z\. Ma, J\. Andréoletti, R\. Singhal, J\. Platnick, V\. Nevirkovets, L\. Basler, A\. Ivanov, S\. Khoury, N\. Gustafsson, M\. Piccardo, H\. Mostaghimi, Q\. Chen, V\. Singh, T\. Q\. Khánh, P\. Rosu, H\. Szlyk, Z\. Brown, H\. Narayan, A\. Menezes, J\. Roberts, W\. Alley, K\. Sun, A\. Patel, M\. Lamparth, A\. Reuel, L\. Xin, H\. Xu, J\. Loader, F\. Martin, Z\. Wang, A\. Achilleos, T\. Preu, T\. Korbak, I\. Bosio, F\. Kazemi, Z\. Chen, B\. Bálint, E\. J\. Y\. Lo, J\. Wang, M\. I\. S\. Nunes, J\. Milbauer, M\. S\. Bari, Z\. Wang, B\. Ansarinejad, Y\. Sun, S\. Durand, H\. Elgnainy, G\. Douville, D\. Tordera, G\. Balabanian, H\. Wolff, L\. Kvistad, H\. Milliron, A\. Sakor, M\. Eron, A\. F\. D\. O\., S\. Shah, X\. Zhou, F\. Kamalov, S\. Abdoli, T\. Santens, S\. Barkan, A\. Tee, R\. Zhang, A\. Tomasiello, G\. B\. D\. Luca, S\. Looi, V\. Le, N\. Kolt, J\. Pan, E\. Rodman, J\. Drori, C\. J\. Fossum, N\. Muennighoff, M\. Jagota, R\. Pradeep, H\. Fan, J\. Eicher, M\. Chen, K\. Thaman, W\. Merrill, M\. Firsching, C\. Harris, S\. Ciobâcă, J\. Gross, R\. Pandey, I\. Gusev, A\. Jones, S\. Agnihotri, P\. Zhelnov, M\. Mofayezi, A\. Piperski, D\. K\. Zhang, K\. Dobarskyi, R\. Leventov, I\. Soroko, J\. Duersch, V\. Taamazyan, A\. Ho, W\. Ma, W\. Held, R\. Xian, A\. R\. Zebaze, M\. Mohamed, J\. N\. Leser, M\. X\. Yuan, L\. Yacar, J\. Lengler, K\. Olszewska, C\. D\. Fratta, E\. Oliveira, J\. W\. Jackson, A\. Zou, M\. Chidambaram, T\. Manik, H\. Haffenden, D\. Stander, A\. Dasouqi, A\. Shen, B\. Golshani, D\. Stap, E\. Kretov, M\. Uzhou, A\. B\. Zhidkovskaya, N\. Winter, M\. O\. Rodriguez, R\. Lauff, D\. Wehr, C\. Tang, Z\. Hossain, S\. Phillips, F\. Samuele, F\. Ekström, A\. Hammon, O\. Patel, F\. Farhidi, G\. Medley, F\. Mohammadzadeh, M\. Peñaflor, H\. Kassahun, A\. Friedrich, R\. H\. Perez, D\. Pyda, T\. Sakal, O\. Dhamane, A\. K\. Mirabadi, E\. Hallman, K\. Okutsu, M\. Battaglia, M\. Maghsoudimehrabani, A\. Amit, D\. Hulbert, R\. Pereira, S\. Weber, Handoko, A\. Peristyy, S\. Malina, M\. Mehkary, R\. Aly, F\. Reidegeld, A\. Dick, C\. Friday, M\. Singh, H\. Shapourian, W\. Kim, M\. Costa, H\. Gurdogan, H\. Kumar, C\. Ceconello, C\. Zhuang, H\. Park, M\. Carroll, A\. R\. Tawfeek, S\. Steinerberger, D\. Aggarwal, M\. Kirchhof, L\. Dai, E\. Kim, J\. Ferret, J\. Shah, Y\. Wang, M\. Yan, K\. Burdzy, L\. Zhang, A\. Franca, D\. T\. Pham, K\. Y\. Loh, J\. Robinson, A\. Jackson, P\. Giordano, P\. Petersen, A\. Cosma, J\. Colino, C\. White, J\. Votava, V\. Vinnikov, E\. Delaney, P\. Spelda, V\. Stritecky, S\. M\. Shahid, J\. Mourrat, L\. Vetoshkin, K\. Sponselee, R\. Bacho, Z\. Yong, F\. de la Rosa, N\. Cho, X\. Li, G\. Malod, O\. Weller, G\. Albani, L\. Lang, J\. Laurendeau, D\. Kazakov, F\. Adesanya, J\. Portier, L\. Hollom, V\. Souza, Y\. A\. Zhou, J\. Degorre, Y\. Yalın, G\. D\. Obikoya, Rai, F\. Bigi, M\. C\. Boscá, O\. Shumar, K\. Bacho, G\. Recchia, M\. Popescu, N\. Shulga, N\. M\. Tanwie, T\. C\. H\. Lux, B\. Rank, C\. Ni, M\. Brooks, A\. Yakimchyk, Huanxu, Liu, S\. Cavalleri, O\. Häggström, E\. Verkama, J\. Newbould, H\. Gundlach, L\. Brito\-Santana, B\. Amaro, V\. Vajipey, R\. Grover, T\. Wang, Y\. Kratish, W\. Li, S\. Gopi, A\. Caciolai, C\. S\. de Witt, P\. Hernández\-Cámara, E\. Rodolà, J\. Robins, D\. Williamson, V\. Cheng, B\. Raynor, H\. Qi, B\. Segev, J\. Fan, S\. Martinson, E\. Y\. Wang, K\. Hausknecht, M\. P\. Brenner, M\. Mao, C\. Demian, P\. Kassani, X\. Zhang, D\. Avagian, E\. J\. Scipio, A\. Ragoler, J\. Tan, B\. Sims, R\. Plecnik, A\. Kirtland, O\. F\. Bodur, D\. P\. Shinde, Y\. C\. L\. Labrador, Z\. Adoul, M\. Zekry, A\. Karakoc, T\. C\. B\. Santos, S\. Shamseldeen, L\. Karim, A\. Liakhovitskaia, N\. Resman, N\. Farina, J\. C\. Gonzalez, G\. Maayan, E\. Anderson, R\. D\. O\. Pena, E\. Kelley, H\. Mariji, R\. Pouriamanesh, W\. Wu, R\. Finocchio, I\. Alarab, J\. Cole, D\. Ferreira, B\. Johnson, M\. Safdari, L\. Dai, S\. Arthornthurasuk, I\. C\. McAlister, A\. J\. Moyano, A\. Pronin, J\. Fan, A\. Ramirez\-Trinidad, Y\. Malysheva, D\. Pottmaier, O\. Taheri, S\. Stepanic, S\. Perry, L\. Askew, R\. A\. H\. Rodríguez, A\. M\. R\. Minissi, R\. Lorena, K\. Iyer, A\. A\. Fasiludeen, R\. Clark, J\. Ducey, M\. Piza, M\. Somrak, E\. Vergo, J\. Qin, B\. Borbás, E\. Chu, J\. Lindsey, A\. Jallon, I\. M\. J\. McInnis, E\. Chen, A\. Semler, L\. Gloor, T\. Shah, M\. Carauleanu, P\. Lauer, T\. Đ\. Huy, H\. Shahrtash, E\. Duc, L\. Lewark, A\. Brown, S\. Albanie, B\. Weber, W\. S\. Vaz, P\. Clavier, Y\. Fan, G\. P\. R\. e Silva, Long, Lian, M\. Abramovitch, X\. Jiang, S\. Mendoza, M\. Islam, J\. Gonzalez, V\. Mavroudis, J\. Xu, P\. Kumar, L\. P\. Goswami, D\. Bugas, N\. Heydari, F\. Jeanplong, T\. Jansen, A\. Pinto, A\. Apronti, A\. Galal, N\. Ze\-An, A\. Singh, T\. Jiang, J\. of Arc Xavier, K\. P\. Agarwal, M\. Berkani, G\. Zhang, Z\. Du, B\. A\. de Oliveira Junior, D\. Malishev, N\. Remy, T\. D\. Hartman, T\. Tarver, S\. Mensah, G\. A\. Loume, W\. Morak, F\. Habibi, S\. Hoback, W\. Cai, J\. Gimenez, R\. G\. Montecillo, J\. Łucki, R\. Campbell, A\. Sharma, K\. Meer, S\. Gul, D\. E\. Gonzalez, X\. Alapont, A\. Hoover, G\. Chhablani, F\. Vargus, A\. Agarwal, Y\. Jiang, D\. Patil, D\. Outevsky, K\. J\. Scaria, R\. Maheshwari, A\. Dendane, P\. Shukla, A\. Cartwright, S\. Bogdanov, N\. Mündler, S\. Möller, L\. Arnaboldi, K\. Thaman, M\. R\. Siddiqi, P\. Saxena, H\. Gupta, T\. Fruhauff, G\. Sherman, M\. Vincze, S\. Usawasutsakorn, D\. Ler, A\. Radhakrishnan, I\. Enyekwe, S\. M\. Salauddin, J\. Muzhen, A\. Maksapetyan, V\. Rossbach, C\. Harjadi, M\. Bahaloohoreh, C\. Sparrow, J\. Sidhu, S\. Ali, S\. Bian, J\. Lai, E\. Singer, J\. L\. Uro, G\. Bateman, M\. Sayed, A\. Menshawy, D\. Duclosel, D\. Bezzi, Y\. Jain, A\. Aaron, M\. Tiryakioglu, S\. Siddh, K\. Krenek, I\. A\. Shah, J\. Jin, S\. Creighton, D\. Peskoff, Z\. EL\-Wasif, R\. P\. V, M\. Richmond, J\. McGowan, T\. Patwardhan, H\. Sun, T\. Sun, N\. Zubić, S\. Sala, S\. Ebert, J\. Kaddour, M\. Schottdorf, D\. Wang, G\. Petruzella, A\. Meiburg, T\. Medved, A\. ElSheikh, S\. A\. Hebbar, L\. Vaquero, X\. Yang, J\. Poulos, V\. Zouhar, S\. Bogdanik, M\. Zhang, J\. Sanz\-Ros, D\. Anugraha, Y\. Dai, A\. N\. Nhu, X\. Wang, A\. A\. Demircali, Z\. Jia, Y\. Zhou, J\. Wu, M\. He, N\. Chandok, A\. Sinha, G\. Luo, L\. Le, M\. Noyé, M\. Perełkiewicz, I\. Pantidis, T\. Qi, S\. S\. Purohit, L\. Parcalabescu, T\. Nguyen, G\. I\. Winata, E\. M\. Ponti, H\. Li, K\. Dhole, J\. Park, D\. Abbondanza, Y\. Wang, A\. Nayak, D\. M\. Caetano, A\. A\. W\. L\. Wong, M\. del Rio\-Chanona, D\. Kondor, P\. Francois, E\. Chalstrey, J\. Zsambok, D\. Hoyer, J\. Reddish, J\. Hauser, F\. Rodrigo\-Ginés, S\. Datta, M\. Shepherd, T\. Kamphuis, Q\. Zhang, H\. Kim, R\. Sun, J\. Yao, F\. Dernoncourt, S\. Krishna, S\. Rismanchian, B\. Pu, F\. Pinto, Y\. Wang, K\. Shridhar, K\. J\. Overholt, G\. Briia, H\. Nguyen, David, S\. Bartomeu, T\. C\. Pang, A\. Wecker, Y\. Xiong, F\. Li, L\. S\. Huber, J\. Jaeger, R\. D\. Maddalena, X\. H\. Lù, Y\. Zhang, C\. Beger, P\. T\. J\. Kon, S\. Li, V\. Sanker, M\. Yin, Y\. Liang, X\. Zhang, A\. Agrawal, L\. S\. Yifei, Z\. Zhang, M\. Cai, Y\. Sonmez, C\. Cozianu, C\. Li, A\. Slen, S\. Yu, H\. K\. Park, G\. Sarti, M\. Briański, A\. Stolfo, T\. A\. Nguyen, M\. Zhang, Y\. Perlitz, J\. Hernandez\-Orallo, R\. Li, A\. Shabani, F\. Juefei\-Xu, S\. Dhingra, O\. Zohar, M\. C\. Nguyen, A\. Pondaven, A\. Yilmaz, X\. Zhao, C\. Jin, M\. Jiang, S\. Todoran, X\. Han, J\. Kreuer, B\. Rabern, A\. Plassart, M\. Maggetti, L\. Yap, R\. Geirhos, J\. Kean, D\. Wang, S\. Mollaei, C\. Sun, Y\. Yin, S\. Wang, R\. Li, Y\. Chang, A\. Wei, A\. Bizeul, X\. Wang, A\. O\. Arrais, K\. Mukherjee, J\. Chamorro\-Padial, J\. Liu, X\. Qu, J\. Guan, A\. Bouyamourn, S\. Wu, M\. Plomecka, J\. Chen, M\. Tang, J\. Deng, S\. Subramanian, H\. Xi, H\. Chen, W\. Zhang, Y\. Ren, H\. Tu, S\. Kim, Y\. Chen, S\. V\. Marjanović, J\. Ha, G\. Luczyna, J\. J\. Ma, Z\. Shen, D\. Song, C\. E\. Zhang, Z\. Wang, G\. Gendron, Y\. Xiao, L\. Smucker, E\. Weng, K\. H\. Lee, Z\. Ye, S\. Ermon, I\. D\. Lopez\-Miguel, T\. Knights, A\. Gitter, N\. Park, B\. Wei, H\. Chen, K\. Pai, A\. Elkhanany, H\. Lin, P\. D\. Siedler, J\. Fang, R\. Mishra, K\. Zsolnai\-Fehér, X\. Jiang, S\. Khan, J\. Yuan, R\. K\. Jain, X\. Lin, M\. Peterson, Z\. Wang, A\. Malusare, M\. Tang, I\. Gupta, I\. Fosin, T\. Kang, B\. Dworakowska, K\. Matsumoto, G\. Zheng, G\. Sewuster, J\. P\. Villanueva, I\. Rannev, I\. Chernyavsky, J\. Chen, D\. Banik, B\. Racz, W\. Dong, J\. Wang, L\. Bashmal, D\. V\. Gonçalves, W\. Hu, K\. Bar, O\. Bohdal, A\. S\. Patlan, S\. Dhuliawala, C\. Geirhos, J\. Wist, Y\. Kansal, B\. Chen, K\. Tire, A\. T\. Yücel, B\. Christof, V\. Singla, Z\. Song, S\. Chen, J\. Ge, K\. Ponkshe, I\. Park, T\. Shi, M\. Q\. Ma, J\. Mak, S\. Lai, A\. Moulin, Z\. Cheng, Z\. Zhu, Z\. Zhang, V\. Patil, K\. Jha, Q\. Men, J\. Wu, T\. Zhang, B\. H\. Vieira, A\. F\. Aji, J\. Chung, M\. Mahfoud, H\. T\. Hoang, M\. Sperzel, W\. Hao, K\. Meding, S\. Xu, V\. Kostakos, D\. Manini, Y\. Liu, C\. Toukmaji, J\. Paek, E\. Yu, A\. E\. Demircali, Z\. Sun, I\. Dewerpe, H\. Qin, R\. Pflugfelder, J\. Bailey, J\. Morris, V\. Heilala, S\. Rosset, Z\. Yu, P\. E\. Chen, W\. Yeo, E\. Jain, R\. Yang, S\. Chigurupati, J\. Chernyavsky, S\. P\. Reddy, S\. Venugopalan, H\. Batra, C\. F\. Park, H\. Tran, G\. Maximiano, G\. Zhang, Y\. Liang, H\. Shiyu, R\. Xu, R\. Pan, S\. Suresh, Z\. Liu, S\. Gulati, S\. Zhang, P\. Turchin, C\. W\. Bartlett, C\. R\. Scotese, P\. M\. Cao, B\. Wu, J\. Karwowski, D\. Scaramuzza, A\. Nattanmai, G\. McKellips, A\. Cheraku, A\. Suhail, E\. Luo, M\. Deng, J\. Luo, A\. Zhang, K\. Jindel, J\. Paek, K\. Halevy, A\. Baranov, M\. Liu, A\. Avadhanam, D\. Zhang, V\. Cheng, B\. Ma, E\. Fu, L\. Do, J\. Lass, H\. Yang, S\. Sunkari, V\. Bharath, V\. Ai, J\. Leung, R\. Agrawal, A\. Zhou, K\. Chen, T\. Kalpathi, Z\. Xu, G\. Wang, T\. Xiao, E\. Maung, S\. Lee, R\. Yang, R\. Yue, B\. Zhao, J\. Yoon, S\. Sun, A\. Singh, E\. Luo, C\. Peng, T\. Osbey, T\. Wang, D\. Echeazu, H\. Yang, T\. Wu, S\. Patel, V\. Kulkarni, V\. Sundarapandiyan, A\. Zhang, A\. Le, Z\. Nasim, S\. Yalam, R\. Kasamsetty, S\. Samal, H\. Yang, D\. Sun, N\. Shah, A\. Saha, A\. Zhang, L\. Nguyen, L\. Nagumalli, K\. Wang, A\. Zhou, A\. Wu, J\. Luo, A\. Telluri, S\. Dillmann, Z\. Wang, J\. Luo, H\. Lunn, A\. Gazizov, H\. Qiu, A\. G\. Hart, R\. B\. Gabrielsson, I\. Akov, A\. Lukoianov, S\. Yue, A\. Wang, and D\. HendrycksHumanity’s last exam\.External Links:2501\.14249,[Document](https://dx.doi.org/https%3A//doi.org/10.1038/s41586-025-09962-4),[Link](https://arxiv.org/abs/2501.14249)Cited by:[§3\.1](https://arxiv.org/html/2609.18417#S3.SS1.SSS0.Px3.p1.1)\.
- Qinet al\.\(2023\)Y\. Qin, S\. Liang, Y\. Ye, K\. Zhu, L\. Yan, Y\. Lu, Y\. Lin, X\. Cong, X\. Tang, B\. Qian, S\. Zhao, L\. Hong, R\. Tian, R\. Xie, J\. Zhou, M\. Gerstein, D\. Li, Z\. Liu, and M\. SunToolLLM: facilitating large language models to master 16000\+ real\-world apis\.External Links:2307\.16789,[Link](https://arxiv.org/abs/2307.16789)Cited by:[§1](https://arxiv.org/html/2609.18417#S1.p1.1),[§5](https://arxiv.org/html/2609.18417#S5.p1.1)\.
- Rafailovet al\.\(2024\)R\. Rafailov, A\. Sharma, E\. Mitchell, S\. Ermon, C\. D\. Manning, and C\. FinnDirect preference optimization: your language model is secretly a reward model\.External Links:2305\.18290,[Link](https://arxiv.org/abs/2305.18290)Cited by:[Appendix F](https://arxiv.org/html/2609.18417#A6.p1.1)\.
- Schicket al\.\(2023\)T\. Schick, J\. Dwivedi\-Yu, R\. Dessì, R\. Raileanu, M\. Lomeli, L\. Zettlemoyer, N\. Cancedda, and T\. ScialomToolformer: language models can teach themselves to use tools\.External Links:2302\.04761,[Link](https://arxiv.org/abs/2302.04761)Cited by:[§1](https://arxiv.org/html/2609.18417#S1.p1.1),[§5](https://arxiv.org/html/2609.18417#S5.p1.1)\.
- Wanget al\.\(2024\)P\. Wang, L\. Li, Z\. Shao, R\. Xu, D\. Dai, Y\. Li, D\. Chen, Y\. Wu, and Z\. SuiMath\-shepherd: verify and reinforce LLMs step\-by\-step without human annotations\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 9426–9439\.External Links:[Link](https://aclanthology.org/2024.acl-long.510/),[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.510)Cited by:[§5](https://arxiv.org/html/2609.18417#S5.p2.1)\.
- Xieet al\.\(2024\)T\. Xie, D\. Zhang, J\. Chen, X\. Li, S\. Zhao, R\. Cao, T\. J\. Hua, Z\. Cheng, D\. Shin, F\. Lei, Y\. Liu, Y\. Xu, S\. Zhou, S\. Savarese, C\. Xiong, V\. Zhong, and T\. YuOSWorld: benchmarking multimodal agents for open\-ended tasks in real computer environments\.External Links:2404\.07972Cited by:[§1](https://arxiv.org/html/2609.18417#S1.p1.1),[§5](https://arxiv.org/html/2609.18417#S5.p1.1)\.
- Xuet al\.\(2025\)S\. Xu, W\. Xie, L\. Zhao, and P\. HeChain of draft: thinking faster by writing less\.External Links:2502\.18600,[Link](https://arxiv.org/abs/2502.18600)Cited by:[§4\.2](https://arxiv.org/html/2609.18417#S4.SS2.p1.1)\.
- Yaoet al\.\(2024\)S\. Yao, N\. Shinn, P\. Razavi, and K\. Narasimhanτ\\tau\-Bench: a benchmark for tool\-agent\-user interaction in real\-world domains\.External Links:2406\.12045,[Link](https://arxiv.org/abs/2406.12045)Cited by:[§5](https://arxiv.org/html/2609.18417#S5.p1.1)\.
- Yaoet al\.\(2023\)S\. Yao, J\. Zhao, D\. Yu, N\. Du, I\. Shafran, K\. Narasimhan, and Y\. CaoReAct: synergizing reasoning and acting in language models\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§1](https://arxiv.org/html/2609.18417#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.18417#S2.SS1.p1.1),[§5](https://arxiv.org/html/2609.18417#S5.p1.1)\.
- Zelikmanet al\.\(2022\)E\. Zelikman, Y\. Wu, J\. Mu, and N\. D\. GoodmanSTaR: bootstrapping reasoning with reasoning\.External Links:2203\.14465,[Link](https://arxiv.org/abs/2203.14465)Cited by:[§5](https://arxiv.org/html/2609.18417#S5.p2.1)\.
- Zenget al\.\(2024\)A\. Zeng, M\. Liu, R\. Lu, B\. Wang, X\. Liu, Y\. Dong, and J\. TangAgentTuning: enabling generalized agent abilities for LLMs\.InFindings of the Association for Computational Linguistics: ACL 2024,L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 3053–3077\.External Links:[Link](https://aclanthology.org/2024.findings-acl.181/),[Document](https://dx.doi.org/10.18653/v1/2024.findings-acl.181)Cited by:[§5](https://arxiv.org/html/2609.18417#S5.p1.1)\.
- Zhenget al\.\(2024\)Y\. Zheng, R\. Zhang, J\. Zhang, Y\. Ye, and Z\. LuoLlamaFactory: unified efficient fine\-tuning of 100\+ language models\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 3: System Demonstrations\),Y\. Cao, Y\. Feng, and D\. Xiong \(Eds\.\),Bangkok, Thailand,pp\. 400–410\.External Links:[Link](https://aclanthology.org/2024.acl-demos.38/),[Document](https://dx.doi.org/10.18653/v1/2024.acl-demos.38)Cited by:[Appendix A](https://arxiv.org/html/2609.18417#A1.p1.1)\.
- Zhouet al\.\(2023\)C\. Zhou, P\. Liu, P\. Xu, S\. Iyer, J\. Sun, Y\. Mao, X\. Ma, A\. Efrat, P\. Yu, L\. YU, S\. Zhang, G\. Ghosh, M\. Lewis, L\. Zettlemoyer, and O\. LevyLIMA: less is more for alignment\.InAdvances in Neural Information Processing Systems,A\. Oh, T\. Naumann, A\. Globerson, K\. Saenko, M\. Hardt, and S\. Levine \(Eds\.\),Vol\.36,pp\. 55006–55021\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2023/file/ac662d74829e4407ce1d126477f4a03a-Paper-Conference.pdf)Cited by:[§5](https://arxiv.org/html/2609.18417#S5.p2.1)\.

## Appendix ATraining Setting

All SFT systems \(Vanilla SFT, Critical\-Path,L1–L3a\) fine\-tune the sameQwen3\-VL\-30B\-A3B\-Thinkingbase model\([Bai et al\., 2025](https://arxiv.org/html/2609.18417#bib.bib1)\)with full\-parameter updates under identical hyperparameters \(Table[4](https://arxiv.org/html/2609.18417#A1.T4)\)\. Systems differ only in their training trajectories\. The training set is split 90/10 into train/eval, and the best checkpoint per system is chosen by held\-out 4\-benchmark accuracy \(Sec\.[H](https://arxiv.org/html/2609.18417#A8)\)\. Implementation builds on LLaMA\-Factory\([Zheng et al\., 2024](https://arxiv.org/html/2609.18417#bib.bib5)\)\.

HyperparameterValueBase modelQwen3\-VL\-30B\-A3B\-ThinkingUpdate typefull\-parameterEpochs4Learning rate×10−65\\\!\\times\\\!10^\{\-6\}LR schedulecosine, 10% warm\-upPrecisionbf16Optimizer shardingDeepSpeed ZeRO\-3Gradient accumulation2ddp\_timeout×1081\.8\\\!\\times\\\!10^\{8\}sTrain / eval split90 / 10

Table 4:Training setting, shared across all SFT systems\.
## Appendix BInference Setting

We use a standard agent loop with one tool call per assistant turn and at most 128 rounds in total\. Serving is viavLLM\(TP=8=8, bf16\)\. The 128\-round limit acts as a safety net\. We report*stuck rate*as the fraction of evaluation samples that hit this limit \(see Sec\.[4](https://arxiv.org/html/2609.18417#S4)and Sec\.[D](https://arxiv.org/html/2609.18417#A4)\)\.

## Appendix CPer\-Dataset Inference Cost

Tables[5](https://arxiv.org/html/2609.18417#A3.T5)–[6](https://arxiv.org/html/2609.18417#A3.T6)expand the per\-sample message and token columns of Table[2](https://arxiv.org/html/2609.18417#S3.T2)into a system×\\timesdataset grid\. Both metrics are averaged across the samples of each benchmark, and the rightmost column reproduces the 4\-benchmark average reported in the main table\.L2ais uniformly the cheapest in messages on every benchmark \(Table[5](https://arxiv.org/html/2609.18417#A3.T5)\)\. Among SFT systems it is the best on all four benchmarks in tokens \(Table[6](https://arxiv.org/html/2609.18417#A3.T6)\)\. Zero\-shot’s apparently low token counts on LiveVQA, HLE, and MMSearch, and its 119\.26 messages on HLE, are a side\-effect of the stuck\-loop rate of 45\.5% on HLE \(Sec\.[D](https://arxiv.org/html/2609.18417#A4)\), which produces short repeated messages until the round limit\.

SystemSimpleVQALiveVQAHLEMMSearchAvg\.Zero\-shot16\.6918\.71119\.26†12\.0341\.67Vanilla SFT12\.5126\.3136\.6020\.3323\.94Critical\-Path19\.2725\.7943\.4120\.8727\.33L111\.0915\.3721\.5815\.5615\.90L214\.0317\.3436\.5321\.1122\.25L326\.0928\.5639\.9026\.2030\.19L2a10\.8414\.9320\.0411\.3914\.30L3a11\.2515\.6227\.6112\.0416\.63Table 5:Per\-sample averagemessage countby system and dataset\.Bold= best in column \(among SFT systems\)\.†Zero\-shot artefact, see prose above\.SystemSimpleVQALiveVQAHLEMMSearchAvg\.Zero\-shot8,27810,757†11,488†5,3858,977†Vanilla SFT10,10628,17626,88018,59620,940Critical\-Path12,53122,80426,82016,00519,540L18,24118,08116,01216,29314,657L29,13314,46121,75018,01815,841L325,53239,35330,83431,39631,779L2a7,66513,26413,8169,16110,977L3a8,48514,43518,70913,43713,766Table 6:Per\-sample averagetoken countby system and dataset, computed viaQwen2Tokenizeron the full assistant\+user trajectory\.Bold= best in column\.†Zero\-shot artefact, see prose above\.Avg\.Acc\. \(%\)Avg\.\#msgsAcc/msgAcc/tok×103\\times 10^\{3\}Zero\-shot44\.6841\.671\.074\.98†Vanilla SFT48\.3523\.942\.022\.31Critical\-Path44\.3227\.331\.622\.27L150\.0415\.903\.153\.41L248\.2222\.252\.173\.04L348\.0730\.191\.591\.51L2a48\.1014\.303\.364\.38L3a47\.5516\.632\.863\.45Table 7:Cost\-effectiveness summary, restating the bottom block of Table[2](https://arxiv.org/html/2609.18417#S3.T2)\.Bold= best in column\. Both cost\-effectiveness ratios are computed from the row’s average accuracy and the corresponding average message / token count\.†Zero\-shot’s Acc/tok ratio inherits the stuck\-loop artefact in Table[6](https://arxiv.org/html/2609.18417#A3.T6)\.L2astrictly beats every SFT baseline on both ratios\.
## Appendix DStuck\-Rate Note

The per\-dataset stuck\-rate breakdown is given in Sec\.[4](https://arxiv.org/html/2609.18417#S4)\. We note two side observations not covered there\. First, the45\.5%45\.5\\%stuck rate of zero\-shot on HLE explains the anomalously low average\-token figure of zero\-shot in Sec\.[C](https://arxiv.org/html/2609.18417#A3), since stuck rollouts end without an<answer\>and are truncated at the 128\-round ceiling\. Second, the residual elevation ofL3aon HLE \(versusL2a\) is consistent with the format\-mismatch interpretation in Sec\.[4](https://arxiv.org/html/2609.18417#S4), since merge\-think rephrasing cannot fully reshape the multi\-response<tool\_response\>block thatL3produces\.

## Appendix EMerge Granularity and Inference\-Time Format Alignment

L2/L2aandL3/L3ause closely related merge primitives but show distinctly different inference behaviour \(Tables[5](https://arxiv.org/html/2609.18417#A3.T5)–[6](https://arxiv.org/html/2609.18417#A3.T6), Sec\.[D](https://arxiv.org/html/2609.18417#A4)\)\. This appendix explains where the gap arises and states the design rule it suggests\.

#### The Format Gap\.

The inference agent loop is strictly serial\. In the great majority of assistant turns, one<tool\_call\>tag carries a single JSON call, and the next user turn returns one<tool\_response\>with that call’s result\. Merging preserves the canonical outer tags on both sides, but concatenates thek≥2k\\\!\\geq\\\!2siblings’*bodies*inside them\. The merged<tool\_call\>now containskkJSON calls, and the merged<tool\_response\>, prefixed “Previous tool call results:”, contains thekkcorresponding results\. A round whose two outer tags holdkkitems each is a configuration the inference loop does not emit \(Fig\.[5](https://arxiv.org/html/2609.18417#A5.F5)\)\.

\(a\) Merged training roundassistant
<think\>…</think\>
<tool\_call\>\{call\_a\} \{call\_b\}</tool\_call\>user
<tool\_response\>
Previous tool call results:
resp\_a; resp\_b
</tool\_response\>\(b\) Inference\-time loopassistant:<think\>…</think\><tool\_call\>call\_x</tool\_call\>user:<tool\_response\>resp\_x</tool\_response\>assistant:<think\>…</think\><tool\_call\>call\_y</tool\_call\>user:<tool\_response\>resp\_y</tool\_response\>Figure 5:Message pattern of a merged training round \(a\) vs\. a typical inference\-time loop \(b\)\. In \(a\) a single<tool\_call\>tag holdskkJSON calls and a single<tool\_response\>tag holds thekkcorresponding results, whereas the loop in \(b\) usually emits one call/result per turn\.
#### Why L2 and L3 Differ\.

Both edits create the multi\-item pattern above, but at different rates and with different downstream structure \(Table[8](https://arxiv.org/html/2609.18417#A5.T8)\)\.*\(i\) Frequency\.*L2requires matching parents*and*children and yields1,0441\{,\}044fusions;L3matches parents only and yields3,2343\{,\}234\(about3×3\\times\), mostly as more merges per sample \(samples affected rise only19\.0%19\.0\\%→\\to21\.3%21\.3\\%\)\.*\(ii\) Downstream alignment\.*L2merges only siblings that feed the same next round, so each merged user turn still has a single consumer—the same “one user turn→\\toone next<think\>” pattern as inference\.L3can bundle siblings that originally fed different children, gaining denser compression at the cost of that one\-to\-one structure; keeping the stricterL2rule is the more inference\-aligned choice\.

Merge group sizeL2L327092,4913230586483122≥5\\geq 52235Total fusions1,0443,234Samples affected1,175 \(19\.0%\)1,322 \(21\.3%\)Table 8:Merge groups by size and fraction of training samples affected, on the 6,196\-trajectory corpus\.
#### Design Rule\.

A round\-level merge should preserve the downstream consumption pattern of the original trajectory\. Concretely, require the merge criterion to match on both parents*and*children, and prefer leaf\-pruning over fusion when downstream consumers differ\. This keeps the training\-time user\-turn distribution close to what the inference loop produces\.

VanillaSFT\+DPO\[CP\]\+DPO\[L1\]ΔV\\Delta\_\{V\}\+DPO\[L2\]ΔV\\Delta\_\{V\}\+DPO\[L3\]ΔV\\Delta\_\{V\}\+DPO\[L2a\]ΔV\\Delta\_\{V\}\+DPO\[L3a\]ΔV\\Delta\_\{V\}*Taskacc\.*SimpleVQA68\.67diverged†70\.00\+1\.3\+1\.370\.67\+2\.0\+2\.067\.33−1\.3\-1\.368\.33−0\.3\-0\.366\.67−2\.0\-2\.0LiveVQA51\.00diverged†52\.00\+1\.0\+1\.053\.67\+2\.7\+2\.750\.67−0\.3\-0\.350\.33−0\.7\-0\.750\.00−1\.0\-1\.0HLE9\.39diverged†12\.42\+3\.0\+3\.012\.12\+2\.7\+2\.712\.73\+3\.3\+3\.312\.42\+3\.0\+3\.010\.91\+1\.5\+1\.5MMSearch64\.33diverged†54\.39−9\.9\-9\.959\.06−5\.3\-5\.357\.89−6\.4\-6\.453\.22−11\.1\-11\.160\.23−4\.1\-4\.1Avg\. Acc\.48\.35diverged†47\.20−1\.2\-1\.248\.88\+0\.5\+0\.547\.16−1\.2\-1\.246\.08−2\.3\-2\.346\.95−1\.4\-1\.4*Inf\.eff\.*Avg\. \#msgs23\.94diverged†11\.60−12\.3\-12\.311\.43−12\.5\-12\.512\.03−11\.9\-11\.911\.25−12\.7\-12\.710\.23−13\.7\-13\.7Avg\. \#toks \(10310^\{3\}\)20\.9diverged†4\.8−16\.1\-16\.14\.5−16\.5\-16\.54\.7−16\.2\-16\.24\.6−16\.4\-16\.44\.0−16\.9\-16\.9*Cost\-eff\.*Acc/msg2\.02diverged†4\.07\+2\.0\+2\.04\.28\+2\.3\+2\.33\.92\+1\.9\+1\.94\.09\+2\.1\+2\.14\.59\+2\.6\+2\.6Acc/10310^\{3\}tok2\.31diverged†9\.85\+7\.5\+7\.510\.98\+8\.7\+8\.79\.99\+7\.7\+7\.710\.10\+7\.8\+7\.811\.63\+9\.3\+9\.3Table 9:Shared\-base DPO from Vanilla SFT\.ΔV\\Delta\_\{V\}: \+DPO minus Vanilla SFT\.†Diverged \(context overflow\)\.Critical\-P\.L1L2L3L2aL3aSFT\+DPOΔ\\DeltaSFT\+DPOΔ\\DeltaSFT\+DPOΔ\\DeltaSFT\+DPOΔ\\DeltaSFT\+DPOΔ\\DeltaSFT\+DPOΔ\\Delta*Taskacc\.*SimpleVQA60\.6768\.00\+7\.3\+7\.370\.6768\.00−2\.7\-2\.768\.6764\.67−4\.0\-4\.066\.6766\.670067\.3368\.67\+1\.3\+1\.364\.3366\.00\+1\.7\+1\.7LiveVQA46\.3348\.67\+2\.3\+2\.353\.6751\.00−2\.7\-2\.751\.3349\.33−2\.0\-2\.050\.3351\.00\+0\.7\+0\.751\.0050\.67−0\.3\-0\.351\.0052\.67\+1\.7\+1\.7HLE11\.218\.48−2\.7\-2\.710\.3010\.91\+0\.6\+0\.610\.3012\.12\+1\.8\+1\.811\.5213\.33\+1\.8\+1\.811\.5214\.55\+3\.0\+3\.08\.7910\.91\+2\.1\+2\.1MMSearch59\.0658\.48−0\.6\-0\.665\.5062\.57−2\.9\-2\.962\.5760\.82−1\.8\-1\.863\.7466\.08\+2\.3\+2\.362\.5759\.06−3\.5\-3\.566\.0863\.74−2\.3\-2\.3Avg\. Acc\.44\.3245\.91\+1\.6\+1\.650\.0448\.12−1\.9\-1\.948\.2246\.73−1\.5\-1\.548\.0749\.27\+1\.2\+1\.248\.1048\.24\+0\.1\+0\.147\.5548\.33\+0\.8\+0\.8*Inf\.eff\.*Avg\. \#msgs27\.3314\.12−13\.2\-13\.215\.9013\.86−2\.0\-2\.022\.2516\.52−5\.7\-5\.730\.1922\.08−8\.1\-8\.114\.3014\.48\+0\.2\+0\.216\.6315\.60−1\.0\-1\.0Avg\. \#toks \(10310^\{3\}\)19\.55\.2−14\.4\-14\.414\.75\.3−9\.4\-9\.415\.86\.1−9\.8\-9\.831\.88\.8−23\.0\-23\.011\.05\.5−5\.5\-5\.513\.85\.8−7\.9\-7\.9*Cost\-eff\.*Acc/msg1\.623\.25\+1\.6\+1\.63\.153\.47\+0\.3\+0\.32\.172\.83\+0\.7\+0\.71\.592\.23\+0\.6\+0\.63\.363\.33−0\.0\-0\.02\.863\.10\+0\.2\+0\.2Acc/10310^\{3\}tok2\.278\.90\+6\.6\+6\.63\.419\.14\+5\.7\+5\.73\.047\.71\+4\.7\+4\.71\.515\.62\+4\.1\+4\.14\.388\.80\+4\.4\+4\.43\.458\.28\+4\.8\+4\.8Table 10:Matched\-base DPO on each setting’s own SFT checkpoint\.Δ\\Delta: \+DPO minus that setting’s SFT\.

## Appendix FDPO on Refined\-vs\-Raw Preference Pairs

Each structural edit yields a free preference pair that shares the final answer, taking the refined trajectory as “chosen” and the raw trajectory as “rejected” \(counts: Critical\-Path4,4464\{,\}446;L1986986;L2/L2a1,1751\{,\}175;L3/L3a1,3221\{,\}322\)\. We plug these pairs into a follow\-up DPO\([Rafailov et al\., 2024](https://arxiv.org/html/2609.18417#bib.bib22)\)stage in two recipes that differ only in the DPO starting checkpoint\.

#### Shared\-Base Recipe\.

Train one Vanilla SFT model on raw trajectories, then run DPO from that shared base while swapping each setting’s preference pairs, Table[9](https://arxiv.org/html/2609.18417#A5.T9)\. All five transferable settings collapse inference cost far below Vanilla SFT\.L2pairs are strongest on accuracy,L3apairs on cost\-efficiency\. Critical\-Path pairs push Vanilla SFT off distribution and the run diverges\.

#### Matched\-Base Recipe\.

Start DPO from each setting’s own SFT checkpoint, Table[10](https://arxiv.org/html/2609.18417#A5.T10)\. Accuracy is better preserved than under the shared base, while tokens still drop, though they remain above the shared\-base band\. Shared\-base is stronger for cost\-efficient deployment\. Matched\-base is safer when raw accuracy is the priority\.

## Appendix GDPO Setting

We train with sigmoid DPO plus an NLL\-on\-chosen anchor \(Table[11](https://arxiv.org/html/2609.18417#A7.T11)\)\.L1uses200200steps andαRPO=0\.7\\alpha\_\{\\mathrm\{RPO\}\}=0\.7; others use300300steps andαRPO=0\.5\\alpha\_\{\\mathrm\{RPO\}\}=0\.5\. Checkpoints every5050steps are selected by mix\-dev accuracy rather than eval\-loss\.

HyperparameterValueLosssigmoid DPO \+ NLL\-on\-chosenβ\\beta\(DPO temp\.\)0\.1αRPO\\alpha\_\{\\mathrm\{RPO\}\}\(NLL wt\.\)0\.5 \(0\.7 forL1\)Max steps300 \(200 forL1\)Learning rate×10−75\\\!\\times\\\!10^\{\-7\}LR schedulecosine, 5% warm\-upMax seq\. length16,384Per\-device batch1Grad\. accum\.2Global batch16Hardware8×808\\times 80GiB GPUsOptimiserDeepSpeed ZeRO\-3AttentionSDPA

Table 11:DPO training setting \(shared across SFT bases\)\.Figure 6:Eval\-loss curves \(left, dense schedule\) vs\. downstream 4\-benchmark accuracy \(right, four evaluated checkpoints\) forL1,L2a,L3a\.
## Appendix HCheckpoint Selection

Eval\-loss on the held\-out 10% of the training data was recorded at every 200 steps \(Table[12](https://arxiv.org/html/2609.18417#A8.T12), top half, and Fig\.[6](https://arxiv.org/html/2609.18417#A7.F6)left\)\. Downstream task accuracy, however, requires running the full agent loop on the four benchmarks for every candidate checkpoint, which is expensive\. Within our compute budget we evaluated four checkpoints for our edits, namely steps 200, 600, 800, and 1396 \(final\), and a single checkpoint at step 600 for Vanilla SFT and Critical\-Path \(selected by the same rule applied to our edits\)\. A denser sweep is left to follow\-up work\.

StepRawCPL1L2aL3a*eval\-loss \(held\-out 10% of training data\)*2000\.62430\.68270\.64450\.64940\.65454000\.62720\.69020\.64940\.65320\.65736000\.62450\.68480\.64400\.64800\.65318000\.66630\.74340\.69520\.69990\.705610000\.67240\.74240\.68970\.69670\.699712000\.71810\.80800\.74360\.75120\.751813500\.71850\.80920\.74350\.75040\.7537*downstream task accuracy \(%, 4\-bench avg\.\)*200——47\.9147\.7346\.5760048\.3544\.3250\.0447\.6846\.90800——46\.4547\.8947\.221396——45\.1148\.1047\.55Table 12:Eval\-loss vs\. downstream accuracy\.#### Best\-Checkpoint Summary\.

By 4\-benchmark average accuracy,L1, Vanilla SFT, and Critical\-Pathpick step600, whileL2a and L3aactually pick step1396\.

#### Eval\-Loss Decouples from Accuracy\.

All three ofL1,L2a,L3afollow a similar eval\-loss trajectory \(minimum near step 600, then rising\), but downstream accuracy diverges\.L1peaks at step 600 \(50\.04%50\.04\\%\), drops to46\.45%46\.45\\%at step 800, and reaches45\.11%45\.11\\%at step 1396 \(−4\.9\-4\.9pp from peak\)\.L2aandL3ainstead climb*monotonically*from step 600 onwards \(→→48\.1047\.68\\\!\\to\\\!47\.89\\\!\\to\\\!48\.10forL2a, and→→47\.5546\.90\\\!\\to\\\!47\.22\\\!\\to\\\!47\.55forL3a, all in %\)\. Picking the checkpoint by eval\-loss alone would have chosen step 600 for everyone, which is exactlyL1’s peak but a poor choice forL2a/L3a\. Practitioners training on merged trajectories should therefore select checkpoints by downstream metrics rather than eval\-loss\. Continuing training beyond step 1396 would extend the 4\-epoch SFT budget and require evaluating additional checkpoints on the full agent loop, which doubles inference cost per added checkpoint\. Eval\-loss has been rising since step 600, so further training carries a non\-trivial overfitting risk despite the still\-climbing downstream metric forL2a/L3a\. A larger compute\-budget sweep is left to follow\-up work\.

## Appendix IQualitative Example

#### Source Sample\.

Training\-corpus sample\_idx=1581\. The user shows an image carrying “SAGE Publishing” branding and a “Live Product Training” banner, and asks:*“Who founded SAGE Publishing, and within the first decade after that, who established the company’s first international office in London?”*The gold answer ends with “Sara Miller McCune \(with George McCune\) founded SAGE in 1965, and SAGE Publications Ltd\. was established in London in 1971”\.

In what follows, each grey box is one round of the original trajectory, and the bar at the top names the round and its role\. Round 1 is the user question \(image \+ text\)\. Rounds 2–4 are assistant turns whosethinkandcallare followed by a user\-sideresp\. Round 5 is the final assistantanswer\. Red\-bordered boxes mark rounds removed by an edit, and teal\-bordered boxes mark rounds produced by merging\.

### Original Trajectory \(8 Messages, 5 Rounds\)

R1\(user\)\[image: SAGE Publishing logo on a Live Product Training banner\] Who founded SAGE Publishing, and within the first decade after that, who established the company’s first international office in London?

R2\(assistant \+ user\)thinkIdentify who founded SAGE Publishing, then look up the London office\. Start with the founder\. callweb\_search\("Who founded SAGE Publishing"\) resp“24\-year\-old Sara Miller McCune founded Sage in 1965 …” \(Wikipedia / About\-SAGE\)

R3\(assistant \+ user\)thinkNow find who set up SAGE’s first international office in London\. callweb\_search\("SAGE first international office London established by who and when"\) resp*Same*About\-SAGE snippet as R2, with no new fact about the London office\.

R4\(assistant \+ user\)thinkThe previous result added nothing new\. Try SAGE’s own 50\-year history page directly\. callvisit\_url\("https://us\.sagepub\.com/\.\.\./sage\-50\-year\-history\.pdf"\) resp“1971: SAGE Publications, Ltd\. established in London\. Sara and George hired David Brooks …”

R5\(assistant\)thinkI have what I need\. answerSara Miller McCune \(with George McCune\) founded SAGE, and the first international office was established in London in 1971\.

### Annotated DAG \(Prompt A\)

Running the strict prompt \(Sec\.[J](https://arxiv.org/html/2609.18417#A10), Prompt A\) on this trajectory, the annotator LLM emits three edges with the following abridged reasons\.

- •R2→\\toR3: “Round 2 produced the concrete founder identification*Sara Miller McCune*, which Round 3 uses to frame the next search\.”
- •R4→\\toR5: “Round 4 produced the key historical fact that*SAGE Publications, Ltd\. was established in London in 1971*, which Round 5 uses to conclude the first international office was in London within the first decade\.”
- •R2→\\toR5: “Round 2 established Sara Miller McCune as the founder, which Round 5 combines with the London\-1971 fact to answer the full question\.”

#### No Edge Originates at R1\.

Prompt A asks the annotator to mark only globally load\-bearing edges, where roundjjuses a specific retrieved fact that roundiiproduced\. The user query itself \(R1\) is the topic of every later round but contributes no*new retrieved fact*that a downstream round consumes, so the strict criterion correctly omits allR1→\\to\*edges\. Two structural guarantees keep this safe\-by\-construction\. First, leaf pruning removes only nodes with out\-degree 0 that are neither round 1 nor the highest\-indexed round, so R1 is never a deletion candidate even when it has no edges in the annotated DAG\. Second, the merge function \(Sec\.[2\.4](https://arxiv.org/html/2609.18417#S2.SS4)\) skips any merge group whose smallest round index is below 2, so R1’s message is left untouched even if it appeared in a merged set\. As a result the merge topology is identical whether or not the annotator emitsR1→\\to\*edges\.

The resulting DAG over annotated edges isR2→\\toR3,R2→\\toR5, andR4→\\toR5\. R3 has out\-degree 0 \(no downstream round uses anything it produced\), which marks it as the L1 leaf\.

### AfterL1\(Leaf Prune\): 6 Messages, 4 Rounds

R3 has out\-degree 0 in the annotated DAG, soL1drops it\. R1, R2, R4, R5 remain unchanged\.

R3\(pruned byL1: out\-degree 0\)think*\(removed\)*call*\(removed\)*resp*\(removed\)*The kept trajectory afterL1isR1→\\toR2→\\toR4→\\toR5, i\.e\. the original boxes for R1, R2, R4, R5 unchanged \(not redrawn\)\.

### AfterL2a\(Strict Merge \+ LLM Rephrase\): 4 Messages, 3 Rounds

In the annotated DAG afterL1, R2 and R4 both have an empty parent set \(Prompt A emits noR1→\\to\*edge\) and the same child set \(\{R5\}\\\{\\text\{R5\}\\\}\)\. The strict\-merge rule fuses them into one roundR2\+4\. Note that even if Prompt A had emittedR1→\\toR2andR1→\\toR4, both rounds would then share parents\{R1\}\\\{\\text\{R1\}\\\}and child\{R5\}\\\{\\text\{R5\}\\\}, and the same merge would fire\. Their twothinkbodies are concatenated and rephrased by the annotator LLM into a single coherent passage \(the<think\>shown below\)\. The two original calls are emitted as two JSON arguments inside a single<tool\_call\>tag, and the tworespbodies are prefixed with “Previous tool call results:” and emitted in one user turn\.

R2\+4\(merged byL2a, covering original R2 and R4\)thinkThe image shows SAGE Publishing’s branding on a Live Product Training banner\. The question has two parts \(who founded SAGE Publishing and who established its first London office within the first decade\) and these are likely connected, so first locate the founder, then look at SAGE’s early history for when and how the London office was set up\.
callweb\_search\("Who founded SAGE Publishing"\),
visit\_url\("https://us\.sagepub\.com/\.\.\./sage\-50\-year\-history\.pdf"\)
respPrevious tool call results:
\(1\) “24\-year\-old Sara Miller McCune founded Sage in 1965 …”
\(2\) “1971: SAGE Publications, Ltd\. established in London\. Sara and George hired David Brooks …”The fullL2atrajectory isR1→\\toR2\+4→\\toR5\(R5 unchanged from the original\)\.

#### Take\-Aways\.

Two distinct effects compose on the same trajectory\. First, Leaf Prune \(L1\) removes the redundant follow\-up search whose evidence was already in R2\. Second, Strict Merge with rephrase \(L2a\) folds the two substantively different sub\-queries \(“who founded” vs\. “where was the first international office”\) into one round of reasoning whose<think\>reads as a single chain of thought, not two disjoint segments concatenated with a semicolon\. The final answer in R5 is unchanged\.

### Two Further DAG\-Only Cases

To isolate each edit’s topological action, we show two more samples from the training corpus as bare DAGs, omitting the per\-round content\. The first sample triggers only L1, the second only L3\.

#### Sample A \(\_idx=6061\), L1 Only\.

The annotated DAG has edgesR1→\\toR2,R1→\\toR3,R3→\\toR4\. Round R2’s out\-degree is 0 and it is not the highest\-indexed round, soL1prunes it\. L2 and L3 fire on sibling groups, and the remaining graph has no siblings sharing parents, so both are no\-ops\. Trajectory shrinks from 6 messages to 4\.

annotated DAGR1R2R3R4L1after L1R1R3R4
#### Sample C \(\_idx=13\), L3 Only\.

The annotated DAG has edgesR1→\\toR2,R1→\\toR4,R2→\\toR3,R3→\\toR5,R4→\\toR5\. No round is a deletable leaf, soL1is a no\-op\. ForL2, R2’s children are\{R3\}\\\{\\text\{R3\}\\\}while R4’s are\{R5\}\\\{\\text\{R5\}\\\}, so the strict criterion \(same parents*and*children\) does not fire either\.L3requires only shared parents, and R2 and R4 both have parent\{R1\}\\\{\\text\{R1\}\\\}, so they merge into one roundR2\+4\. The remaining edges becomeR1→\\toR2\+4,R2\+4→\\toR3,R2\+4→\\toR5,R3→\\toR5\. Trajectory shrinks from 8 messages to 6\.

annotated DAGR1R2R4R3R5L3after L3R1R2\+4R3R5Together with Sample B \(\_idx=1581\) above, the three samples isolate the three edit primitives\. Real trajectories often chain them, as Sample B does \(L1 then L2a\)\.

## Appendix JPrompt Templates

This appendix gathers LLM prompts used in the paper:

- •Prompt A, DAG annotation: extracts the round\-level dependency DAG used by all our edits \(Sec\.[2\.2](https://arxiv.org/html/2609.18417#S2.SS2)\)\.
- •Prompt B, Critical\-Path KEEP/REMOVE: the LLM\-deletion baseline against which we compare \(Sec\.[3\.1](https://arxiv.org/html/2609.18417#S3.SS1)\)\.
- •Prompt C, Merge\-think rephrase: turns raw\-concatenated<think\>blocks into a single coherent passage forL2a/L3a\(Sec\.[2\.4](https://arxiv.org/html/2609.18417#S2.SS4)\)\.
- •Prompt D, LLM\-as\-judge: scores the model’s final<answer\>against the gold answer\. Used as the task\-accuracy metric in all tables \(Sec\.[3\.1](https://arxiv.org/html/2609.18417#S3.SS1)\)\.

Prompt A: DAG annotation\# Role
You are an expert in logical reasoning and causal analysis\. Your task is to analyze a multi\-round AI Agent trajectory and deconstruct it into a Directed Acyclic Graph \(DAG\) that capturesonly the dependencies that actually matter for producing the final answer\.\# Task Description
The trajectory is organized by ‘‘Rounds\.’’ You will be told which round contains the final answer \(typically the last round\)\. Your job is to decide, for every candidate edge Round i \-\> Round j, whether Round j would have been*impossible or clearly wrong*without Round i’s concrete contribution \-\-\- judgedglobally against the final answer, not by local narrative flow\.\# Definition of a Dependency Edge \(Round i \-\> Round j\)
An edge existsonly if ALL of the following are true:
1\.Concrete carry\-over: A specific fact, number, entity, URL, identifier, or inferred conclusion that Round j actually*uses*was first produced in Round i \-\-\- and cannot be obtained from any other earlier round or from the original query alone\.
2\.Globally load\-bearing: Removing Round i would break Round j’s ability to make progress toward the final answer in Round N\. If Round j could have been produced by skipping Round i \(perhaps with minor rewording\), do NOT add the edge\.
3\.Not mere narrative / chronological continuity: Do NOT add an edge just because Round j’s <think\> text mentions, reacts to, or rhetorically references Round i\.\# Anti\-Patterns \-\-\- Do NOT Add an Edge When:
•Dead\-end rounds: Round i returned information not used anywhere downstream \(including Round N\)\.
•Redundant confirmation: Round j merely re\-verifies a fact already established in an earlier round\.
•Parallel independent lookups: Round i and j are both sub\-queries derived from a common ancestor, with no information flowing from one to the other\.
•Pure stylistic/continuity reference: ‘‘continuing from the previous step’’ without consuming any output of Round i\.\# Be Aggressive About Pruning Edges
Default to NOT adding an edge\. When in doubt, ask: ‘‘If I deleted Round i from the transcript, would Round j still reach the same conclusion \(perhaps via trivial rewording\)?’’ If yes \-\-\- do not add the edge\.\# Output Format
JSON edge list, each entry naming the concrete fact/artifact carried:
\[\{"source": "Round X", "target": "Round Y",
"reason": "Round X produced <fact\> which
Round Y uses to <purpose\>"\}, \.\.\.\]\# Input Trajectory
\{INPUT\_TRAJ\_STRING\}Prompt B: Critical\-Path KEEP/REMOVE baseline \(English translation of the production prompt\)System\.You are a rigorous, conservative cleaning assistant for multi\-modal Agent research data\. Your task is to prune unnecessary reasoning / tool\-call rounds so the training data is more efficient, while keeping the remaining trajectory logically coherent and self\-consistent\. Output compact JSON only; do not wrap in markdown code fences\.User\.
\# Background
Below is one Agent trajectory\. Each middle round consists of anassistantturn \(with <think\> and <tool\_call\>\) followed by auserturn \(tool\_response\)\. The final assistant turn emits <answer\>\.\# Goal
Decide which middle rounds can bedeleted entirely, such that the trajectory becomes more concise without changing the correctness of the final answer\.\# Typical removable rounds
1\.Failed / empty result: tool\_response is ‘‘Page not found’’, empty, or irrelevant, and no later round adjusts strategy based on it\.
2\.Redundant repeat: this tool\_call overlaps a prior one, information already obtained\.
3\.Abandoned branch: this round tried a direction but no later think/tool\_call uses its output, and the final answer is unrelated\.
4\.Extra re\-verification: the key fact already entered the answer via an earlier round; this round only reconfirms\.
5\.Over\-long reasoning prep: the tool\_call output of this round does not appear in the evidence chain for the final answer\.\# Non\-removable rounds
Concrete artefacts of this round \(number, year, name, link, image description, …\) appeardirectly in the final answer; OR some kept round’s think / tool\_call explicitly references this round’s output; OR this round is a non\-skippable step in the reasoning chain\.\# Coherence hard constraint
After deleting a round, the next kept round’s <think\> prefix must not read as a reference to the deleted content \(e\.g\., ‘‘let’s try another link’’, ‘‘that didn’t work’’, ‘‘another search’’, ‘‘hmm’’\)\. To repair such*dangling references*, you may supply aminimal rewrite\(patch\_think\) for the affected kept round\. The rewrite must preserve its conclusion and subsequent tool\_call, only altering the opening one or two sentences\. null means no rewrite needed\.\# Conservative principle
When in doubt,keepthe round\. If no round can be safely deleted, return remove: \[\]\.\# Output Format
Strict JSON, no markdown:
\{
"status": "reduce" \| "non\_reducible" \| "illogical",
"remove": \[<round\_id\_int\>, \.\.\.\],
"patch\_think": \{ "<id\>": "<rewritten think\>", \.\.\. \},
"reason": "<short justification\>"
\}
• reduce: at least one deletable round\.
• non\_reducible: all rounds load\-bearing or patch\_think cannot fix coherence\.
• illogical: trajectory itself is broken \(answer mismatched, final answer missing, …\)\.
• remove must not include Round 1 or the final answer round\.\# Input Trajectory
\{INPUT\_TRAJ\_STRING\}Prompt C: Merge\-think rephrase \(used by L2a, L3a\)Your task is to merge multiple <think\></think\> segments from a trajectory of model outputs into a single coherent reasoning passage\.Context: These <think\> segments originally came from separate turns in a multi\-turn user\-\-assistant conversation, where each <think\> block was followed by tool\_call arguments\. It has now been determined that the think contents and the tool\_calls can each be consolidated into a single turn\. Your job is to merge the think segments into one logical, coherent reasoning passage\.Requirements:
• Preserve the original intent and logical flow of the reasoning\.
• If the original think segments reference or describe tool\_calls \(tool invocations\), you must retain that meaning \-\-\- do not drop references to which tools are being used or why\.
• Smooth out the transitions: the current think contents are crudely concatenated with semicolons \(;\) and line breaks\. Rewrite them so they read as one continuous, natural chain of thought rather than disjoint fragments\.
• Do not add new reasoning that wasn’t present in the original; only restructure and connect what is already there\.
• Compress for token efficiency without losing information\. Eliminate redundancy, repeated context, filler phrases, and verbose restatements\. Merge overlapping points \-\-\- but every distinct piece of information, decision, and tool\-call reference must still be present in the output\.Input \(concatenated think contents\):
\{CONCATED\_THINK\}Output format: Plain text only\. Do not wrap the output in any special tags \(no <think\>, no XML, no markdown code fences\)\.Prompt D: LLM\-as\-judge \(gpt\-5\-nano\)Your job is to look at a question, a gold target, and a predicted answer, and then assign a grade of either \[CORRECT, INCORRECT, NOT\_ATTEMPTED\]\. First, examples of each grade; then a new example\.CORRECT examples\(predicted answer is equivalent to the gold target\):
Q: ‘‘What are the names of Barack Obama’s children?’’; Gold: ‘‘Malia Obama and Sasha Obama’’\. Accept: ‘‘sasha and malia obama’’; ‘‘most people would say Malia and Sasha, but I’m not sure…’’; ‘‘…Malia Ann and Natasha Marian, but commonly Malia and Sasha…’’\. A predicted answer is CORRECT iff it fully contains the gold information, contradicts nothing in it, and only semantic meaning matters \(case / punctuation / order do not\)\. Hedging is allowed as long as the gold is fully included\.INCORRECT examples\(predicted answer contradicts the gold\):
‘‘Malia\.’’; ‘‘Malia, Sasha, and Susan\.’’; ‘‘Obama has no children\.’’; ‘‘either Malia and Sasha\. Or Malia and Jackie…’’\. A factual statement contradicting the gold is INCORRECT, even if hedged\.NOT\_ATTEMPTED examples\(predicted answer neither contains nor contradicts the gold\):
‘‘I don’t know\.’’; ‘‘I need more context\.’’; ‘‘He has two children\. I know one is Malia, but I’m not sure about the other\.’’Additional rules:
• Numbers must be correct to the last significant figure in the gold \(‘‘120k’’→\\toaccept 115k\-\-124k; reject 100k or 113k\)\.
• Gold may carry extra info beyond the question; predicted need only cover what the question asked\.
• Do not punish omissions clearly inferred from the question \(‘‘San Francisco’’ for ‘‘San Francisco, California’’\)\.
• Tolerate typos in names if clearly the same person\.Here is a new example\. Simply reply with A, B, or C \(no other text\)\.
Question: \{query\}
Gold target: \{reference\_answer\}
Predicted answer: \{generated\_answer\}Grade as one of:
A: CORRECT
B: INCORRECT
C: NOT\_ATTEMPTED
Return only the single letter\.

Similar Articles

Offline Preference-Based Trajectory Evaluation

arXiv cs.LG

This paper proposes offline preference-based trajectory evaluation for agentic systems, which compares trajectories via temporal preferences rather than binary success metrics. It shows that this approach reduces ties from roughly 75% to 35%, improving discriminative power and data efficiency across diverse benchmarks.