GraphForge:通过图锚定工作区合成训练可用的智能体
摘要
GraphForge 提出了一种证据图框架,将智能体的任务生成与验证共同锚定在真实文件之上,从而能够为工作场景中的智能体合成可验证的训练轨迹。在 2,169 条 GraphForge 轨迹上对 Qwen3.6-27B 进行微调后,模型在 GDPVal、Workspace-Bench-Lite 和 SpreadsheetBench II 上均有所提升,而采用拒绝微调(rejection fine-tuning)则能进一步取得增益。
查看缓存全文
缓存时间: 2026/10/01 09:45
# GraphForge: Training Working Agents with Graph-Anchored Workspace Synthesis
Source: [https://arxiv.org/html/2609.38923](https://arxiv.org/html/2609.38923)
Qisheng SuAffiliation:University of Science and Technology of ChinaAffiliation:Shanghai Innovation InstituteEmail:[nicksu@mail\.ustc\.edu\.cn](mailto:)Guanru ZhuAffiliation:Fudan UniversityHuicheng JiangAffiliation:Fudan UniversityQiuyinzhe ZhangAffiliation:University of Science and Technology of ChinaAffiliation:Shanghai AI LaboratoryKou ShiAffiliation:University of Science and Technology of ChinaZhen FangAffiliation:University of Science and Technology of ChinaZiao ZhangAffiliation:University of Science and Technology of ChinaQingnan RenAffiliation:University of Science and Technology of ChinaZehui ChenAffiliation:University of Science and Technology of ChinaTao GuiAffiliation:Fudan UniversityAffiliation:Shanghai AI LaboratoryFeng ZhaoAffiliation:University of Science and Technology of China
###### Abstract
Working agents need to read diverse files, coordinate tools, and produce deliverables\. Training such agents requires tasks built on many real files with verifiable results, but few pipelines exist to synthesize this kind of data\. Existing pipelines either generate files with models, which lack realism and diversity, or build tasks on real files without task\-specific verifiers, leaving result quality unchecked\. We introduce GraphForge, an evidence\-graph based framework that grounds both the task and its verification in real files\. Starting from occupation\-grounded seeds for controlled diversity, GraphForge assembles a workspace of real files for each seed and builds an evidence graph over their relations\. Since the task statement and rubrics are both derived from this graph, task requirements are backed by the workspace files and each criterion is anchored to the files needed to verify it\. An initial rollout further tests executability, and a revision agent repairs the task and rubrics against the original files before trajectories are collected\. Fine\-tuning Qwen3\.6\-27B on 2,169 GraphForge trajectories brings GDPVal to 1445\.7 \(\+65\.7\) under OpenHands, and Workspace\-Bench\-Lite and SpreadsheetBench II to 63\.7 \(\+7\.7\) and 24\.0 \(\+13\.7\) under Claude Code\. Rejection fine\-tuning on the SFT model’s own rollouts, with candidates selected by the evidence\-anchored rubrics, yields further improvements on all three benchmarks, suggesting that the rubrics provide a useful selection signal\. The data and models are available at[https://huggingface\.co/collections/groundhogLLM/graphforge](https://huggingface.co/collections/groundhogLLM/graphforge)\.
## 1Introduction
Large language model agents are moving from conversation to real work\. Working agents such as OpenClaw\([OpenClaw Team, 2026](https://arxiv.org/html/2609.38923#bib.bib19)\)and Hermes\-Agent\([Hermes\-Agent Team, 2026](https://arxiv.org/html/2609.38923#bib.bib20)\)act as persistent digital assistants, handling long\-horizon tasks across file systems, databases, and terminal shells\. Benchmarks such as Claw\-Eval\([Ye et al\., 2026](https://arxiv.org/html/2609.38923#bib.bib18)\), GDPVal\([Patwardhan et al\., 2025](https://arxiv.org/html/2609.38923#bib.bib16)\), and Workspace\-Bench\([Tang et al\., 2026](https://arxiv.org/html/2609.38923#bib.bib15)\)evaluate these agents on realistic work tasks\. Working agents read diverse files, coordinate tools, and produce deliverables that others can use\. For agents more broadly, a common training approach is supervised fine\-tuning on synthesized trajectories produced by a strong teacher model\([Chu et al\., 2026](https://arxiv.org/html/2609.38923#bib.bib10);[Dong et al\., 2026](https://arxiv.org/html/2609.38923#bib.bib3);[Shi et al\., 2026](https://arxiv.org/html/2609.38923#bib.bib11)\)\. Applying this approach to working agents requires tasks built on many real files, with verifiable results\.
Two recent pipelines synthesize training data for working agents\. EnvCraft\([Zeng et al\., 2026](https://arxiv.org/html/2609.38923#bib.bib1)\)generates files with a model and checks each task with a Python script over the workspace state, which cannot read file contents and may overlook errors in document deliverables\. NexForge\([Zhao et al\., 2026](https://arxiv.org/html/2609.38923#bib.bib2)\)builds tasks on real files, but provides no task\-specific rubrics or verifiers, so result quality cannot be systematically checked\. It remains hard to construct diverse tasks on real files and to equip them with reliable verification rubrics\.
To make progress on both fronts, we introduce GraphForge, an evidence\-graph based framework for synthesizing working agent training data\. Our design gives seeds and files separate roles\. Tasks invented freely by a model tend to collapse toward frequent occupations and generic task types\. Grounded seeds therefore fix the task direction, keeping diversity controllable, while the concrete task and its verification are derived from the files\. Our seeds are drawn from O\*NET occupations and their official work activities, and for each seed, an agent crawls real files to form a workspace, over which a model builds an evidence graph of cross\-file relations\. The graph is then compiled into task statements and rubrics, so task requirements are backed by the workspace files and each criterion is anchored to the files needed to verify it\. An initial rollout further tests executability, and a revision agent repairs the task and rubrics against the original files before trajectories are collected\.
We validate our framework by training Qwen3\.6\-27B on the synthesized data\. Supervised fine\-tuning on 2,169 trajectories from GraphForge brings GDPVal to 1445\.7 \(\+65\.7\), Workspace\-Bench\-Lite to 63\.7 \(\+7\.7\), and SpreadsheetBench II to 24\.0 \(\+13\.7\)\. The same data also improves Qwen3\.6\-35B\-A3B, suggesting that GraphForge trajectories generalize across base models\. Rejection fine\-tuning on the SFT model’s own rollouts on new queries disjoint from the SFT data yields further improvements on all three benchmarks, suggesting that the evidence\-anchored rubrics provide a useful selection signal\.
Our contributions are as follows\.
- •We propose GraphForge, a framework that grounds both the task and its verification in real files\. Starting from occupation\-grounded seeds, task statements and rubrics are compiled from an evidence graph over real files, and each task is validated through an initial rollout before trajectory collection\.
- •Training Qwen3\.6\-27B and Qwen3\.6\-35B\-A3B on GraphForge data yields large gains on GDPVal, Workspace\-Bench\-Lite, and SpreadsheetBench II, and our analysis suggests that rubric\-guided selection adds signal beyond training on the model’s own rollouts\.
- •We release the synthesized data and trained models to support future research on working agents\.
Figure 1:Overview of the GraphForge pipeline\.Starting from an O\*NET\-derived seed, an agent assembles a workspace of real files with hidden roles, and a model builds an evidence graph that is compiled into a task specification whose rubric criteria are anchored to graph nodes\. An initial teacher rollout supports a one\-step revision of the task specification, and the final trajectory is scored by an evidence\-anchored judge before admission\.
## 2Related Work
Agent task and environment synthesis\.Recent work synthesizes tasks and environments for agent training across several domains, including general tool use\([Dong et al\., 2026](https://arxiv.org/html/2609.38923#bib.bib3);[Wang et al\., 2026](https://arxiv.org/html/2609.38923#bib.bib4);[Shi et al\., 2025](https://arxiv.org/html/2609.38923#bib.bib5)\), computer use\([Xie et al\., 2026](https://arxiv.org/html/2609.38923#bib.bib6)\), software engineering\([Yang et al\., 2025](https://arxiv.org/html/2609.38923#bib.bib7);[Jain et al\., 2025](https://arxiv.org/html/2609.38923#bib.bib8)\), and terminal operation\([Fan et al\., 2026](https://arxiv.org/html/2609.38923#bib.bib9);[Chu et al\., 2026](https://arxiv.org/html/2609.38923#bib.bib10);[Hua et al\., 2026](https://arxiv.org/html/2609.38923#bib.bib12);[Shi et al\., 2026](https://arxiv.org/html/2609.38923#bib.bib11);[Wu et al\., 2026](https://arxiv.org/html/2609.38923#bib.bib13);[Raoof et al\., 2026](https://arxiv.org/html/2609.38923#bib.bib14)\)\. In these domains, task outcomes can be verified programmatically\. The most relevant pipelines to our work are EnvCraft\([Zeng et al\., 2026](https://arxiv.org/html/2609.38923#bib.bib1)\)and NexForge\([Zhao et al\., 2026](https://arxiv.org/html/2609.38923#bib.bib2)\), discussed in Section[1](https://arxiv.org/html/2609.38923#S1)\. GraphForge addresses their shared gap by compiling an evidence graph over real crawled files into rubrics with evidence anchors, so that trajectory selection and final evaluation are both grounded in source files\.
Evaluation of working agents\.A number of recent benchmarks evaluate agents on realistic work tasks, each with a different emphasis\. GDPVal\([Patwardhan et al\., 2025](https://arxiv.org/html/2609.38923#bib.bib16)\)covers 1,320 tasks across 44 occupations, grades deliverables with expert\-written rubrics, and reports Elo ratings from pairwise comparisons against human work\. Workspace\-Bench\([Tang et al\., 2026](https://arxiv.org/html/2609.38923#bib.bib15)\)places agents in realistic workspaces with tens of thousands of files and evaluates cross\-file dependency reasoning with fine\-grained rubrics\. APEX\-Agents\([Vidgen et al\., 2026](https://arxiv.org/html/2609.38923#bib.bib17)\)focuses on professional services, with tasks created by investment banking analysts, management consultants, and corporate lawyers inside data\-rich simulated worlds\. Claw\-Eval\([Ye et al\., 2026](https://arxiv.org/html/2609.38923#bib.bib18)\)grades agents with trajectory\-aware evidence, recording execution traces, audit logs, and environment snapshots to score fine\-grained rubric items along completion, safety, and robustness\. SpreadsheetBench II\([Zhu et al\., 2026](https://arxiv.org/html/2609.38923#bib.bib27)\)evaluates spreadsheet agents on end\-to\-end business workflows across generation, debugging, and visualization, with expert\-annotated tasks over multi\-sheet workbooks built from authentic business data\. We evaluate our trained models on GDPVal, Workspace\-Bench, and SpreadsheetBench II, following the official protocol of each benchmark\.
Table 1:Comparison of training\-data synthesis pipelines for agents\.The upper block lists general\-domain pipelines, the middle block lists working\-agent pipelines, and the bottom row shows our pipeline\.DatasetDomainTask environmentVerificationScaleOpen dataGeneral\-Domain PipelinesAgentSynthcomputer usedesktop VMper\-step execution check6K tasks✓TaskCrafttool useweb and document toolsgolden answer36K tasks✓SWE\-smithsoftware eng\.code repositoriesfail\-to\-pass tests50K tasks✓CLI\-Universeterminaldocker environmentsfail\-to\-pass tests6K trajs✗Working\-Agent PipelinesEnvCraftworkingsynthesized workspacesstate\-check scripts20K tasks✓NexForge1workingreal filesnone5\.6K tasks✗Our PipelineGraphForgeworkingreal filesevidence\-anchored agent judge2\.1K tasks✓
1NexForge releases trained models but not the synthesized tasks or trajectories\.
## 3Method
GraphForge turns occupational task seeds into verifiable training trajectories through five stages, from seed construction to trajectory admission\.Our design gives seeds and files separate roles\.Seeds fix the occupational direction, so diversity is controlled at seed selection and rebalanced after workspace materialization\. The concrete task and its verification are derivedonly aftera real workspace has been instantiated, so task requirements and evaluation signals are both grounded in source evidence\.
### 3\.1O\*NET\-grounded task\-form seeds
Our seeds come from the O\*NET database, which provides occupations, their task statements, and a controlled vocabulary of Detailed Work Activities \(DWAs\) with an official task\-to\-DWA mapping\([National Center for O\*NET Development, 2025](https://arxiv.org/html/2609.38923#bib.bib21)\)\. As not all tasks are digitally executable, we keep only those annotated DIGITAL by AI4Work\([Wang and others, 2026](https://arxiv.org/html/2609.38923#bib.bib22)\)\. Each retained task is mapped to its DWA through the official relation, and the DWA serves as our controlled task type\. After filtering, 246 occupations across 16 sectors and 43 sub\-sectors remain, covering 891 DWA task types and 3,419 valid occupation\-task\-type pairs\.
Each seed is a tuple
si=\(oi,ai,wi,pi,ui,eiocc\),s\_\{i\}=\(o\_\{i\},a\_\{i\},w\_\{i\},p\_\{i\},u\_\{i\},e\_\{i\}^\{\\mathrm\{occ\}\}\),whereoio\_\{i\}is an occupation,aia\_\{i\}a DWA task type,wiw\_\{i\}an occupation\-specific work demand,pip\_\{i\}the dominant execution pattern,uiu\_\{i\}the expected input file family, andeiocce\_\{i\}^\{\\mathrm\{occ\}\}the retrieved occupational evidence\. For each candidate pair\(oi,ai\)\(o\_\{i\},a\_\{i\}\), we retrieve professional passages and keep only work demandswiw\_\{i\}directly supported by the cited evidence; unsupported demands are dropped rather than filled to a quota\. A second step assigns each demand a dominant execution patternpip\_\{i\}from a vocabulary of 16, from quantitative modeling and reconciliation to policy design and artifact revision\.
To avoid concentrating the corpus on frequent occupations or generic analysis tasks, we select seeds by marginal coverage over the dimensions\(o,a,p,u\)\(o,a,p,u\)\. Let𝒟\\mathcal\{D\}denote these dimensions andnd\(v\)n\_\{d\}\(v\)the number of already selected seeds with valuevvin dimensiondd\. The gain of a candidatessis
Δ\(s∣𝒮\)=∑d∈𝒟11\+nd\(vd\(s\)\)\.\\Delta\(s\\mid\\mathcal\{S\}\)=\\sum\_\{d\\in\\mathcal\{D\}\}\\frac\{1\}\{1\+n\_\{d\}\(v\_\{d\}\(s\)\)\}\.We greedily pick the candidate with the highest gain, breaking ties by a stable task\-ID hash\. After files are downloaded and validated, the same rule is applied again to the actual input and output families\. Balanced subsets add equal quotas overpp, sector round\-robin overoo, and a cap on repeated normalized demandsww\.
### 3\.2Real\-file workspace construction
A seedsis\_\{i\}is case\-neutral\.Its occupationoio\_\{i\}, work demandwiw\_\{i\}, and execution patternpip\_\{i\}specify who does the work, what demand is addressed, and how it is mainly carried out, but the seed names no company, event, dataset, or result\. A search agent instantiatessis\_\{i\}by finding a coherent public case and retrieving the files needed to do the work, forming a workspaceWi=\{fi1,…,fim\}W\_\{i\}=\\\{f\_\{i1\},\\dots,f\_\{im\}\\\}\. Files are downloaded in native formats, parsed with format\-specific tools, and exact duplicates and invalid files are removed\.
Each retained filef∈Wif\\in W\_\{i\}gets a hidden roleρ\(f\)∈\{core,supporting,confuser,ambient\}\\rho\(f\)\\in\\\{\\text\{core\},\\text\{supporting\},\\text\{confuser\},\\text\{ambient\}\\\}\. Core files drive the main computation or decision, supporting files provide policy or context, confusers are plausible but inapplicable alternatives, and ambient files add realistic redundancy\. These roles guide assembly and are never shown to the working agent\. Files may span organizations, mixing related public evidence with same\-domain distractors as long as the task stays coherent and answerable\.
### 3\.3Evidence graphs and verifiable rubrics
Given the workspaceWiW\_\{i\}, GraphForge builds an evidence graphGi=\(Vi,Ei\)G\_\{i\}=\(V\_\{i\},E\_\{i\}\)\. A nodev∈Viv\\in V\_\{i\}records a source file, the fact or field it provides, and its role in the task\. An edgee∈Eie\\in E\_\{i\}records a cross\-file dependency needed to interpret, compare, reconcile, or derive information\. The graph is not ground truth but an intermediate representation whose claimsmust remain recoverable from the original files\.
The graph is compiled into a task specification
𝒞i=\(qi,𝒟i,ℛi\+,ℛi−,Gi\),\\mathcal\{C\}\_\{i\}=\(q\_\{i\},\\mathcal\{D\}\_\{i\},\\mathcal\{R\}\_\{i\}^\{\+\},\\mathcal\{R\}\_\{i\}^\{\-\};G\_\{i\}\),containing a natural task statementqiq\_\{i\}, deliverable requirements𝒟i\\mathcal\{D\}\_\{i\}, positive criteriaℛi\+\\mathcal\{R\}\_\{i\}^\{\+\}, and negative penalty criteriaℛi−\\mathcal\{R\}\_\{i\}^\{\-\}\. Each positive criterion
ck=\(dk,rk,zk,wk,Ak,ϕk\),Ak⊆Vi,c\_\{k\}=\(d\_\{k\},r\_\{k\},z\_\{k\},w\_\{k\},A\_\{k\},\\phi\_\{k\}\),\\qquad A\_\{k\}\\subseteq V\_\{i\},specifies the target deliverabledkd\_\{k\}, the requirementrkr\_\{k\}, the expected value or computationzkz\_\{k\}, a weightwkw\_\{k\}, evidence anchorsAkA\_\{k\}, and a verification procedureϕk\\phi\_\{k\}\. Negative criteria describe concrete prohibited outcomes and are penalized only when the violation is directly evidenced\.
This design separates execution from verification\. The working agent sees onlyqiq\_\{i\}andWiW\_\{i\}, not node IDs or hidden rolesρ\(f\)\\rho\(f\)\. The judge receives the anchorsAkA\_\{k\}and verification instructionsϕk\\phi\_\{k\}, telling it which files and deliverable parts to inspect\.Anchors thus guide both rubric generation and judgingwithout leaking a solution procedure into the task statement\.
### 3\.4Execution\-conditioned one\-step revision
Static inspection cannot catch every ambiguity in a long\-horizon task\. We therefore run each initial task specification𝒞i0\\mathcal\{C\}\_\{i\}^\{0\}once with a strong teacher modelπT\\pi\_\{T\}, producing an initial trajectoryτi0\\tau\_\{i\}^\{0\}and its deliverables\. A revision agent then receives the original filesWiW\_\{i\}, the evidence graphGiG\_\{i\}, the full task and rubrics, and this execution, and checks whether the task is natural and executable, whether required quantities are supported, whether every criterion is correctly anchored, and whether the verification instructions suffice to inspect the artifacts\.
The revision agent returns only the components that need to change\. The compiler keeps all untouched fields, validates references and schema constraints, and emits a revised specification𝒞i1\\mathcal\{C\}\_\{i\}^\{1\}\. We rerun the teacher only when the task statement changes or the initial trajectory is missing\. If only rubric bindings change, the initial execution is reused and judged against the revised specification\. Formally, the trajectory retained for taskiiis
τi=\{τi0,H\(qi1\)=H\(qi0\)andτi0exists,πT\(Wi,qi1\),otherwise,\\tau\_\{i\}=\\begin\{cases\}\\tau\_\{i\}^\{0\},&H\(q\_\{i\}^\{1\}\)=H\(q\_\{i\}^\{0\}\)\\ \\text\{and\}\\ \\tau\_\{i\}^\{0\}\\ \\text\{exists\},\\\\\[2\.0pt\] \\pi\_\{T\}\(W\_\{i\},q\_\{i\}^\{1\}\),&\\text\{otherwise\},\\end\{cases\}whereH\(⋅\)H\(\\cdot\)denotes the normalized hash of the task statement\. The first branch applies exactly when the revision leaves the task statement unchanged and a reusable initial trajectory is available\.
### 3\.5Artifact\-level admission and trajectory cleaning
A trajectoryτi\\tau\_\{i\}is admitted only after its promised deliverables are materialized\. Deterministic checks verify required filenames, readable formats, required sheets and formulas in spreadsheets, and task\-specific structural constraints\. The agent judge then scores every criterion while consulting the referenced source files and produced artifacts\. Each positive criterion contributes its weight times the fraction of the requirement met, and each negative criterion a penalty proportional to the evidenced degree of violation:
Qi\(τ\)=∑k∈ℛi\+wik\+aik\(τ\)−∑j∈ℛi−λijvij\(τ\)∑k∈ℛi\+wik\+,Q\_\{i\}\(\\tau\)=\\frac\{\\sum\_\{k\\in\\mathcal\{R\}\_\{i\}^\{\+\}\}w^\{\+\}\_\{ik\}\\,a\_\{ik\}\(\\tau\)\-\\sum\_\{j\\in\\mathcal\{R\}\_\{i\}^\{\-\}\}\\lambda\_\{ij\}\\,v\_\{ij\}\(\\tau\)\}\{\\sum\_\{k\\in\\mathcal\{R\}\_\{i\}^\{\+\}\}w^\{\+\}\_\{ik\}\},wherewik\+\>0w^\{\+\}\_\{ik\}\>0is the weight of a positive criterion,aik∈\[0,1\]a\_\{ik\}\\in\[0,1\]the fraction of the requirement met,λij\>0\\lambda\_\{ij\}\>0the penalty strength of a negative criterion, andvij∈\[0,1\]v\_\{ij\}\\in\[0,1\]the evidenced degree of violation, withvij=0v\_\{ij\}=0when no violation is found\. At this admission stage, violations are binary, sovij∈\{0,1\}v\_\{ij\}\\in\\\{0,1\\\}\. Scores are used raw and may fall below zero\. We further discard trajectories with degenerate tool\-use behavior, such as repeated non\-polling calls, excessive tool use, high tool\-failure rates, and repeated truncation\.
## 4Main Results
We train Qwen3\.6\-27B and Qwen3\.6\-35B\-A3B on GraphForge data and evaluate the resulting models on GDPVal\-AA\([Patwardhan et al\., 2025](https://arxiv.org/html/2609.38923#bib.bib16)\), the 220\-task gold subset of the full GDPVal benchmark, Workspace\-Bench\-Lite\([Tang et al\., 2026](https://arxiv.org/html/2609.38923#bib.bib15)\), and SpreadsheetBench II\([Zhu et al\., 2026](https://arxiv.org/html/2609.38923#bib.bib27)\)\. The first stage is SFT on admitted teacher trajectories\. The second stage is RFT on rubric\-selected trajectories generated by the SFT model itself\.
### 4\.1Training data
The SFT corpus contains 2,169 admitted trajectories withQi\(τi\)\>0\.90Q\_\{i\}\(\\tau\_\{i\}\)\>0\.90, selected from 3,638 materialized tasks through the construction funnel in Table[4](https://arxiv.org/html/2609.38923#S4.T4)\(b\)\. The corpus covers 466 distinct O\*NET task types, 15 of the 16 occupational sectors, and all 16 execution patterns\.Tasks invented freely by a model tend to collapse toward frequent occupations and generic task types\.Our seeds fix the occupation, task type, execution pattern, and input family before any file is retrieved, and coverage\-based selection keeps the corpus broad\. Figure[2\(a\)](https://arxiv.org/html/2609.38923#S4.F2.sf1)shows the joint coverage of sectors and patterns, with 164 of the 256 sector–pattern combinations realized\. The distribution is not uniform\. Data analysis and reporting accounts for 14\.2%, research and source synthesis for 13\.2%, and no other pattern exceeds 9%\. This shape comes from occupational demand and admission filtering\. Figure[2\(b\)](https://arxiv.org/html/2609.38923#S4.F2.sf2)shows the materialized input file families\. PDF appears in 96\.7% of the trajectories, and most tasks draw on several file families\.
Figure[3](https://arxiv.org/html/2609.38923#S4.F3)reports trajectory length\. A trajectory contains 50\.0 assistant steps on average \(median 48, 95th percentile 82\) and 162\.0k tokens on average \(median 158\.0k, 95th percentile 224\.8k\)\. 28 sequences \(1\.3%\) reach the 262,144\-token training ceiling\.
\(a\)Joint coverage of occupational sectors and execution patterns\. Cell color gives the number of trajectories on a log scale, and gray cells are unobserved combinations\.
\(b\)Materialized input file families\. One task can use several file families\.
Figure 2:Diversity of the SFT corpus across occupational sectors, execution patterns, and input file families\.Figure 3:Distribution of assistant steps and total tokenized length in the 2,169\-example SFT corpus\.
### 4\.2Training and evaluation setup
Training\.We useGLM\-5\.2for all components of the GraphForge pipeline, including the workspace construction agent, the evidence graph and task specification generation, the revision agent, the teacher rollouts, and the evidence\-anchored judge\. SFT trains on the admitted teacher trajectories, with the same corpus and optimization recipe for the 27B and 35B variants\. Full configurations are in Appendix[A](https://arxiv.org/html/2609.38923#A1)\. For RFT data selection, we sampleK=4K=4rollouts per query from the SFT model on 2,000 queries, 125 per execution pattern\. The evidence\-anchored judge scores the four candidates of one query jointly against the rubrics of the task specification\. We keep the highest\-scoring valid trajectory when its score exceeds 0\.95, apply the behavior filter, and drop queries where all candidates fail or the trajectory is overlength, giving 462 trajectories\. The three RFT arms in Table[3](https://arxiv.org/html/2609.38923#S4.T3)share the same 462 query IDs, candidate pools, and optimization budget, and differ only in the selection rule\.
Evaluation\.We evaluate on three working agent benchmarks, GDPVal\-AA\([Patwardhan et al\., 2025](https://arxiv.org/html/2609.38923#bib.bib16)\), Workspace\-Bench\-Lite\([Tang et al\., 2026](https://arxiv.org/html/2609.38923#bib.bib15)\)and SpreadsheetBench II\([Zhu et al\., 2026](https://arxiv.org/html/2609.38923#bib.bib27)\)\. Every model runs under pass@1 with a fixed workspace interface\. Each scaffold \(OpenHands, Codex, Claude Code\) uses a fixed configuration, and comparisons between models are always made within the same scaffold\. A missing or invalid deliverable, or a failure of the model to complete the task counts as a loss\. Model\-side timeouts and infrastructure failures are retried\.
For GDPVal\-AA we maintain an internal Elo pool\. Each \(model, scaffold\) pair is a separate node, and we fit a Bradley–Terry model over the connected comparison graph\([Bradley and Terry, 1952](https://arxiv.org/html/2609.38923#bib.bib24);[Chiang et al\., 2024](https://arxiv.org/html/2609.38923#bib.bib26)\)\. Cross\-scaffold bridge comparisons place the OpenHands and Codex nodes on a common scale\. A tie contributes one half\-win and one half\-loss\. Each nodeiihas a strength parameterθi\\theta\_\{i\}, and
Pr\(i≻j\)=σ\(θi−θj\),Eloi=1667\+400ln10\(θi−θanchor\)\.\\Pr\(i\\succ j\)=\\sigma\(\\theta\_\{i\}\-\\theta\_\{j\}\),\\qquad\\mathrm\{Elo\}\_\{i\}=1667\+\\frac\{400\}\{\\ln 10\}\\,\(\\theta\_\{i\}\-\\theta\_\{\\mathrm\{anchor\}\}\)\.\(1\)The model is invariant to additive shifts of the strengths, so we anchor the scale by fixing the Elo of GLM\-5\.3 \(OpenHands\) to 1667\. We report scores on the conventional Elo scale\([Elo, 2008](https://arxiv.org/html/2609.38923#bib.bib23);[Boubdir et al\., 2023](https://arxiv.org/html/2609.38923#bib.bib25)\), in which a 400\-point gap corresponds to ten\-to\-one odds\. The factor400/ln10400/\\ln 10converts the fitted strengths to this scale\.
### 4\.3Overall comparison
Table 2:Overall comparison on working agent benchmarks\.GDPVal reports Elo under the OpenHands and Codex scaffolds, with each \(model, scaffold\) pair fitted as a separate node and the scale anchored at GLM\-5\.3 \(OpenHands\) = 1667\. Workspace\-Bench\-Lite reports micro scores and SpreadsheetBench II reports execution accuracy, both under the Claude Code and Codex scaffolds\. GDPVal bootstrap confidence intervals are reported in Appendix[C](https://arxiv.org/html/2609.38923#A3)\.ModelGDPValWorkspace\-Bench\-LiteSpreadsheetBench IIOpenHandsCodexClaude CodeCodexClaude CodeCodexFrontier ModelsClaude Opus 51774\.11753\.170\.168\.933\.6—GPT\-5\.6\-sol1687\.11710\.8—60\.5—32\.7Qwen3\.8\-Max1719\.01771\.067\.466\.634\.934\.9GLM\-5\.31667\.01543\.767\.761\.432\.131\.5Kimi\-K31615\.51664\.465\.860\.635\.837\.7DeepSeek\-V4\-Pro1531\.51576\.758\.157\.929\.335\.5Open\-Weight BaselineNex\-N2\-Mini\-35B1288\.81342\.333\.131\.66\.510\.3Our ModelsQwen3\.6\-35B\-A3B1260\.61283\.055\.953\.42\.84\.7\+ SFT \(GraphForge\)1362\.3\(\+101\.7\)1384\.4\(\+101\.4\)59\.7\(\+3\.8\)60\.0\(\+6\.6\)19\.3\(\+16\.5\)18\.7\(\+14\.0\)Qwen3\.6\-27B1380\.01364\.056\.061\.410\.315\.6\+ SFT \(GraphForge\)1445\.7\(\+65\.7\)1427\.4\(\+63\.4\)63\.7\(\+7\.7\)65\.2\(\+3\.8\)24\.0\(\+13\.7\)24\.6\(\+9\.0\)
Table[2](https://arxiv.org/html/2609.38923#S4.T2)compares GraphForge with frontier models and with Nex\-N2\-Mini\-35B, the open\-weight model released with NexForge\([Zhao et al\., 2026](https://arxiv.org/html/2609.38923#bib.bib2)\)\. Supervised training on GraphForge data produces large gains on all three benchmarks\. On GDPVal, SFT improves the 35B base model by 101\.7 Elo on OpenHands and 101\.4 Elo on Codex\. The 27B SFT model reaches 1445\.7 Elo on OpenHands and 1427\.4 Elo on Codex, improving over its base model by 65\.7 and 63\.4 points\. On Workspace\-Bench\-Lite, SFT improves the two base models by up to 6\.6 and 7\.7 points\. On SpreadsheetBench II, the gains reach 16\.5 and 13\.7 points\.
The gain also transfers across agent scaffolds\. All GraphForge trajectories are rolled out with the Codex scaffold, while the evaluation covers OpenHands and Codex on GDPVal and Claude Code and Codex on Workspace\-Bench\-Lite and SpreadsheetBench II\. The SFT model improves over the base model under every scaffold\. This suggests that the corpus teaches working skills that transfer across scaffolds, rather than habits tied to the rollout scaffold\.
### 4\.4Rubric\-guided rejection fine\-tuning
Table 3:RFT ablation on top of the SFT model\.Parentheses report the change from the SFT model\. GDPVal values are SFT\-anchored Elo from direct paired comparisons with SFT\. GDPVal bootstrap confidence intervals are reported in Appendix[D](https://arxiv.org/html/2609.38923#A4)\.ModelGDPValWorkspace\-Bench\-LiteSpreadsheetBench IIOpenHandsCodexClaude CodeCodexClaude CodeCodexOur ModelsQwen3\.6\-35B\-A3B \(reference\)1260\.61283\.055\.953\.42\.84\.735B SFT \(GraphForge\)1362\.31384\.459\.760\.019\.318\.7RFT Variants\+ RFT1369\.5\(\+7\.2\)1395\.4\(\+11\.0\)63\.7\(\+4\.0\)64\.0\(\+4\.0\)20\.3\(\+1\.0\)19\.6\(\+0\.9\)\+ RFT \(unanchored\)1396\.4\(\+34\.1\)1409\.7\(\+25\.3\)62\.2\(\+2\.5\)61\.7\(\+1\.7\)17\.5\(\-1\.8\)18\.7\(0\.0\)\+ RFT \(random\-of\-4\)1353\.3\(\-9\.0\)1374\.9\(\-9\.5\)60\.8\(\+1\.1\)62\.8\(\+2\.8\)19\.0\(\-0\.3\)15\.6\(\-3\.1\)
We compare three offline rejection fine\-tuning arms initialized from the same SFT checkpoint\. From 2,000 newly synthesized queries, we sample up to four trajectories per query with the SFT model, forming a shared candidate pool\. The anchored arm keeps, for each query, the rubric\-best trajectory when its judge score exceeds 0\.95\. After validity and behavior filtering, 462 queries remain, each contributing one trajectory\. The other two arms reuse the same 462 queries and their candidate pools\. The unanchored arm ranks the candidates with a judge that sees the rubric text but not the explicit evidence anchors and verification instructions\. The random\-of\-4 arm selects uniformly at random from the eligible candidates\. All arms share the candidate eligibility rules, the 462 training examples, and the optimization budget, and each arm branches independently from the same checkpoint\. This matched design isolates within\-query trajectory selection rather than the full task\-admission pipeline\.
For scoring, each arm is compared directly against the SFT model under the same scaffold on the same tasks\. We convert the resulting win rate into an Elo difference and add it to the frozen main\-table SFT score, so all arms are reported on the same scale as Table[2](https://arxiv.org/html/2609.38923#S4.T2)\. These paired comparisons are independent of the joint pool used for the main table\.
Table[3](https://arxiv.org/html/2609.38923#S4.T3)reports the results\. On Workspace\-Bench\-Lite and SpreadsheetBench II, the ordering follows the design intent\. Anchored selection gives the largest gains over SFT, unanchored selection gives smaller or negative gains, and random selection is the weakest overall\. On GDPVal, anchored RFT improves over SFT \(\+7\.2 and \+11\.0 Elo\) and random selection degrades performance \(\-9\.0 and \-9\.5\), while the unanchored arm attains higher point estimates \(\+34\.1 and \+25\.3\)\. However, none of these GDPVal differences is statistically resolved at this scale\. GDPVal Elo differences at 220 tasks are therefore noisy rather than decisive\. We read the results as follows\. Selection quality matters across benchmarks, since random selection is consistently the weakest arm\. The advantage of evidence anchoring is reflected on Workspace\-Bench\-Lite and SpreadsheetBench II, while GDPVal Elo is too noisy to separate the two judge variants\. RFT is compatible with continued improvement and does not damage the SFT model\.
### 4\.5Ablation Studies
Contamination and transfer\.We audit all 2,150 training workspaces against the 220 GDPVal tasks at three levels of granularity \(Table[4](https://arxiv.org/html/2609.38923#S4.T4)a\), since GDPVal draws its tasks from the same O\*NET taxonomy as our seeds and poses the highest overlap risk\. At the file level, none of the 39,201 training files coincides with any of the 260 GDPVal files\. At the text level, the top\-20 most similar 13\-gram pairs between the two corpora contain no substantive shared content\. At the occupation level, only 13 of the 44 GDPVal occupations are covered by our training taxonomy, so most evaluation tasks are occupation\-disjoint from the training data\.
To test whether the SFT gain is concentrated near covered content, we split GDPVal tasks by occupation coverage and compare Base vs\. SFT win rates \(Table[4](https://arxiv.org/html/2609.38923#S4.T4)c\)\. SFT improves over the base model on both groups, and the win rate on the 155 uncovered tasks \(0\.739, 95% CI \[0\.671, 0\.803\]\) is no lower than on the 65 covered tasks \(0\.692, 95% CI \[0\.585, 0\.800\]\)\. Together with the gains on Workspace\-Bench\-Lite and SpreadsheetBench II in Table[2](https://arxiv.org/html/2609.38923#S4.T2), this indicates thatthe improvement reflects transferable working skills rather than memorization of benchmark content\. Full details are in Appendix[B](https://arxiv.org/html/2609.38923#A2)\.
Judge sensitivity\.We probe whether the evidence\-anchored judge grounds its scores in the referenced files \(Table[4](https://arxiv.org/html/2609.38923#S4.T4)d\)\. Deleting the worksheet that a criterion cites drops the corresponding score by 0\.377 on average, and non\-target criteria remain essentially unchanged \(mean\|ΔQ\|=0\.016\|\\Delta Q\|=0\.016\)\. This suggests that the judge reads the cited evidence and that its response is localized to the affected criterion\. In contrast, fine\-grained corruptions of rows, numbers, and citations cause only small changes \(below 0\.03 in magnitude\)\. Since the corrupted cells are part of the cited evidence, an ideal judge should catch these perturbations as well\.We attribute this gap to the capability limit of GLM\-5\.2 as an agentic judge, which reliably detects structural evidence failures but struggles to verify fine\-grained content\.
Table 4:Audits of the training corpus and the judge\.\(a\) Train–test overlap at the file, text, and occupation level\. \(b\) Corpus construction funnel\. \(c\) SFT transfer on occupation\-covered and occupation\-uncovered GDPVal tasks, with task\-bootstrap confidence intervals\. \(d\) Controlled judge perturbations\.\(a\) Overlap audit LevelComparisonResultFile39,201 train vs\. 260 GDPVal files0 sharedTextTop\-20 13\-gram pairs0 substantiveOccupationGDPVal occupations covered13/44
\(c\) Grouped transfer, Base vs\. SFT GroupW/T/LWin rate95% CICovered \(65\)43/4/180\.692\[0\.585, 0\.800\]Uncovered \(155\)109/11/350\.739\[0\.671, 0\.803\]All \(220\)152/15/530\.725\[0\.668, 0\.782\]
\(b\) Construction funnel StageCountShareMaterialized tasks3,638100\.0%Unchanged rollouts reused2,96781\.6%Tasks admitted atQ\>0\.90Q\>0\.902,15359\.2%Validated trajectories12,16959\.6%Unique workspaces2,15059\.1%
1A task can contribute multiple trajectories when a rollout is compacted into separate training sequences\.
\(d\) Controlled judge sensitivity PerturbationTarget\-criterionΔQ\\Delta QDeleted worksheet−0\.377\-0\.377Row corruption−0\.018\-0\.018Numeric corruption−0\.022\-0\.022Citation corruption−0\.013\-0\.013Non\-target criteria, mean\|ΔQ\|\|\\Delta Q\|:0\.0160\.016
## 5Conclusion
We have presented GraphForge, an evidence\-graph based framework that synthesizes working agent training data from real files\. In GraphForge, occupational seeds fix the task direction, and an evidence graph over the instantiated workspace supplies both the task and its verification, so task requirements are backed by source files and each criterion is anchored to the files needed to verify it\. Training Qwen3\.6\-27B on 2,169 synthesized trajectories brings GDPVal to 1445\.7 \(\+65\.7\) under OpenHands, and Workspace\-Bench\-Lite and SpreadsheetBench II to 63\.7 \(\+7\.7\) and 24\.0 \(\+13\.7\) under Claude Code, with gains holding across the OpenHands, Codex, and Claude Code scaffolds\. The same data also improves Qwen3\.6\-35B\-A3B, suggesting that GraphForge trajectories generalize across base models\. Further analysis with rubric\-guided rejection fine\-tuning yields additional gains over SFT and supports the value of evidence\-anchored selection\. We hope the released models make it easier to build working agents that operate faithfully on real files\.
## 6Discussion of Limitations
Our study has several limitations\. First, our current corpus contains 2,169 trajectories, and we have not studied how the benefits of GraphForge scale with larger data budgets\. Second, both the agent judge and the synthesis pipeline are powered by GLM\-5\.2\. Stronger frontier models could improve the quality of the synthesized data and the reliability of the judging, and exploring the ceiling of our framework with such models remains future work\. Third, our experiments cover two base models from the same family, and we do not study how GraphForge transfers to other model families\. Looking ahead, we plan to scale GraphForge to more task families and file types, and to study how evidence\-anchored verification interacts with longer\-horizon agent scaffolds\.
### AI use statement
We used generative AI tools to polish the English writing and to assist with coding tasks such as debugging\. We did not use generative AI tools to design the method, run experiments, or write the scientific claims\. We reviewed all AI\-assisted content\. LLM\-polished text was checked by the authors for accuracy, and LLM\-generated code was verified and tested by the authors\. We take full responsibility for the final content of this work\.
### Ethics statement
This work does not involve human subjects, so no IRB approval is required\. Training workspaces are assembled from publicly available documents, such as corporate filings and public reports, and are used for research purposes only\. Exact duplicates and invalid files are removed during collection\. We release the synthesized tasks, rubrics, trajectories, and trained checkpoints\. Workspace files are public documents, and we provide their source links rather than redistributing file contents\. Public documents may mention individuals in their original context, and we do not collect, curate, or infer any personally identifiable information beyond what already appears in these public sources\. We do not foresee harmful applications of our method\. We have no conflicts of interest to disclose\.
### Reproducibility statement
The method details including architecture, training procedure, and hyperparameters are given\. The synthesized data and trained checkpoints are released at[https://huggingface\.co/collections/groundhogLLM/graphforge](https://huggingface.co/collections/groundhogLLM/graphforge)\. All experiments use fixed random seeds, and the hardware and software setup is reported\.
## References
- Boubdiret al\.\(2023\)M\. Boubdir, E\. Kim, B\. Ermis, S\. Hooker, and M\. FadaeeElo uncovered: robustness and best practices in language model evaluation\.External Links:2311\.17295,[Link](https://arxiv.org/abs/2311.17295)Cited by:[§4\.2](https://arxiv.org/html/2609.38923#S4.SS2.p3.2)\.
- Bradley and Terry \(1952\)R\. A\. Bradley and M\. E\. TerryRank analysis of incomplete block designs: I\. the method of paired comparisons\.Biometrika39\(3/4\),pp\. 324–345\.Cited by:[§4\.2](https://arxiv.org/html/2609.38923#S4.SS2.p3.1)\.
- Chianget al\.\(2024\)W\. Chiang, L\. Zheng, Y\. Sheng, A\. N\. Angelopoulos, T\. Li, D\. Li, H\. Zhang, B\. Zhu, M\. Jordan, J\. E\. Gonzalez, and I\. StoicaChatbot arena: an open platform for evaluating llms by human preference\.External Links:2403\.04132,[Link](https://arxiv.org/abs/2403.04132)Cited by:[§4\.2](https://arxiv.org/html/2609.38923#S4.SS2.p3.1)\.
- Chuet al\.\(2026\)Z\. Chu, J\. Hu, X\. Jiang, P\. Zou, H\. Li, C\. Peng, P\. O’Hearn, E\. T\. Barr, M\. Harman, F\. Sarro, and H\. YeTerminalWorld: benchmarking agents on real\-world terminal tasks\.External Links:2605\.22535,[Link](https://arxiv.org/abs/2605.22535)Cited by:[§1](https://arxiv.org/html/2609.38923#S1.p1.1),[§2](https://arxiv.org/html/2609.38923#S2.p1.1)\.
- Donget al\.\(2026\)G\. Dong, J\. Lu, J\. Huang, W\. Zhong, L\. Liu, S\. Huang, Z\. Li, Y\. Zhao, X\. Song, X\. Li, J\. Jin, Y\. Zhu, H\. Wang, F\. Lei, Q\. Luo, M\. Chen, Z\. Chen, J\. Feng, J\. Wen, and Z\. DouAgent\-world: scaling real\-world environment synthesis for evolving general agent intelligence\.External Links:2604\.18292,[Link](https://arxiv.org/abs/2604.18292)Cited by:[§1](https://arxiv.org/html/2609.38923#S1.p1.1),[§2](https://arxiv.org/html/2609.38923#S2.p1.1)\.
- Elo \(2008\)A\. E\. EloThe rating of chessplayers, past and present\.Bronx, NY : Ishi Press International\.Cited by:[§4\.2](https://arxiv.org/html/2609.38923#S4.SS2.p3.2)\.
- Fanet al\.\(2026\)Z\. Fan, T\. Yu, Y\. Cai, J\. Guan, Y\. Yang, D\. Hu, J\. Zhou, X\. Wu, Z\. Han, F\. Zhang, and L\. WangToward scalable terminal task synthesis via skill graphs\.External Links:2604\.25727,[Link](https://arxiv.org/abs/2604.25727)Cited by:[§2](https://arxiv.org/html/2609.38923#S2.p1.1)\.
- Hermes\-Agent Team \(2026\)Hermes\-Agent TeamHermes\-agent\.Note:[https://github\.com/nousresearch/hermes\-agent](https://github.com/nousresearch/hermes-agent)Cited by:[§1](https://arxiv.org/html/2609.38923#S1.p1.1)\.
- Huaet al\.\(2026\)Z\. Hua, Y\. Yao, W\. Xie, Y\. Zhao, M\. Liu, R\. Qiu, Z\. Huang, Z\. Wang, Y\. Ji, Y\. Ye, L\. Zhu, X\. Lei, H\. Li, Z\. Ma, Z\. Wang, Z\. Zhang, and J\. LiuCLI\-universe: towards verifiable task synthesis engine for terminal agents\.External Links:2606\.22883,[Link](https://arxiv.org/abs/2606.22883)Cited by:[§2](https://arxiv.org/html/2609.38923#S2.p1.1)\.
- Jainet al\.\(2025\)N\. Jain, J\. Singh, M\. Shetty, L\. Zheng, K\. Sen, and I\. StoicaR2E\-gym: procedural environments and hybrid verifiers for scaling open\-weights swe agents\.External Links:2504\.07164,[Link](https://arxiv.org/abs/2504.07164)Cited by:[§2](https://arxiv.org/html/2609.38923#S2.p1.1)\.
- National Center for O\*NET Development \(2025\)National Center for O\*NET DevelopmentO\*NET 30\.3 Database\.Note:[https://www\.onetcenter\.org/database\.html](https://www.onetcenter.org/database.html)Cited by:[§3\.1](https://arxiv.org/html/2609.38923#S3.SS1.p1.1)\.
- OpenClaw Team \(2026\)OpenClaw TeamOpenClaw\.Note:[https://github\.com/openclaw/openclaw](https://github.com/openclaw/openclaw)Cited by:[§1](https://arxiv.org/html/2609.38923#S1.p1.1)\.
- Patwardhanet al\.\(2025\)T\. Patwardhan, R\. Dias, E\. Proehl, G\. Kim, M\. Wang, O\. Watkins, S\. P\. Fishman, M\. Aljubeh, P\. Thacker, L\. Fauconnet, N\. S\. Kim, P\. Chao, S\. Miserendino, G\. Chabot, D\. Li, M\. Sharman, A\. Barr, A\. Glaese, and J\. TworekGDPval: evaluating ai model performance on real\-world economically valuable tasks\.External Links:2510\.04374,[Link](https://arxiv.org/abs/2510.04374)Cited by:[§1](https://arxiv.org/html/2609.38923#S1.p1.1),[§2](https://arxiv.org/html/2609.38923#S2.p2.1),[§4\.2](https://arxiv.org/html/2609.38923#S4.SS2.p2.1),[§4](https://arxiv.org/html/2609.38923#S4.p1.1)\.
- Raoofet al\.\(2026\)N\. Raoof, R\. Zhuang, M\. Nezhurina, E\. Guha, A\. Tejaswi, R\. Marten, C\. F\. Ruan, T\. Griggs, A\. G\. Shaw, H\. Bansal, E\. K\. Buchanan, A\. Gazizov, R\. Heckel, C\. Hegde, S\. Jajee, D\. Khazi, E\. Koukoumidis, X\. Li, H\. Liu, S\. Natarajan, H\. Raj, N\. Roberts, E\. Shen, N\. Singhi, M\. Siu, A\. Suvarna, H\. Xing, P\. Yubeaton, R\. Zhang, L\. L\. Chen, X\. Chen, S\. Dillmann, S\. Gabriel, X\. Jiang, A\. Kashyap, B\. Li, Y\. Park, M\. Pham, S\. Sanghavi, L\. Shi, K\. Sun, Y\. Wang, Z\. Xu, E\. Zhang, S\. Zhao, W\. Zhao, J\. Jitsev, A\. Dimakis, B\. Feuer, and L\. SchmidtOpenThoughts\-agent: data recipes for agentic models\.External Links:2606\.24855,[Link](https://arxiv.org/abs/2606.24855)Cited by:[§2](https://arxiv.org/html/2609.38923#S2.p1.1)\.
- Shiet al\.\(2025\)D\. Shi, J\. Cao, Q\. Chen, W\. Sun, W\. Li, H\. Lu, F\. Dong, T\. Qin, K\. Zhu, M\. Liu, J\. Yang, G\. Zhang, J\. Liu, C\. Zhang, J\. Wang, Y\. E\. Jiang, and W\. ZhouTaskCraft: automated generation of agentic tasks\.External Links:2506\.10055,[Link](https://arxiv.org/abs/2506.10055)Cited by:[§2](https://arxiv.org/html/2609.38923#S2.p1.1)\.
- Shiet al\.\(2026\)K\. Shi, Z\. Wang, Q\. Su, S\. Huang, Z\. Zhang, Z\. Fang, Q\. Ren, J\. Liu, Y\. Zeng, Y\. Zhao, L\. Chen, Z\. Chen, and F\. ZhaoFACET: preserving source intent and executable state in terminal task synthesis\.External Links:2608\.18580,[Link](https://arxiv.org/abs/2608.18580)Cited by:[§1](https://arxiv.org/html/2609.38923#S1.p1.1),[§2](https://arxiv.org/html/2609.38923#S2.p1.1)\.
- Tanget al\.\(2026\)Z\. Tang, X\. Zhou, Y\. Liu, L\. Li, Y\. Wu, W\. Wang, H\. Huang, W\. Zhou, J\. Zhou, J\. Song, S\. Yu, J\. Wang, Z\. Zhou, H\. Zhou, Y\. Lv, J\. Li, J\. Liu, R\. Chen, C\. Liu, G\. Li, J\. Kang, and F\. WuWorkspace\-bench 1\.0: benchmarking ai agents on workspace tasks with large\-scale file dependencies\.External Links:2605\.03596,[Link](https://arxiv.org/abs/2605.03596)Cited by:[§1](https://arxiv.org/html/2609.38923#S1.p1.1),[§2](https://arxiv.org/html/2609.38923#S2.p2.1),[§4\.2](https://arxiv.org/html/2609.38923#S4.SS2.p2.1),[§4](https://arxiv.org/html/2609.38923#S4.p1.1)\.
- Vidgenet al\.\(2026\)B\. Vidgen, A\. Mann, A\. Fennelly, J\. W\. Stanly, L\. Rothman, M\. Burstein, J\. Benchek, D\. Ostrofsky, A\. Ravichandran, D\. Sur, N\. Venugopal, A\. Hsia, I\. Robinson, C\. Huang, O\. Varones, D\. Khan, M\. Haines, A\. Bridges, J\. Boyle, K\. Twist, Z\. Richards, C\. Mahapatra, B\. Foody, and O\. NitskiAPEX\-agents\.External Links:2601\.14242,[Link](https://arxiv.org/abs/2601.14242)Cited by:[§2](https://arxiv.org/html/2609.38923#S2.p2.1)\.
- Wanget al\.\(2026\)Z\. Wang, C\. Xu, B\. Liu, Y\. Wang, S\. Han, Z\. Yao, H\. Yao, and Y\. HeAgent world model: infinity synthetic environments for agentic reinforcement learning\.External Links:2602\.10090,[Link](https://arxiv.org/abs/2602.10090)Cited by:[§2](https://arxiv.org/html/2609.38923#S2.p1.1)\.
- Wanget al\.\(2026\)Z\. Z\. Wanget al\.How well does agent development reflect real\-world work?\.arXiv preprint arXiv:2603\.01203\.Cited by:[§3\.1](https://arxiv.org/html/2609.38923#S3.SS1.p1.1)\.
- Wuet al\.\(2026\)S\. Wu, Y\. Li, Y\. Song, W\. Zhang, Y\. Wang, R\. Batista\-Navarro, X\. Yang, M\. Tang, B\. Dai, J\. Yang, and C\. LinLarge\-scale terminal agentic trajectory generation from dockerized environments\.External Links:2602\.01244,[Link](https://arxiv.org/abs/2602.01244)Cited by:[§2](https://arxiv.org/html/2609.38923#S2.p1.1)\.
- Xieet al\.\(2026\)J\. Xie, D\. Xu, X\. Zhao, and D\. SongAgentSynth: scalable task generation for generalist computer\-use agents\.External Links:2506\.14205,[Link](https://arxiv.org/abs/2506.14205)Cited by:[§2](https://arxiv.org/html/2609.38923#S2.p1.1)\.
- Yanget al\.\(2025\)J\. Yang, K\. Lieret, C\. E\. Jimenez, A\. Wettig, K\. Khandpur, Y\. Zhang, B\. Hui, O\. Press, L\. Schmidt, and D\. YangSWE\-smith: scaling data for software engineering agents\.External Links:2504\.21798,[Link](https://arxiv.org/abs/2504.21798)Cited by:[§2](https://arxiv.org/html/2609.38923#S2.p1.1)\.
- Yeet al\.\(2026\)B\. Ye, R\. Li, Q\. Yang, Y\. Liu, L\. Yao, H\. Lv, Z\. Xie, C\. An, L\. Li, L\. Kong, Q\. Liu, Z\. Sui, and T\. YangClaw\-eval: towards trustworthy evaluation of autonomous agents\.External Links:2604\.06132,[Link](https://arxiv.org/abs/2604.06132)Cited by:[§1](https://arxiv.org/html/2609.38923#S1.p1.1),[§2](https://arxiv.org/html/2609.38923#S2.p2.1)\.
- Zenget al\.\(2026\)Y\. Zeng, S\. You, J\. Feng, Y\. Liu, X\. Ding, Y\. Hou, H\. Cong, Y\. Wang, W\. Ning, W\. Xu, and B\. CaiEnvCraft: synthesizing executable environments in agentic rl for claw\-like agent\.External Links:2609\.05576,[Link](https://arxiv.org/abs/2609.05576)Cited by:[§1](https://arxiv.org/html/2609.38923#S1.p2.1),[§2](https://arxiv.org/html/2609.38923#S2.p1.1)\.
- Zhaoet al\.\(2026\)J\. Zhao, Z\. Lei, Z\. Xi, R\. Zheng, H\. Yan, J\. Zhou, Q\. Chen, and L\. HeNexForge: scaling agent capabilities through requirement\-driven task synthesis for llms\.External Links:2607\.14186,[Link](https://arxiv.org/abs/2607.14186)Cited by:[§1](https://arxiv.org/html/2609.38923#S1.p2.1),[§2](https://arxiv.org/html/2609.38923#S2.p1.1),[§4\.3](https://arxiv.org/html/2609.38923#S4.SS3.p1.1)\.
- Zhuet al\.\(2026\)J\. Zhu, Y\. Zhang, Z\. Ma, B\. Zhang, A\. Schoepf, D\. Woloch, P\. Y\. Wang, G\. R\. Yang, S\. Jacob, S\. Nagisetty, A\. Chundru, J\. Lin, S\. Mateega, and J\. ZhangSpreadsheetBench 2: evaluating agents on end\-to\-end business spreadsheet workflows\.External Links:2606\.29955,[Link](https://arxiv.org/abs/2606.29955)Cited by:[§2](https://arxiv.org/html/2609.38923#S2.p2.1),[§4\.2](https://arxiv.org/html/2609.38923#S4.SS2.p2.1),[§4](https://arxiv.org/html/2609.38923#S4.p1.1)\.
## Appendix ATraining Configurations
Table[5](https://arxiv.org/html/2609.38923#A1.T5)reports the configurations used for the SFT and controlled RFT experiments\. The 27B and 35B SFT models use the same corpus and optimization recipe\. All three RFT arms start from the same 35B SFT checkpoint and differ only in trajectory selection\.
Table 5:SFT and RFT training configurations\.The anchored, unanchored, and random RFT arms use the same 462 query IDs, candidate pools, behavior filter, and optimization budget\.ConfigurationSFTRFT armsInitializationQwen3\.6 base \(27B or 35B\-A3B\)35B\-A3B SFT checkpointTraining examples2,169 trajectories462 matched trajectories per armData admissionone\-step revision,Qi\>0\.90Q\_\{i\}\>0\.90best\-of\-4,Qi\>0\.95Q\_\{i\}\>0\.95, behavior\-cleanEpochs31OptimizerMuonMuonPeak learning rate2×10−52\\times 10^\{\-5\}1×10−61\\times 10^\{\-6\}Minimum learning rate1×10−61\\times 10^\{\-6\}1×10−71\\times 10^\{\-7\}Learning\-rate scheduleCosineCosineWarmup ratio0\.100\.10Weight decay0\.050\.05Global batch size88Maximum \(packed\) length262,144 tokensSequence packingEnabledEnabledAll trajectories are trained with the full assistant reasoning and tool\-interaction history preserved\. Samples whose complete serialized sequence exceeds 262,144 tokens are excluded rather than truncated\. For RFT, the three arms use identical query IDs and training hyperparameters\. Only the rule used to choose one trajectory from each four\-candidate pool changes\.
## Appendix BContamination Audit
File\-level overlap\.The 2,169 training sequences come from 2,150 unique workspaces containing 49,750 file instances and 39,201 unique SHA\-256 hashes\. We compared these hashes against the files actually referenced by the 220 GDPVal tasks, which contain 261 file instances and 260 unique hashes\. No hash is shared between the two corpora\.
Text\-level overlap\.We extracted text from every file in a supported format and added the 220 task prompts\. Extraction succeeded for 39,075 training files and 229 GDPVal reference files\. Files that failed extraction or use non\-text formats still participated in the hash check above\. We retrieved candidate pairs with normalized 13\-gram signatures and manually reviewed the top 20 pairs ranked by containment\. None of them shares task requirements, entities, business facts, or deliverable content\. Five pairs share only generic numeric sequences, and fifteen share only PowerPoint master placeholder text\.
Occupation coverage\.13 of the 44 GDPVal occupations also appear in the training taxonomy\. These occupations account for 65 tasks, while the remaining 155 tasks belong to occupations that the corpus does not cover\. This overlap follows from the shared O\*NET taxonomy rather than from shared files or tasks\.
Grouped comparison\.To check whether the improvement concentrates on covered occupations, we compared the base model and the SFT model directly on all 220 tasks\. Each pair of deliverables was presented to the judge in balanced order, with the base model shown first on 110 tasks and the SFT model shown first on the other 110\. On 212 tasks both models produced a deliverable and the judge decided the outcome\. Two tasks where only the base model produced an empty deliverable count as wins for the SFT model, and six tasks where only the SFT model produced an empty deliverable count as losses\. Table[4](https://arxiv.org/html/2609.38923#S4.T4)reports the results\. The SFT model wins at least as often on uncovered tasks as on covered ones, and the win\-rate difference of 0\.047 has a task\-bootstrap 95% confidence interval of\[−0\.079,0\.175\]\[\-0\.079,0\.175\]\. Together with the file\-level and text\-level audits above, we find no sign that the improvement relies on proximity to benchmark content\.
## Appendix CGDPVal Elo Confidence Intervals
Table[6](https://arxiv.org/html/2609.38923#A3.T6)reports 95% bootstrap confidence intervals for the GDPVal Elo scores in Table[2](https://arxiv.org/html/2609.38923#S4.T2)\. The main\-table pool excludes all RFT arms\. It also retains Claude Opus 4\.8 \(OpenHands\) as a bridge node, which is needed to keep the comparison graph connected and is not reported in Table[2](https://arxiv.org/html/2609.38923#S4.T2)\. Intervals are obtained by resampling the W/T/L outcomes on each comparison edge 10,000 times and refitting the complete Bradley–Terry graph, with GLM\-5\.3 \(OpenHands\) fixed at 1667 in every draw\.
Table 6:GDPVal Elo with 95% bootstrap confidence intervals\.OpenHands and Codex results are separate \(model, scaffold\) nodes\.ModelOpenHands Elo \[95% CI\]Codex Elo \[95% CI\]Claude Opus 51774\.1 \[1750\.2, 1798\.3\]1753\.1 \[1697\.4, 1815\.7\]GPT\-5\.6\-sol1687\.1 \[1663\.4, 1710\.7\]1710\.8 \[1657\.8, 1769\.3\]Qwen3\.8\-Max1719\.0 \[1694\.9, 1741\.8\]1771\.0 \[1714\.6, 1831\.9\]GLM\-5\.31667\.0 \[1667\.0, 1667\.0\]1543\.7 \[1491\.6, 1595\.8\]Kimi\-K31615\.5 \[1592\.6, 1638\.4\]1664\.4 \[1613\.9, 1717\.2\]DeepSeek\-V4\-Pro1531\.5 \[1507\.5, 1555\.4\]1576\.7 \[1527\.0, 1628\.2\]Nex\-N2\-Mini\-35B1288\.8 \[1175\.9, 1378\.9\]1342\.3 \[1296\.9, 1384\.5\]Qwen3\.6\-35B\-A3B1260\.6 \[1195\.1, 1316\.3\]1283\.0 \[1241\.1, 1322\.4\]\+ SFT \(GraphForge\)1362\.3 \[1306\.3, 1413\.6\]1384\.4 \[1345\.8, 1422\.0\]Qwen3\.6\-27B1380\.0 \[1326\.6, 1432\.1\]1364\.0 \[1324\.6, 1401\.4\]\+ SFT \(GraphForge\)1445\.7 \[1376\.2, 1513\.8\]1427\.4 \[1389\.1, 1465\.2\]
## Appendix DRFT Elo Confidence Intervals
Table[7](https://arxiv.org/html/2609.38923#A4.T7)reports conditional 95% bootstrap intervals for the RFT arms in Table[3](https://arxiv.org/html/2609.38923#S4.T3)\. Each scaffold uses a star graph whose three edges compare the RFT arms directly with SFT, and the SFT point estimate from the frozen main table serves as a fixed reporting anchor\. With ties counted as half wins, the Bradley–Terry solution on each edge is400log10\(p/\(1−p\)\)400\\log\_\{10\}\(p/\(1\-p\)\)relative to SFT\. For intervals, we jointly resample task UUIDs across the three edges for 10,000 draws and recompute the scores, so the intervals reflect task\-sampling uncertainty conditional on the fixed SFT anchor\. All difference intervals include zero\.
Table 7:RFT direct\-comparison Elo with conditional 95% bootstrap confidence intervals\.ModelOpenHands Elo \[95% CI\]Codex Elo \[95% CI\]35B SFT1362\.3 \(fixed\)1384\.4 \(fixed\)\+ RFT1369\.5 \[1319\.1, 1416\.5\]1395\.4 \[1349\.5, 1440\.1\]\+ RFT \(unanchored\)1396\.4 \[1348\.0, 1446\.3\]1409\.7 \[1363\.8, 1458\.1\]\+ RFT \(random\-of\-4\)1353\.3 \[1306\.3, 1401\.9\]1374\.9 \[1330\.3, 1420\.8\]相似文章
SkillGym:通过自动可验证环境生成训练技能使用智能体
SkillGym 提出了一条自动化流水线:从互联网上爬取可复现的技能,构建可验证且难度可控的环境,并收集 19k 条已验证轨迹来训练技能使用智能体。在这些轨迹上对 Qwen3.5 系列模型(2B 到 122B)进行微调后,其在四个技能使用基准上的表现均有所提升;其中 9B 的 SFT 模型在两项基准上超过了 397B 的未训练模型。
CLI-Universe:面向终端代理的可验证任务合成引擎
CLI-Universe是一个合成引擎,通过多维能力分类体系和证据引导的研究生成可验证的终端代理任务,并产生包含6000条轨迹的精炼数据集。在该数据集上微调Qwen3-32B,在Terminal-Bench 2.0上达到了33.4%,为参数量在32B及以下的开源模型树立了新的最优水平。
NexForge:通过需求驱动任务合成扩展LLM的智能体能力
NexForge是一个需求驱动的框架,能够为LLM后训练合成多样化、可执行的智能体任务和专家轨迹,优于先前的方法,并在Terminal-Bench上实现了最先进的开源智能体性能。
Lean4Agent: 代理工作流与轨迹的形式化建模与验证
介绍Lean4Agent,一个使用Lean4对代理工作流和轨迹进行形式化建模与验证的框架,展示了在SWE-Bench和ELAIP-Bench上的性能提升。
Workflow-GYM:面向真实世界专业领域中计算机使用代理任务的长期评估
Workflow-GYM 是一个用于评估 AI 代理在专业领域中长期 GUI 任务的基准。实验表明,即使是最先进的模型也仅能达到约 30% 的成功率,揭示了重大挑战。