DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable Environment

arXiv cs.CL Papers

Summary

DeepSearch-Evolve introduces a self-distillation framework for web agents using a verifiable environment (DeepSearch-World) with 420K multi-hop QA tasks, achieving competitive performance without distillation from stronger models.

arXiv:2607.07820v1 Announce Type: new Abstract: Training tool-use agents to improve from their own experience remains challenging, as supervised fine-tuning relies on fixed teacher-distilled trajectories, while sparse-reward reinforcement learning provides weak supervision for long-horizon interactions. We present DeepSearch-Evolve, a self-distillation framework for web agents built on DeepSearch-World, a deterministic and verifiable environment with reproducible search and page-reading tools. DeepSearch-World contains 420K multi-hop QA tasks constructed from entity-level random walks and supports key agentic cognitive behaviors useful for self-evolving, including progress verification, grounded reflection, and failure recovery. DeepSearch-Evolve iteratively performs trajectory generation, filtering, data mixing, and fine-tuning to train stronger agents. Without distillation from more capable models, DeepSearch-World-9B achieves competitive performance compared with open-source agents, reaching 31.2% on BrowseComp, 61.5% on GAIA, and 93.4% on HotpotQA, showing that verifiable environments enable scalable self-evolution for long-horizon web agents. We will release the environment, 420K training pool, validation set, model, and code to facilitate future research on self-improving deep search agents.
Original Article
View Cached Full Text

Cached at: 07/10/26, 06:11 AM

# DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable Environment
Source: [https://arxiv.org/abs/2607.07820](https://arxiv.org/abs/2607.07820)
[View PDF](https://arxiv.org/pdf/2607.07820)

> Abstract:Training tool\-use agents to improve from their own experience remains challenging, as supervised fine\-tuning relies on fixed teacher\-distilled trajectories, while sparse\-reward reinforcement learning provides weak supervision for long\-horizon interactions\. We present DeepSearch\-Evolve, a self\-distillation framework for web agents built on DeepSearch\-World, a deterministic and verifiable environment with reproducible search and page\-reading tools\. DeepSearch\-World contains 420K multi\-hop QA tasks constructed from entity\-level random walks and supports key agentic cognitive behaviors useful for self\-evolving, including progress verification, grounded reflection, and failure recovery\. DeepSearch\-Evolve iteratively performs trajectory generation, filtering, data mixing, and fine\-tuning to train stronger agents\. Without distillation from more capable models, DeepSearch\-World\-9B achieves competitive performance compared with open\-source agents, reaching 31\.2% on BrowseComp, 61\.5% on GAIA, and 93\.4% on HotpotQA, showing that verifiable environments enable scalable self\-evolution for long\-horizon web agents\. We will release the environment, 420K training pool, validation set, model, and code to facilitate future research on self\-improving deep search agents\.

## Submission history

From: Xinyu Geng \[[view email](https://arxiv.org/show-email/31d1c647/2607.07820)\] **\[v1\]**Wed, 8 Jul 2026 18:03:41 UTC \(2,480 KB\)

Similar Articles