Tag
This paper identifies "co-cheating" in self-evolving search agents, where a proposer and solver converge on shared errors that inflate internal reward without improving true correctness. It proposes CrossFit, a cross-fitted verification scheme that substantially reduces false agreement and boosts performance on seven search benchmarks with Qwen3.5-4B/9B models.
FinFIRST introduces a benchmark for evaluating financial search agents by jointly assessing answers and supporting evidence through atomic rubrics, with 123 expert-authored tasks spanning difficulty levels.
The paper introduces Retrieval-Grounded Voting (RGV) to address the limitations of confidence-based voting in multi-turn search agents by using lexical overlap with retrieved documents, achieving up to 5.4% accuracy gains.
CAFE is a framework that couples a search agent and critic via shared parameters to learn in-trajectory corrective feedback, improving search performance and reducing hallucinations across benchmarks.
Introduces SSPO, a step-level self-distilled policy optimization method for training deep search agents, which uses evidence anchors and advantage weights to improve credit assignment beyond sparse outcome rewards. SSPO outperforms GRPO on benchmarks like BrowseComp and GAIA with only ~5% overhead per step.
This paper systematically studies context interference in multi-turn LLM-based search agents, finding that interference primarily arises from the latest retrieved documents, and introduces a distill-based context refiner to mitigate it. Incorporating context refinement into RL training pipelines significantly improves reliability and efficiency.
Search-G1 proposes a representation-based intrinsic reward framework for search-augmented language agents, using intervention-calibrated readouts to balance retrieval necessity and evidence reliance, improving search efficiency without costly annotations.
This paper introduces SearchAuditBench, a benchmark of 1,243 failed long-horizon search-agent trajectories with expert annotations, and SearchAuditor, a multi-perspective auditing framework that localizes, attributes, and repairs agent failures. Experiments show SearchAuditor outperforms baselines, achieving a 32.3% end-to-end pass rate with frontier models like GPT-5.5.
This paper proposes Answer-Backtracked Credit Assignment (ABC), a framework that converts sparse trajectory-level outcomes into dense step-level supervision for training long-horizon search agents. The resulting ABSeeker model, built on Qwen3.5-4B, achieves strong results on BrowseComp benchmarks, outperforming same-scale agents and matching larger models.
This paper introduces SESA, a self-evolving skill-augmented search agent that co-evolves task generation and skill memory via tool-augmented search self-play. It improves accuracy across seven QA benchmarks over baselines while supporting memory-free deployment.
This paper introduces Harness-G, a graph-structured retrieval framework that reformulates free-form query generation as finite action selection to reduce retrieval aliasing in RL-powered search agents. Across six QA benchmarks, Harness-G outperforms the strongest baseline Graph-R1 by 10.74 points at 1.5B and 3.98 points at 3B scale.
The article argues that AI visibility dashboards, which claim to track brand presence in AI search responses, are unreliable and lack predictive validity, comparing them to weighing smoke due to the inconsistency of AI outputs and the absence of meaningful correlation with business outcomes.
DiscoBench is a new benchmark that evaluates whether LLM-powered search agents can proactively identify ambiguity in user queries, ask clarifying questions, and recover correct reasoning paths through multi-turn interaction.
Proposes ProMSA, a progressive multimodal search agent for knowledge-based visual question answering that adaptively selects search strategies and optimizes through sequence-level reinforcement learning, achieving consistent gains on E-VQA and InfoSeek.
DailyReport is an open-ended benchmark for evaluating search agents on daily search tasks, featuring 150 tasks and 3,546 rubrics for interpretable, user-centric evaluation.
This paper introduces EvoBrowseComp, a dynamic benchmark of 400 English and 400 Chinese complex questions that are synthesized via live-web traversal to evaluate search agents without test-set contamination, ensuring robustness against parametric memorization.
LoHoSearch is a new benchmark for evaluating long-horizon search agents, built from a knowledge graph of 7 million Wikipedia entities. It introduces questions with large search spaces and structural complexity to exceed human-authored difficulty ceilings, and shows that the best model achieves only 34.74% accuracy.
EvoBrowseComp is an evolving benchmark with 800 contamination-free questions for evaluating search agents, designed to prevent parametric memorization and maintain temporal freshness through a three-agent framework.
FORT-Searcher introduces a framework for synthesizing shortcut-resistant training data for deep search agents by identifying and mitigating four shortcut risks. The resulting agent, trained via supervised fine-tuning, achieves state-of-the-art performance among comparable open-source search agents.
Harness-1 is a 20B search agent trained with reinforcement learning using a stateful search harness, achieving strong results on retrieval benchmarks and outperforming other open search subagents.