search-agents

Tag

Cards List
#search-agents

False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents

Hugging Face Daily Papers ↗ · yesterday Cached

This paper identifies "co-cheating" in self-evolving search agents, where a proposer and solver converge on shared errors that inflate internal reward without improving true correctness. It proposes CrossFit, a cross-fitted verification scheme that substantially reduces false agreement and boosts performance on seven search benchmarks with Qwen3.5-4B/9B models.

0 favorites 0 likes
#search-agents

FinFIRST: Benchmarking Search Agents for Financial Information Retrieval, Sourcing and Traceability

arXiv cs.CL ↗ · 2026-09-23 Cached

FinFIRST introduces a benchmark for evaluating financial search agents by jointly assessing answers and supporting evidence through atomic rubrics, with 123 expert-authored tasks spanning difficulty levels.

0 favorites 0 likes
#search-agents

Beyond Confidence: Test-Time Scaling for Multi-Turn Search Agents via Retrieval Grounding

arXiv cs.AI ↗ · 2026-08-26 Cached

The paper introduces Retrieval-Grounded Voting (RGV) to address the limitations of confidence-based voting in multi-turn search agents by using lexical overlap with retrieved documents, achieving up to 5.4% accuracy gains.

0 favorites 0 likes
#search-agents

CAFE: Self-Improving Search Agents Need Co-Evolving Feedback

Hugging Face Daily Papers ↗ · 2026-08-25 Cached

CAFE is a framework that couples a search agent and critic via shared parameters to learn in-trajectory corrective feedback, improving search performance and reducing hallucinations across benchmarks.

0 favorites 0 likes
#search-agents

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents

arXiv cs.LG ↗ · 2026-08-14 Cached

Introduces SSPO, a step-level self-distilled policy optimization method for training deep search agents, which uses evidence anchors and advantage weights to improve credit assignment beyond sparse outcome rewards. SSPO outperforms GRPO on benchmarks like BrowseComp and GAIA with only ~5% overhead per step.

0 favorites 0 likes
#search-agents

Mitigating Context Interference for Reliable and Efficient Search Agents

arXiv cs.CL ↗ · 2026-08-12 Cached

This paper systematically studies context interference in multi-turn LLM-based search agents, finding that interference primarily arises from the latest retrieved documents, and introduces a distill-based context refiner to mitigate it. Incorporating context refinement into RL training pipelines significantly improves reliability and efficiency.

0 favorites 0 likes
#search-agents

Search-G1: Grounded Search Agents via Representation-Based Intrinsic Rewards

arXiv cs.CL ↗ · 2026-08-11 Cached

Search-G1 proposes a representation-based intrinsic reward framework for search-augmented language agents, using intervention-calibrated readouts to balance retrieval necessity and evidence reliance, improving search efficiency without costly annotations.

0 favorites 0 likes
#search-agents

SearchAuditor: Auditing and Attributing Failures in Long-Horizon Search Agents

arXiv cs.AI ↗ · 2026-08-07 Cached

This paper introduces SearchAuditBench, a benchmark of 1,243 failed long-horizon search-agent trajectories with expert annotations, and SearchAuditor, a multi-perspective auditing framework that localizes, attributes, and repairs agent failures. Experiments show SearchAuditor outperforms baselines, achieving a 32.3% end-to-end pass rate with frontier models like GPT-5.5.

0 favorites 0 likes
#search-agents

ABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked Credit Assignment

arXiv cs.AI ↗ · 2026-08-06 Cached

This paper proposes Answer-Backtracked Credit Assignment (ABC), a framework that converts sparse trajectory-level outcomes into dense step-level supervision for training long-horizon search agents. The resulting ABSeeker model, built on Qwen3.5-4B, achieves strong results on BrowseComp benchmarks, outperforming same-scale agents and matching larger models.

0 favorites 0 likes
#search-agents

Self-Play Meets Skill Evolution: Self-Evolving Search Agents that Pose, Solve, and Remember

arXiv cs.AI ↗ · 2026-08-03 Cached

This paper introduces SESA, a self-evolving skill-augmented search agent that co-evolves task generation and skill memory via tool-augmented search self-play. It improves accuracy across seven QA benchmarks over baselines while supporting memory-free deployment.

0 favorites 0 likes
#search-agents

Harness-G: A Graph-Structured Harness for Search Agents

Hugging Face Daily Papers ↗ · 2026-07-30 Cached

This paper introduces Harness-G, a graph-structured retrieval framework that reformulates free-form query generation as finite action selection to reduce retrieval aliasing in RL-powered search agents. Across six QA benchmarks, Harness-G outperforms the strongest baseline Graph-R1 by 10.74 points at 1.5B and 3.98 points at 3B scale.

0 favorites 0 likes
#search-agents

Weighing smoke: why AI visibility dashboards are mostly useless

Hacker News Top ↗ · 2026-07-07 Cached

The article argues that AI visibility dashboards, which claim to track brand presence in AI search responses, are unreliable and lack predictive validity, comparing them to weighing smoke due to the inconsistency of AI outputs and the absence of meaningful correlation with business outcomes.

0 favorites 0 likes
#search-agents

When Search Agents Should Ask: DiscoBench for Clarification-Aware Deep Search

arXiv cs.CL ↗ · 2026-06-29 Cached

DiscoBench is a new benchmark that evaluates whether LLM-powered search agents can proactively identify ambiguity in user queries, ask clarifying questions, and recover correct reasoning paths through multi-turn interaction.

0 favorites 0 likes
#search-agents

ProMSA:Progressive Multimodal Search Agents for Knowledge-Based Visual Question Answering

Hugging Face Daily Papers ↗ · 2026-06-26 Cached

Proposes ProMSA, a progressive multimodal search agent for knowledge-based visual question answering that adaptively selects search strategies and optimizes through sequence-level reinforcement learning, achieving consistent gains on E-VQA and InfoSeek.

0 favorites 0 likes
#search-agents

DailyReport: An Open-ended Benchmark for Evaluating Search Agents on Daily Search Tasks

arXiv cs.AI ↗ · 2026-06-12 Cached

DailyReport is an open-ended benchmark for evaluating search agents on daily search tasks, featuring 150 tasks and 3,546 rubrics for interpretable, user-centric evaluation.

0 favorites 0 likes
#search-agents

EvoBrowseComp: Benchmarking Search Agents on Evolving Knowledge

arXiv cs.CL ↗ · 2026-06-12 Cached

This paper introduces EvoBrowseComp, a dynamic benchmark of 400 English and 400 Chinese complex questions that are synthesized via live-web traversal to evaluate search agents without test-set contamination, ensuring robustness against parametric memorization.

0 favorites 0 likes
#search-agents

LoHoSearch: Benchmarking Long-Horizon Search Agents Beyond the Human Difficulty Ceiling

arXiv cs.CL ↗ · 2026-06-12 Cached

LoHoSearch is a new benchmark for evaluating long-horizon search agents, built from a knowledge graph of 7 million Wikipedia entities. It introduces questions with large search spaces and structural complexity to exceed human-authored difficulty ceilings, and shows that the best model achieves only 34.74% accuracy.

0 favorites 0 likes
#search-agents

EvoBrowseComp: Benchmarking Search Agents on Evolving Knowledge

Hugging Face Daily Papers ↗ · 2026-06-11 Cached

EvoBrowseComp is an evolving benchmark with 800 contamination-free questions for evaluating search agents, designed to prevent parametric memorization and maintain temporal freshness through a three-agent framework.

0 favorites 0 likes
#search-agents

FORT-Searcher: Synthesizing Shortcut-Resistant Search Tasks for Training Deep Search Agents

Hugging Face Daily Papers ↗ · 2026-06-10 Cached

FORT-Searcher introduces a framework for synthesizing shortcut-resistant training data for deep search agents by identifying and mitigating four shortcut risks. The resulting agent, trained via supervised fine-tuning, achieves state-of-the-art performance among comparable open-source search agents.

0 favorites 0 likes
#search-agents

@patpcj: Thanks again for your interest in our work! Links here so they don’t get buried under “show more”: Paper : https://arxi…

X AI KOLs Following ↗ · 2026-06-08 Cached

Harness-1 is a 20B search agent trained with reinforcement learning using a stateful search harness, achieving strong results on retrieval benchmarks and outperforming other open search subagents.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback