EvoDuet: Bilevel Co-Evolution of Web Searching and Task Solving for Scientific Discovery
Summary
EvoDuet introduces a bilevel optimization framework that co-evolves solutions and web search queries with fixed LLM parameters, using a retrieval gate to decide when to fetch new documents versus reusing stored ones. Across 21 tasks it boosts OpenEvolve's normalized discovery gain notably (e.g., 61.3% to 82.3% with Gemini-3.8-Flash), surpassing prior best scores on eight tasks and generalizing across different evolutionary search scaffolds.
View Cached Full Text
Cached at: 10/01/26, 04:20 AM
Paper page - EvoDuet: Bilevel Co-Evolution of Web Searching and Task Solving for Scientific Discovery
Source: https://huggingface.co/papers/2609.40340
Abstract
Evolutionarysearchwithlargelanguagemodels(LLMs)canstallwhenprogressrequiresexternalknowledgethemodellacks.Supplyingrelevantdocumentshelps,butsimplyaddingwebsearchtoolcankeepreturningthesamepagesassolutionschange.WeintroduceEvoDuet,abi-leveloptimizationmethodthatco-evolvessolutionsandsearchquerieswithfixedmodelparameters.Ateachiteration,aretrievalgateletstheLLMassessitsknowledgegapandchoosetoretrievenewdocuments,reusestoredones,orproceedwithoutthem.Aninnerlooprefinesqueriesandranksdocumentsbythesolutionscorestheyarepredictedtoyield;anouterloopgeneratescandidatesinparallelfromthesedocumentsandrecordstheevaluatedoutcomesforlatersearches.Across21optimizationtaskswithonecandidateperiteration,EvoDuetraisesOpenEvolve’snormalizeddiscoverygainfrom74.1%to78.0%withGPT-5.6-Lunaandfrom61.3%to82.3%withGemini-3.8-Flash,whereasQwen3.5-9Bdoesnotbenefit.Ourbestrunssurpassthepreviouslyreportedbestscoresoneighttasks,includingSwapReductiononQ20andRosetta,andmatchthemonthreemore.EvoDuetalsoimproveswithotherscaffolds(e.g.,Top-K,EvoX)onSums/DiffsandDenoising,demonstratingitsapplicabilityacrossevolutionarysearchscaffolds.
View arXiv pageView PDFProject pageGitHub0Add to collection
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.40340 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.40340 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.40340 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
EvoBrowseComp: Benchmarking Search Agents on Evolving Knowledge
EvoBrowseComp is an evolving benchmark with 800 contamination-free questions for evaluating search agents, designed to prevent parametric memorization and maintain temporal freshness through a three-agent framework.
EvoBrowseComp: Benchmarking Search Agents on Evolving Knowledge
This paper introduces EvoBrowseComp, a dynamic benchmark of 400 English and 400 Chinese complex questions that are synthesized via live-web traversal to evaluate search agents without test-set contamination, ensuring robustness against parametric memorization.
EvoSci: A Bio-Inspired Multi-Agent Framework for the Evolution of Scientific Discovery
EvoSci proposes a bio-inspired multi-agent framework that integrates evolutionary algorithms with knowledge graph modeling to iteratively generate, evaluate, and refine research ideas, achieving top performance in peer-review evaluations.
EvoMaster: A Foundational Agent Framework for Building Evolving Autonomous Scientific Agents at Scale
EvoMaster is a scalable, self-evolving agent framework for large-scale scientific discovery that enables iterative hypothesis refinement and knowledge accumulation across experimental cycles. It achieves state-of-the-art results on four benchmarks including Humanity's Last Exam (41.1%) and MLE-Bench Lite (75.8%), outperforming general-purpose baselines by up to 316%.
MetaEvo: A Meta-Optimization Framework for Experience-Driven Agent Evolution
MetaEvo proposes a two-stage framework for continual evolution of LLM-based agents, using preference-based optimization to enhance principle abstraction and modular architecture for experience reuse, outperforming strong baselines on reasoning benchmarks.