EvoDuet: Bilevel Co-Evolution of Web Searching and Task Solving for Scientific Discovery

Hugging Face Daily Papers Papers

Summary

EvoDuet introduces a bilevel optimization framework that co-evolves solutions and web search queries with fixed LLM parameters, using a retrieval gate to decide when to fetch new documents versus reusing stored ones. Across 21 tasks it boosts OpenEvolve's normalized discovery gain notably (e.g., 61.3% to 82.3% with Gemini-3.8-Flash), surpassing prior best scores on eight tasks and generalizing across different evolutionary search scaffolds.

Evolutionary search with large language models (LLMs) can stall when progress requires external knowledge the model lacks. Supplying relevant documents helps, but simply adding web search tool can keep returning the same pages as solutions change. We introduce EvoDuet, a bi-level optimization method that co-evolves solutions and search queries with fixed model parameters. At each iteration, a retrieval gate lets the LLM assess its knowledge gap and choose to retrieve new documents, reuse stored ones, or proceed without them. An inner loop refines queries and ranks documents by the solution scores they are predicted to yield; an outer loop generates candidates in parallel from these documents and records the evaluated outcomes for later searches. Across 21 optimization tasks with one candidate per iteration, EvoDuet raises OpenEvolve's normalized discovery gain from 74.1% to 78.0% with GPT-5.6-Luna and from 61.3% to 82.3% with Gemini-3.8-Flash, whereas Qwen3.5-9B does not benefit. Our best runs surpass the previously reported best scores on eight tasks, including Swap Reduction on Q20 and Rosetta, and match them on three more. EvoDuet also improves with other scaffolds (e.g., Top-K, EvoX) on Sums/Diffs and Denoising, demonstrating its applicability across evolutionary search scaffolds.
Original Article
View Cached Full Text

Cached at: 10/01/26, 04:20 AM

Paper page - EvoDuet: Bilevel Co-Evolution of Web Searching and Task Solving for Scientific Discovery

Source: https://huggingface.co/papers/2609.40340

Abstract

Evolutionarysearchwithlargelanguagemodels(LLMs)canstallwhenprogressrequiresexternalknowledgethemodellacks.Supplyingrelevantdocumentshelps,butsimplyaddingwebsearchtoolcankeepreturningthesamepagesassolutionschange.WeintroduceEvoDuet,abi-leveloptimizationmethodthatco-evolvessolutionsandsearchquerieswithfixedmodelparameters.Ateachiteration,aretrievalgateletstheLLMassessitsknowledgegapandchoosetoretrievenewdocuments,reusestoredones,orproceedwithoutthem.Aninnerlooprefinesqueriesandranksdocumentsbythesolutionscorestheyarepredictedtoyield;anouterloopgeneratescandidatesinparallelfromthesedocumentsandrecordstheevaluatedoutcomesforlatersearches.Across21optimizationtaskswithonecandidateperiteration,EvoDuetraisesOpenEvolve’snormalizeddiscoverygainfrom74.1%to78.0%withGPT-5.6-Lunaandfrom61.3%to82.3%withGemini-3.8-Flash,whereasQwen3.5-9Bdoesnotbenefit.Ourbestrunssurpassthepreviouslyreportedbestscoresoneighttasks,includingSwapReductiononQ20andRosetta,andmatchthemonthreemore.EvoDuetalsoimproveswithotherscaffolds(e.g.,Top-K,EvoX)onSums/DiffsandDenoising,demonstratingitsapplicabilityacrossevolutionarysearchscaffolds.

View arXiv pageView PDFProject pageGitHub0Add to collection

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.40340 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.40340 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.40340 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

EvoBrowseComp: Benchmarking Search Agents on Evolving Knowledge

arXiv cs.CL

This paper introduces EvoBrowseComp, a dynamic benchmark of 400 English and 400 Chinese complex questions that are synthesized via live-web traversal to evaluate search agents without test-set contamination, ensuring robustness against parametric memorization.