CoEvoWhen: Policy-Tool Coevolution for Ultra-Long Video Temporal Grounding
Summary
CoEvoWhen proposes a policy-tool coevolution framework that jointly evolves high-level policies and executable media tools from VLM agentic reasoning trajectories to improve ultra-long video temporal grounding without updating model parameters.
View Cached Full Text
Cached at: 10/01/26, 04:20 AM
Paper page - CoEvoWhen: Policy-Tool Coevolution for Ultra-Long Video Temporal Grounding
Source: https://huggingface.co/papers/2609.40048
Abstract
Ultra-longvideotemporalgroundingrequiresbalancinglong-rangeevidencesearchwithfine-grainedeventunderstandingunderalimitedvisualbudget,yetexistingagenticmethodsstillrelylargelyonpredefinedpoliciesandtoolcapabilities.Motivatedbythis,weproposeanovelpolicy-toolcoevolutionframeworkthatjointlyevolveshigh-levelpoliciesandexecutablemediatoolsfromtheagenticreasoningtrajectoriesofaVLM,formingareusableskillwithoutupdatingmodelparameters.Duringevolution,anexternalskillupdaterdistillstransferabletaskexperienceinlong-videotemporalgrounding,accordinglyrefiningtheorchestrationoflong-rangeimage-basedandfine-grainedvideo-basedobservations.Alongsidethesepolicyupdates,theupdateremploysitscodingcapabilitiestoupgradeexistingtoolsorcreatenewones,adaptingthetoolstolong-videoevidenceacquisition.Equippedwiththeevolvedskill,theVLMautonomouslyorchestratestoolsundertheguidanceoftheevolvedpolicy,coordinatingimageandvideoobservationsforagenticinferencewithoutrelyingonaseparate,strongerplanningmodel.ExtensiveexperimentsspanningfivebenchmarksandthreeVLMsshowthatpolicy-toolcoevolutionconsistentlyimprovestemporalgroundingaccuracyinultra-longvideoswhilereducingvisualtokencostatinference,andthattheevolvedskillyieldssubstantialperformancegainsongenerallong-videoQAwithoutadditionaltask-specificevolution,demonstratingtheeffectivenessandgeneralizabilityofourframeworkforlong-videounderstanding.
View arXiv pageView PDFProject pageGitHub0Add to collection
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.40048 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.40048 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.40048 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning
EvoTrainer introduces an autonomous training framework that co-evolves LLM policies and training harnesses through empirical feedback, outperforming human-engineered RL baselines on mathematical reasoning, code generation, and long-horizon software engineering tasks.
VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System
Introduces VideoCoCo, an agentic dual-engine framework that uses executable Blender code as a chain-of-thought intermediate representation for physically-consistent video generation, achieving state-of-the-art scores on PhyGenBench and VBench-2.0.
CoRe: Co-Evolving Reward Models for Mitigating Latent Reward Hacking in Video Diffusion Models
CoRe introduces a co-evolving latent reward framework that continually refits the reward model on the generator's current samples while anchoring to real-video preferences, preventing latent reward hacking and quality collapse in video diffusion models like Wan2.1-T2V-1.3B.
CoEvolve: Training LLM Agents via Agent-Data Mutual Evolution
CoEvolve proposes an agent-data mutual evolution framework for training LLM agents through closed-loop, interaction-driven learning that adapts both the agent and its training data distribution. The method extracts feedback signals from rollout trajectories to guide LLM-based task synthesis, demonstrating significant improvements (15-19% absolute gains) across multiple Qwen models on AppWorld and BFCL benchmarks.
PACEvolve++: Improving Test-time Learning for Evolutionary Search Agents
The paper introduces PACEvolve++, a reinforcement learning framework that improves test-time policy adaptation for evolutionary search agents by decoupling hypothesis generation from execution.