AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities

Hugging Face Daily Papers Papers

Summary

AVE-Compass is a new benchmark for holistic evaluation of audio-video editing abilities, with 145 videos, 196 editing instructions, and 2,688 checklist items. It also proposes AVE-Agent, a modular agent framework that improves cross-modal editing via self-reflection and evaluator feedback.

While instruction-based video editing has advanced rapidly, real-world videos contain tightly coupled audio and visual signals, and editing one modality often requires coordinated changes in the other. Existing benchmarks primarily evaluate visual transformations on silent clips or isolated audio editing, leaving complex audio-visual editing and cross-modal consistency underexplored. We introduce AVE-Compass, a comprehensive benchmark with 145 curated source videos, 196 audio-visually coupled editing instructions, and 2,688 fine-grained checklist items. It evaluates Instruction Following, Fidelity Preserving, Realism, and Editing Intent through checklist-based MLLM judging and a dedicated realism rubric, complemented by automated cross-modal, video, and audio metrics. Extensive evaluation shows that state-of-the-art models still struggle to execute cross-modal instructions while preserving non-target content. We further propose AVE-Agent, a modular agent framework that decomposes complex instructions into dependent subtasks and iteratively improves editing results through self-reflection and evaluator feedback. AVE-Agent improves instruction execution, Fidelity Preserving, and audio-visual alignment in joint editing while maintaining competitive perceptual quality.
Original Article
View Cached Full Text

Cached at: 08/06/26, 05:49 AM

Paper page - AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities

Source: https://huggingface.co/papers/2607.24821

Abstract

Whileinstruction-basedvideoeditinghasadvancedrapidly,real-worldvideoscontaintightlycoupledaudioandvisualsignals,andeditingonemodalityoftenrequirescoordinatedchangesintheother.Existingbenchmarksprimarilyevaluatevisualtransformationsonsilentclipsorisolatedaudioediting,leavingcomplexaudio-visualeditingandcross-modalconsistencyunderexplored.WeintroduceAVE-Compass,acomprehensivebenchmarkwith145curatedsourcevideos,196audio-visuallycouplededitinginstructions,and2,688fine-grainedchecklistitems.ItevaluatesInstructionFollowing,FidelityPreserving,Realism,andEditingIntentthroughchecklist-basedMLLMjudgingandadedicatedrealismrubric,complementedbyautomatedcross-modal,video,andaudiometrics.Extensiveevaluationshowsthatstate-of-the-artmodelsstillstruggletoexecutecross-modalinstructionswhilepreservingnon-targetcontent.WefurtherproposeAVE-Agent,amodularagentframeworkthatdecomposescomplexinstructionsintodependentsubtasksanditerativelyimproveseditingresultsthroughself-reflectionandevaluatorfeedback.AVE-Agentimprovesinstructionexecution,FidelityPreserving,andaudio-visualalignmentinjointeditingwhilemaintainingcompetitiveperceptualquality.

View arXiv pageView PDFProject pageGitHubAdd to collection

Get this paper in your agent:

hf papers read 2607\.24821

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2607.24821 in a model README.md to link it from this page.

Datasets citing this paper1

#### NJU-LINK/AVE-Compass-v2 Viewer• Updated8 days ago • 196 • 925 • 2

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2607.24821 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

CoVEBench: Can Video Editing Models Handle Complex Instructions?

Hugging Face Daily Papers

Introduces CoVEBench, a new benchmark for evaluating compositional video editing capabilities, addressing limitations in handling complex multi-step instructions. The benchmark includes 416 videos, 626 instructions, and 9,990 checklist items, revealing that current models struggle with compositional editing tasks.

VEFX-Bench: A Holistic Benchmark for Generic Video Editing and Visual Effects

Hugging Face Daily Papers

VEFX-Bench introduces a large-scale human-annotated video editing dataset (5,049 examples) with multi-dimensional quality labels and a specialized reward model for standardized evaluation of video editing systems. The paper addresses the lack of comprehensive benchmarks in AI-assisted video creation by providing VEFX-Dataset, VEFX-Reward, and a 300-video-prompt benchmark that reveals gaps in current editing models.