AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities
Summary
AVE-Compass is a new benchmark for holistic evaluation of audio-video editing abilities, with 145 videos, 196 editing instructions, and 2,688 checklist items. It also proposes AVE-Agent, a modular agent framework that improves cross-modal editing via self-reflection and evaluator feedback.
View Cached Full Text
Cached at: 08/06/26, 05:49 AM
Paper page - AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities
Source: https://huggingface.co/papers/2607.24821
Abstract
Whileinstruction-basedvideoeditinghasadvancedrapidly,real-worldvideoscontaintightlycoupledaudioandvisualsignals,andeditingonemodalityoftenrequirescoordinatedchangesintheother.Existingbenchmarksprimarilyevaluatevisualtransformationsonsilentclipsorisolatedaudioediting,leavingcomplexaudio-visualeditingandcross-modalconsistencyunderexplored.WeintroduceAVE-Compass,acomprehensivebenchmarkwith145curatedsourcevideos,196audio-visuallycouplededitinginstructions,and2,688fine-grainedchecklistitems.ItevaluatesInstructionFollowing,FidelityPreserving,Realism,andEditingIntentthroughchecklist-basedMLLMjudgingandadedicatedrealismrubric,complementedbyautomatedcross-modal,video,andaudiometrics.Extensiveevaluationshowsthatstate-of-the-artmodelsstillstruggletoexecutecross-modalinstructionswhilepreservingnon-targetcontent.WefurtherproposeAVE-Agent,amodularagentframeworkthatdecomposescomplexinstructionsintodependentsubtasksanditerativelyimproveseditingresultsthroughself-reflectionandevaluatorfeedback.AVE-Agentimprovesinstructionexecution,FidelityPreserving,andaudio-visualalignmentinjointeditingwhilemaintainingcompetitiveperceptualquality.
View arXiv pageView PDFProject pageGitHubAdd to collection
Get this paper in your agent:
hf papers read 2607\.24821
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2607.24821 in a model README.md to link it from this page.
Datasets citing this paper1
#### NJU-LINK/AVE-Compass-v2 Viewer• Updated8 days ago • 196 • 925 • 2
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2607.24821 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
LongAV-Compass: Towards Unified Evaluation of Minute-Scale Audio-Visual Generation Across T2AV, I2AV, and V2AV
LongAV-Compass is a comprehensive benchmark for evaluating minute-long audio-visual generation across text, image, and video conditioning modalities, assessing quality, consistency, and alignment over extended temporal sequences.
MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation
This paper introduces MultiRef-Compass, a comprehensive benchmark for multi-reference-to-audio-video generation, comprising 350 curated samples and an evaluation protocol with four dimensions including Basic Quality, Reference Consistency, Audio-Visual Consistency, and Instruction Following.
CoVEBench: Can Video Editing Models Handle Complex Instructions?
Introduces CoVEBench, a new benchmark for evaluating compositional video editing capabilities, addressing limitations in handling complex multi-step instructions. The benchmark includes 416 videos, 626 instructions, and 9,990 checklist items, revealing that current models struggle with compositional editing tasks.
VEFX-Bench: A Holistic Benchmark for Generic Video Editing and Visual Effects
VEFX-Bench introduces a large-scale human-annotated video editing dataset (5,049 examples) with multi-dimensional quality labels and a specialized reward model for standardized evaluation of video editing systems. The paper addresses the lack of comprehensive benchmarks in AI-assisted video creation by providing VEFX-Dataset, VEFX-Reward, and a 300-video-prompt benchmark that reveals gaps in current editing models.
MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation
MSAVBench is the first comprehensive benchmark and adaptive evaluation framework for multi-shot audio-video generation, assessing 19 models across diverse tasks and achieving high alignment with human judgment.