MMAE: A Massive Multitask Audio Editing Benchmark
Summary
MMAE is a comprehensive benchmark for instruction-based audio editing across multiple modalities and complexity levels, revealing significant gaps in current model capabilities.
View Cached Full Text
Cached at: 06/08/26, 07:14 AM
Paper page - MMAE: A Massive Multitask Audio Editing Benchmark
Source: https://huggingface.co/papers/2606.07229 Published on Jun 5
#2 Paper of the day Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
Abstract
MMAE presents a comprehensive benchmark for instruction-based audio editing across multiple modalities and complexity levels, revealing significant gaps in current model capabilities.
We introduce MMAE, a MassiveMultitask Audio Editingbenchmark, serving as the first comprehensive evaluation testbed designed for general-purposeinstruction-based audio editing. Spurred by the shift toward intelligent creation, interactive editing has rapidly expanded from visual domains, pioneered by models like Nano-banana 2 for images and Gemini-Omni for video, into audio. However, the current evaluation infrastructure lags severely, remaining highly fragmented and restricted to specific subdomains or basic operations. Unlike existing benchmarks that are limited in scope, MMAE extends to a broad spectrum of real-world scenarios, encompassing 7 distinctaudio modalities, including sound, speech, music, and their mixtures. Furthermore, we establish a comprehensive taxonomy spanning 6 levels oftask complexity, from basic modifications tomulti-hop reasoningandmulti-round editing, 2 levels of granularity, and 8 distinctoperation types. Meticulously curated through human-agent collaboration, MMAE comprises 2,000 high-fidelity samples paired with a pioneeringrubric-based evaluationframework. By decomposing free-form tasks into 17,741 verifiable criteria, this robust rubric-based paradigm enables a precise, multi-dimensional assessment of both instruction following and context consistency. Our extensive evaluation of leading models reveals that current systems remain far from achieving reliable edits. Strikingly, theExact Match Rate(EMR) consistently falls below 5% and plummets to an absolute 0% in complex, mixed-modality tasks, exposing critical bottlenecks in precise execution and structural robustness. We hope MMAE will serve as a catalyst for future advances in the intelligent creation community, providing a clear diagnostic roadmap and establishing a standardized, long-lasting evaluation paradigm for next-generation audio editing systems.
View arXiv pageView PDFGitHub27Add to collection
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2606.07229 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2606.07229 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2606.07229 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
@TencentHunyuan: Can AI truly edit audio, not just generate it? Tencent Hy, in collaboration with SJTU, SII, NTU, TJU, ZODA, PKU, FDU, a…
MMAE is a comprehensive benchmark for multitask audio editing that evaluates AI's ability to precisely modify existing audio clips via natural language instructions, with current models achieving under 5% exact match rate.
AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities
AVE-Compass is a new benchmark for holistic evaluation of audio-video editing abilities, with 145 videos, 196 editing instructions, and 2,688 checklist items. It also proposes AVE-Agent, a modular agent framework that improves cross-modal editing via self-reflection and evaluator feedback.
SpeechEditBench: A Bilingual Multi-Attribute Benchmark for Instruction-Guided Speech Editing
SpeechEditBench is a bilingual multi-attribute benchmark for evaluating instruction-guided speech editing across seven atomic tasks and compositional tasks, using an anchor-based evaluation protocol with three metrics. Evaluation of mainstream Speech LLMs reveals no single model excels across all dimensions, and compositional editing remains highly challenging.
MVEB: Massive Video Embedding Benchmark
This paper introduces MVEB, a large-scale benchmark for evaluating video embeddings across 23 tasks, finding that no single model dominates and that audio's contribution depends on dataset annotation provenance. It integrates into the MTEB ecosystem for unified multimodal evaluation.
MPIE-Bench: Benchmarking Anatomically Plausible Multi-Person Interaction Editing
Introduces MPIE-Bench, a 2,500-sample benchmark for multi-person interaction image editing, along with MPIE-Eval, a mesh-based evaluation method that tracks human judgment more closely than VLM checklists across ten editors.