Comprehensive Benchmarking of Long-Form Speech Generation in Diverse Scenarios

Hugging Face Daily Papers Papers

Summary

Swanbench-Speech is a comprehensive benchmark for evaluating long-form speech generation across diverse scenarios, using multi-dimensional metrics covering acoustics, semantics, and expressiveness, revealing limitations of current models.

Recent advances in speech generation have enabled high-fidelity synthesis, yet systematic evaluation of models under long-context conditions remains largely underexplored. A comprehensive evaluation benchmark for long-form speech is indispensable for two reasons: 1) existing test scenarios are often confined to limited domains, creating a significant gap with the diverse downstream applications; 2) existing metrics overlook critical long-text factors such as consistency and coherence, failing to generalize reliably. To this end, we propose Swanbench-Speech, a comprehensive benchmark that decomposes long-form speech quality into specific, disentangled dimensions. SwanBench-Speech has three key properties. 1) Rich speech scenarios: Focusing on long-form speech generation and dialog generation, SwanBench-Speech covers acoustics, semantics, and expressiveness challenges, and consists of 1,101 samples spanning 17 common speech scenarios; 2) Comprehensive evaluation dimensions: Along the acoustics, semantics, and expressiveness axes, SwanBench-Speech defines an automated evaluation protocol with seven metrics to provide a comprehensive, accurate, and standardized assessment; 3) Valuable Insights: Through extensive experiments, we reveal that current models still struggle in highly expressive scenarios and exhibit a notable gap in consistency and hierarchy compared to real recordings.
Original Article
View Cached Full Text

Cached at: 06/01/26, 07:18 AM

Paper page - Comprehensive Benchmarking of Long-Form Speech Generation in Diverse Scenarios

Source: https://huggingface.co/papers/2605.28618 Authors:

,

,

,

,

,

,

,

,

,

,

,

,

,

Abstract

Swanbench-Speech addresses the lack of comprehensive long-form speech evaluation by providing a benchmark with diverse scenarios, multi-dimensional metrics, and insights into model limitations.

Recentadvancesinspeechgenerationhaveenabledhigh-fidelitysynthesis,yetsystematicevaluationofmodelsunderlong-contextconditionsremainslargelyunderexplored.Acomprehensiveevaluationbenchmarkforlong-formspeechisindispensablefortworeasons:1)existingtestscenariosareoftenconfinedtolimiteddomains,creatingasignificantgapwiththediversedownstreamapplications;2)existingmetricsoverlookcriticallong-textfactorssuchasconsistencyandcoherence,failingtogeneralizereliably.Tothisend,weproposeSwanbench-Speech,acomprehensivebenchmarkthatdecomposeslong-formspeechqualityintospecific,disentangleddimensions.SwanBench-Speechhasthreekeyproperties.1)Richspeechscenarios:Focusingonlong-formspeechgenerationanddialoggeneration,SwanBench-Speechcoversacoustics,semantics,andexpressivenesschallenges,andconsistsof1,101samplesspanning17commonspeechscenarios;2)Comprehensiveevaluationdimensions:Alongtheacoustics,semantics,andexpressivenessaxes,SwanBench-Speechdefinesanautomatedevaluationprotocolwithsevenmetricstoprovideacomprehensive,accurate,andstandardizedassessment;3)ValuableInsights:Throughextensiveexperiments,werevealthatcurrentmodelsstillstruggleinhighlyexpressivescenariosandexhibitanotablegapinconsistencyandhierarchycomparedtorealrecordings.

View arXiv pageView PDFProject pageAdd to collection

Get this paper in your agent:

hf papers read 2605\.28618

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2605.28618 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2605.28618 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2605.28618 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

SpeechEditBench: A Bilingual Multi-Attribute Benchmark for Instruction-Guided Speech Editing

Hugging Face Daily Papers

SpeechEditBench is a bilingual multi-attribute benchmark for evaluating instruction-guided speech editing across seven atomic tasks and compositional tasks, using an anchor-based evaluation protocol with three metrics. Evaluation of mainstream Speech LLMs reveals no single model excels across all dimensions, and compositional editing remains highly challenging.

OpenSTBench: Beyond Semantic Evaluation for Speech Translation

Hugging Face Daily Papers

OpenSTBench is a unified multidimensional evaluation framework for speech translation systems that jointly assesses translation quality, speech quality, speaker preservation, emotion fidelity, and latency across both S2TT and S2ST systems in offline and streaming settings. The framework addresses the gap left by fragmented evaluation protocols and provides a reproducible benchmark for comparing heterogeneous speech translation systems.