Comprehensive Benchmarking of Long-Form Speech Generation in Diverse Scenarios
Summary
Swanbench-Speech is a comprehensive benchmark for evaluating long-form speech generation across diverse scenarios, using multi-dimensional metrics covering acoustics, semantics, and expressiveness, revealing limitations of current models.
View Cached Full Text
Cached at: 06/01/26, 07:18 AM
Paper page - Comprehensive Benchmarking of Long-Form Speech Generation in Diverse Scenarios
Source: https://huggingface.co/papers/2605.28618 Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
Abstract
Swanbench-Speech addresses the lack of comprehensive long-form speech evaluation by providing a benchmark with diverse scenarios, multi-dimensional metrics, and insights into model limitations.
Recentadvancesinspeechgenerationhaveenabledhigh-fidelitysynthesis,yetsystematicevaluationofmodelsunderlong-contextconditionsremainslargelyunderexplored.Acomprehensiveevaluationbenchmarkforlong-formspeechisindispensablefortworeasons:1)existingtestscenariosareoftenconfinedtolimiteddomains,creatingasignificantgapwiththediversedownstreamapplications;2)existingmetricsoverlookcriticallong-textfactorssuchasconsistencyandcoherence,failingtogeneralizereliably.Tothisend,weproposeSwanbench-Speech,acomprehensivebenchmarkthatdecomposeslong-formspeechqualityintospecific,disentangleddimensions.SwanBench-Speechhasthreekeyproperties.1)Richspeechscenarios:Focusingonlong-formspeechgenerationanddialoggeneration,SwanBench-Speechcoversacoustics,semantics,andexpressivenesschallenges,andconsistsof1,101samplesspanning17commonspeechscenarios;2)Comprehensiveevaluationdimensions:Alongtheacoustics,semantics,andexpressivenessaxes,SwanBench-Speechdefinesanautomatedevaluationprotocolwithsevenmetricstoprovideacomprehensive,accurate,andstandardizedassessment;3)ValuableInsights:Throughextensiveexperiments,werevealthatcurrentmodelsstillstruggleinhighlyexpressivescenariosandexhibitanotablegapinconsistencyandhierarchycomparedtorealrecordings.
View arXiv pageView PDFProject pageAdd to collection
Get this paper in your agent:
hf papers read 2605\.28618
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2605.28618 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2605.28618 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2605.28618 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue
SwanVoice is a zero-shot text-to-speech model designed for expressive long-form monologue and dialogue synthesis, combining VAE, flow-matching DiT, and diffusion post-training to achieve higher richness and hierarchy scores than existing baselines.
SEA-SpeechBench: A Large-Scale Multitask Benchmark for Speech Understanding Across Southeast Asia
SEA-SpeechBench is the first large-scale multitask benchmark for evaluating speech understanding in 11 Southeast Asian languages, highlighting performance gaps in current models.
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks
This paper introduces SwanTale, a unified multi-speaker speech and audio generation model supporting both zero-shot and instruct tasks, along with SwanData-Caption for data annotation and SwanVAE for high-quality multi-audio-modality generation.
SpeechEditBench: A Bilingual Multi-Attribute Benchmark for Instruction-Guided Speech Editing
SpeechEditBench is a bilingual multi-attribute benchmark for evaluating instruction-guided speech editing across seven atomic tasks and compositional tasks, using an anchor-based evaluation protocol with three metrics. Evaluation of mainstream Speech LLMs reveals no single model excels across all dimensions, and compositional editing remains highly challenging.
OpenSTBench: Beyond Semantic Evaluation for Speech Translation
OpenSTBench is a unified multidimensional evaluation framework for speech translation systems that jointly assesses translation quality, speech quality, speaker preservation, emotion fidelity, and latency across both S2TT and S2ST systems in offline and streaming settings. The framework addresses the gap left by fragmented evaluation protocols and provides a reproducible benchmark for comparing heterogeneous speech translation systems.