EmoRES-TTS: Residual-Enhanced Vector Steering for Emotional Speech Generation

Hugging Face Daily Papers Papers

Summary

The paper introduces EmoRES-TTS, a training-free method for emotional speech generation that enhances controllability by decomposing emotion vectors into shared and residual components, achieving superior performance over existing methods on benchmarks.

Emotion-conditioned text-to-speech (TTS) models may fail to express the requested emotion reliably, and improving controllability by additional training is costly in both computation and emotion-labeled speech training data. We therefore study vector steering, a training-free approach that modifies the internal representations of a frozen model. CoCoEmo, a conventional vector steering method for emotion TTS, treats each emotion vector as an indivisible direction controlled by a single global strength, limiting adherence to the requested emotion. In this work, we first discover that an emotion vector can be decomposed into a shared component that moves speech away from neutral expression and a residual component that directs generation toward the requested emotion. Building on this finding, we propose Emotion Residual-Enhanced Steering for TTS (EmoRES), a novel method that controls the two components without retraining the backbone. On IEMOCAP, EmoRES outperforms CoCoEmo across all four objective emotion metrics on the IndexTTS-2 and CosyVoice2 backbones. Rank correlation improves by 26.13 and 12.97 percentage points, corresponding to relative gains of 118.8% and 33.1%, while emotion hit rate improves by 12.95 and 6.92 points, corresponding to relative gains of 20.1% and 9.8%. Human evaluation further shows a relative improvement up to 35.0% in the rate at which listeners correctly identified the dominant requested emotion and up to a 17.3% improvement in fidelity, while listeners prefer EmoRES for naturalness in up to 63.8% of pairwise comparisons. Component ablations further demonstrate that effective control benefits from preserving the shared component while strengthening the residual of the emotion steering vectors.
Original Article
View Cached Full Text

Cached at: 09/30/26, 08:20 AM

Paper page - EmoRES-TTS: Residual-Enhanced Vector Steering for Emotional Speech Generation

Source: https://huggingface.co/papers/2609.38157

Abstract

Emotion-conditionedtext-to-speech(TTS)modelsmayfailtoexpresstherequestedemotionreliably,andimprovingcontrollabilitybyadditionaltrainingiscostlyinbothcomputationandemotion-labeledspeechtrainingdata.Wethereforestudyvectorsteering,atraining-freeapproachthatmodifiestheinternalrepresentationsofafrozenmodel.CoCoEmo,aconventionalvectorsteeringmethodforemotionTTS,treatseachemotionvectorasanindivisibledirectioncontrolledbyasingleglobalstrength,limitingadherencetotherequestedemotion.Inthiswork,wefirstdiscoverthatanemotionvectorcanbedecomposedintoasharedcomponentthatmovesspeechawayfromneutralexpressionandaresidualcomponentthatdirectsgenerationtowardtherequestedemotion.Buildingonthisfinding,weproposeEmotionResidual-EnhancedSteeringforTTS(EmoRES),anovelmethodthatcontrolsthetwocomponentswithoutretrainingthebackbone.OnIEMOCAP,EmoRESoutperformsCoCoEmoacrossallfourobjectiveemotionmetricsontheIndexTTS-2andCosyVoice2backbones.Rankcorrelationimprovesby26.13and12.97percentagepoints,correspondingtorelativegainsof118.8%and33.1%,whileemotionhitrateimprovesby12.95and6.92points,correspondingtorelativegainsof20.1%and9.8%.Humanevaluationfurthershowsarelativeimprovementupto35.0%intherateatwhichlistenerscorrectlyidentifiedthedominantrequestedemotionanduptoa17.3%improvementinfidelity,whilelistenerspreferEmoRESfornaturalnessinupto63.8%ofpairwisecomparisons.Componentablationsfurtherdemonstratethateffectivecontrolbenefitsfrompreservingthesharedcomponentwhilestrengtheningtheresidualoftheemotionsteeringvectors.

View arXiv pageView PDFProject pageGitHub1Add to collection

Get this paper in your agent:

hf papers read 2609\.38157

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.38157 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.38157 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.38157 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

Controllable Affective Generation via Latent Vector Steering

arXiv cs.CL

This paper proposes EmoVec, a lightweight framework for controllable affective generation in large language models via latent vector steering, enabling continuous control over emotional intensity without model weight updates.

Adding emotion control tags to Qwen3-TTS

Reddit r/LocalLLaMA

The article describes fine-tuning Qwen3-TTS to add emotion control tags, overcoming training challenges like codec prefix inconsistencies and generation concurrency issues, and discovering that emotion can be manipulated via affine transformations in speaker embeddings.