EmoRES-TTS: Residual-Enhanced Vector Steering for Emotional Speech Generation
Summary
The paper introduces EmoRES-TTS, a training-free method for emotional speech generation that enhances controllability by decomposing emotion vectors into shared and residual components, achieving superior performance over existing methods on benchmarks.
View Cached Full Text
Cached at: 09/30/26, 08:20 AM
Paper page - EmoRES-TTS: Residual-Enhanced Vector Steering for Emotional Speech Generation
Source: https://huggingface.co/papers/2609.38157
Abstract
Emotion-conditionedtext-to-speech(TTS)modelsmayfailtoexpresstherequestedemotionreliably,andimprovingcontrollabilitybyadditionaltrainingiscostlyinbothcomputationandemotion-labeledspeechtrainingdata.Wethereforestudyvectorsteering,atraining-freeapproachthatmodifiestheinternalrepresentationsofafrozenmodel.CoCoEmo,aconventionalvectorsteeringmethodforemotionTTS,treatseachemotionvectorasanindivisibledirectioncontrolledbyasingleglobalstrength,limitingadherencetotherequestedemotion.Inthiswork,wefirstdiscoverthatanemotionvectorcanbedecomposedintoasharedcomponentthatmovesspeechawayfromneutralexpressionandaresidualcomponentthatdirectsgenerationtowardtherequestedemotion.Buildingonthisfinding,weproposeEmotionResidual-EnhancedSteeringforTTS(EmoRES),anovelmethodthatcontrolsthetwocomponentswithoutretrainingthebackbone.OnIEMOCAP,EmoRESoutperformsCoCoEmoacrossallfourobjectiveemotionmetricsontheIndexTTS-2andCosyVoice2backbones.Rankcorrelationimprovesby26.13and12.97percentagepoints,correspondingtorelativegainsof118.8%and33.1%,whileemotionhitrateimprovesby12.95and6.92points,correspondingtorelativegainsof20.1%and9.8%.Humanevaluationfurthershowsarelativeimprovementupto35.0%intherateatwhichlistenerscorrectlyidentifiedthedominantrequestedemotionanduptoa17.3%improvementinfidelity,whilelistenerspreferEmoRESfornaturalnessinupto63.8%ofpairwisecomparisons.Componentablationsfurtherdemonstratethateffectivecontrolbenefitsfrompreservingthesharedcomponentwhilestrengtheningtheresidualoftheemotionsteeringvectors.
View arXiv pageView PDFProject pageGitHub1Add to collection
Get this paper in your agent:
hf papers read 2609\.38157
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.38157 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.38157 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.38157 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Controllable Affective Generation via Latent Vector Steering
This paper proposes EmoVec, a lightweight framework for controllable affective generation in large language models via latent vector steering, enabling continuous control over emotional intensity without model weight updates.
We built NeuTTS-2E, an open-source on-device TTS model with 7 controllable emotions
NeuTTS-2E is an open-source on-device TTS model that supports seven controllable emotions.
Adding emotion control tags to Qwen3-TTS
The article describes fine-tuning Qwen3-TTS to add emotion control tags, overcoming training challenges like codec prefix inconsistencies and generation concurrency issues, and discovering that emotion can be manipulated via affine transformations in speaker embeddings.
EmoStance: Response-Side Affective-Orientation Control for Empathetic Response Generation via Emoji Weak Supervision
EmoStance is a method for empathetic response generation that uses emoji weak supervision to model response-side affective orientation, improving contextual specificity and perceived responsiveness in dialogues.
VA-DPO: Valence-Arousal Direct Preference Optimization for Controllable Emotion Generation in Language Models
This paper proposes VA-DPO, a method for controllable emotion generation in language models using continuous valence-arousal dimensions, which improves over prompting techniques without degrading model performance.