Dual-Loop Self-Evolution via Verifiable Emotion Feedback for Multi-Turn Empathetic Dialogue
Summary
This paper proposes a dual-loop self-evolution framework for multi-turn empathetic dialogue, using verifiable emotion feedback to optimize policy and adapt training distribution. On SAGE, it improves Qwen3-8B Overall from 53.87 to 79.24, outperforming uniform emotion-reward RL by 7.23 points.
View Cached Full Text
Cached at: 08/12/26, 08:36 AM
# Dual-Loop Self-Evolution via Verifiable Emotion Feedback for Multi-Turn Empathetic Dialogue
Source: [https://arxiv.org/html/2608.10626](https://arxiv.org/html/2608.10626)
Yi Wei1,2\\equalcontrib, Shuo Jiang1\\equalcontrib, Huaixia Dou1, Jie Zhu1\\corresponding, Junhui Li3, Lifan Guo1, Feng Chen1, Chi Zhang1
###### Abstract
Large language models have demonstrated conversational capabilities, yet empathetic competence remains challenging\. Empathetic support is inherently multi\-turn and path\-dependent: users disclose concerns gradually, emotions evolve over time, and early responses shape trust and receptivity\. Reinforcement learning with verifiable emotion rewards provides scalable supervision for long\-horizon interactions\. However, existing methods evolve the dialogue policy while keeping its training interaction distribution fixed, creating a mismatch between policy competence and training experience\. We introduce a dual\-loop self\-evolution framework driven by verifiable emotion feedback\. With the user simulator and verifier frozen, the inner loop optimizes the multi\-turn policy using continuous emotion rewards, while the outer loop uses the same outcomes to estimate policy\-relative interaction utility and adapt experience\. To obtain estimates from sparse, stochastic rollouts, the framework holds the scenario and interaction state constant within each group and prioritizes conditions whose group pass rates lie near the policy’s competence boundary\. A hierarchical controller shares evidence across support intents, while uncertainty\-guided exploration and uniform rehearsal prevent premature exclusion\. The resulting distribution generates trajectories, closing both loops without increasing the rollout budget\. On SAGE, our framework raises Qwen3\-8B Overall from 53\.87 to 79\.24 and outperforms protocol\-matched uniform emotion\-reward reinforcement learning by 7\.23 points\.
## 1Introduction
Large language models \(LLMs\) have become capable conversational systems, yet reliable empathetic support remains difficult\(Rashkinet al\.[2019](https://arxiv.org/html/2608.10626#bib.bib38); Liuet al\.[2021](https://arxiv.org/html/2608.10626#bib.bib12); Chenget al\.[2022](https://arxiv.org/html/2608.10626#bib.bib21)\)\. Empathy is inherently multi\-turn: concerns surface incrementally over turns, affect rises and falls with the conversation, and each response constrains what the user is willing to reveal next\(Zhouet al\.[2023](https://arxiv.org/html/2608.10626#bib.bib11); Wanet al\.[2025](https://arxiv.org/html/2608.10626#bib.bib25); Wuet al\.[2026](https://arxiv.org/html/2608.10626#bib.bib47); Zhanget al\.[2026](https://arxiv.org/html/2608.10626#bib.bib5)\)\. A reply that appears appropriate in isolation may still derail the subsequent conversation\. An effective policy must therefore learn how its actions influence an evolving user over a complete trajectory, rather than merely imitate locally plausible supportive text\.
Figure 1:Training paradigms\. Static curricula lag behind competence; unconstrained self\-play creates moving targets; ours adapts the interaction distribution with stable role\-play\.Learning such behavior requires coherent long\-horizon experience and reliable outcome supervision\. Human\-authored support dialogues are costly to collect, difficult to standardize, and limited in the range of user reactions they expose\(Liuet al\.[2021](https://arxiv.org/html/2608.10626#bib.bib12); Chenget al\.[2022](https://arxiv.org/html/2608.10626#bib.bib21)\)\. Offline augmentation improves coverage but remains fixed once generated\(Zhenget al\.[2023](https://arxiv.org/html/2608.10626#bib.bib22); Yeet al\.[2025](https://arxiv.org/html/2608.10626#bib.bib42)\)\. Interactive user simulation offers a scalable alternative: a policy can observe how different responses change later disclosure and emotion, provided the simulated user preserves its persona, context, and hidden need across turns\(Yoonet al\.[2024](https://arxiv.org/html/2608.10626#bib.bib27); Gromadaet al\.[2025](https://arxiv.org/html/2608.10626#bib.bib37); Douet al\.[2025](https://arxiv.org/html/2608.10626#bib.bib28); Zhanget al\.[2026](https://arxiv.org/html/2608.10626#bib.bib5)\)\. Verifiable emotion rewards further turn these reactions into trajectory\-level supervision, enabling reinforcement learning \(RL\) to optimize long\-term affective consequences rather than only match static references\(Wanget al\.[2026b](https://arxiv.org/html/2608.10626#bib.bib6)\)\.
However, existing methods still treat the training interaction distribution as a predefined and fixed external condition\(Wanget al\.[2026b](https://arxiv.org/html/2608.10626#bib.bib6)\)\. Emotion rewards continually update the dialogue policy, but do not change how subsequent experience is allocated\. Consequently, each condition receives the same expected rollout budget whether the current policy has mastered it, repeatedly fails on it, or still finds it informative\. As competence evolves, the learning value of an interaction changes with it, while the allocation of expensive multi\-turn experience remains static\. This creates a growing mismatch between what the policy can currently do and what it is asked to practice\.
Simply enlarging the interaction pool or prescribing an easy\-to\-hard schedule does not resolve this mismatch\. Interaction difficulty is not an intrinsic, permanent property of a scenario\. The same hesitant or emotionally activated user may be excessive for an early policy, informative for an improving policy, and redundant for a mature one\. A curriculum fixed before training cannot remain aligned with this moving competence boundary\(Bengioet al\.[2009](https://arxiv.org/html/2608.10626#bib.bib1); Kumaret al\.[2010](https://arxiv.org/html/2608.10626#bib.bib2)\)\. Adaptive curriculum and replay methods recognize that training utility is learner\-dependent\(Graveset al\.[2017](https://arxiv.org/html/2608.10626#bib.bib46); Schaulet al\.[2016](https://arxiv.org/html/2608.10626#bib.bib3); Jianget al\.[2021](https://arxiv.org/html/2608.10626#bib.bib4); Shiet al\.[2026](https://arxiv.org/html/2608.10626#bib.bib13); Jianget al\.[2025](https://arxiv.org/html/2608.10626#bib.bib14)\), but existing formulations do not directly address sparse, stochastic outcomes from long multi\-turn interactions with a role\-playing user\.
Self\-play offers an apparently natural solution by evolving both the assistant and its simulated partner\(Chenet al\.[2024](https://arxiv.org/html/2608.10626#bib.bib43); Daiet al\.[2026](https://arxiv.org/html/2608.10626#bib.bib7)\)\. In empathetic dialogue, however, the simulator participates in both trajectory generation and outcome formation\. Unconstrained adaptation can make users arbitrarily resistant, overly accepting, or inconsistent with their original persona and hidden need\(Yoonet al\.[2024](https://arxiv.org/html/2608.10626#bib.bib27); Douet al\.[2025](https://arxiv.org/html/2608.10626#bib.bib28)\)\. Reward changes would then reflect both assistant improvement and simulator drift, obscuring credit assignment and turning evaluation into a moving target\. A more difficult user is not necessarily a more useful or faithful teacher\.
To this end, we propose a dual\-loop self\-evolution framework that jointly advances the dialogue policy and the experience distribution through shared verifiable emotion feedback; Figure[1](https://arxiv.org/html/2608.10626#S1.F1)contrasts this design with static curricula and unconstrained self\-play\. The loops are nested but have distinct targets: the inner loop uses continuous emotion rewards to improve*how the policy responds*, while the outer loop reuses group outcomes to identify*what the policy should practice next*and reallocates subsequent interactions accordingly\. The updated distribution produces new trajectories that again supervise both processes, turning a fixed interaction pool into a capability\-aligned online curriculum grounded in comparable multi\-turn outcomes\.
Experiments show that our framework achieves 79\.24 SAGE Overall, surpassing uniform emotion\-reward RL by 7\.23 points, or a 10\.0% relative gain\. The improvement requires no additional rollouts and holds across complementary evaluation protocols and component ablations\.
Our contributions are threefold:
- •We identify*curriculum–policy misalignment*as a central challenge in multi\-turn emotion\-reward RL: policy competence evolves while its interaction distribution remains fixed\.
- •We introduce a dual\-loop self\-evolution framework in which shared emotion feedback jointly drives policy improvement and capability\-aligned adaptation of the training experience distribution\.
- •Multi\-run, cross\-benchmark, and component\-level evaluations demonstrate substantial performance gains and improved rollout\-budget efficiency over uniform emotion\-reward training\.
## 2Related Work
#### Empathetic dialogue and affective reinforcement learning\.
Empathetic dialogue has progressed from emotion\-grounded generation to long\-horizon support over latent needs and strategies\(Rashkinet al\.[2019](https://arxiv.org/html/2608.10626#bib.bib38); Linet al\.[2019](https://arxiv.org/html/2608.10626#bib.bib39); Majumderet al\.[2020](https://arxiv.org/html/2608.10626#bib.bib40); Liuet al\.[2021](https://arxiv.org/html/2608.10626#bib.bib12); Chenget al\.[2022](https://arxiv.org/html/2608.10626#bib.bib21)\), with work on planning, augmentation, interpretable reasoning, strategy prediction, and personalization\(Zhouet al\.[2023](https://arxiv.org/html/2608.10626#bib.bib11); Zhenget al\.[2023](https://arxiv.org/html/2608.10626#bib.bib22); Zhanget al\.[2024b](https://arxiv.org/html/2608.10626#bib.bib23); Kanget al\.[2024](https://arxiv.org/html/2608.10626#bib.bib24); Wanet al\.[2025](https://arxiv.org/html/2608.10626#bib.bib25); Shiet al\.[2025](https://arxiv.org/html/2608.10626#bib.bib41); Yeet al\.[2025](https://arxiv.org/html/2608.10626#bib.bib42)\)\. RL encourages value\-sensitive support\(Kimet al\.[2025](https://arxiv.org/html/2608.10626#bib.bib33); Zhaoet al\.[2025](https://arxiv.org/html/2608.10626#bib.bib26); Wanget al\.[2025](https://arxiv.org/html/2608.10626#bib.bib34)\), while LLM simulators enable scalable interaction but require faithful role\-play and assessment\(Yoonet al\.[2024](https://arxiv.org/html/2608.10626#bib.bib27); Gromadaet al\.[2025](https://arxiv.org/html/2608.10626#bib.bib37); Douet al\.[2025](https://arxiv.org/html/2608.10626#bib.bib28)\)\. SAGE supplies evolving emotional personas\(Zhanget al\.[2026](https://arxiv.org/html/2608.10626#bib.bib5)\), and verifiable emotion rewards optimize policies\(Wanget al\.[2026b](https://arxiv.org/html/2608.10626#bib.bib6)\); we instead align the interaction distribution with the evolving policy\.
#### Adaptive curricula and self\-evolving training\.
Curriculum, self\-paced, and replay methods allocate computation by learner\-dependent utility\(Bengioet al\.[2009](https://arxiv.org/html/2608.10626#bib.bib1); Kumaret al\.[2010](https://arxiv.org/html/2608.10626#bib.bib2); Jianget al\.[2015](https://arxiv.org/html/2608.10626#bib.bib32); Schaulet al\.[2016](https://arxiv.org/html/2608.10626#bib.bib3); Jianget al\.[2021](https://arxiv.org/html/2608.10626#bib.bib4)\); LLM self\-improvement adds self\-play, self\-judgment, and online optimization\(Chenet al\.[2024](https://arxiv.org/html/2608.10626#bib.bib43); Yuanet al\.[2024](https://arxiv.org/html/2608.10626#bib.bib44); Ahmadianet al\.[2024](https://arxiv.org/html/2608.10626#bib.bib45)\)\. Recent work adapts prompts or rollouts using success, uncertainty, variance, and progress\(Shiet al\.[2026](https://arxiv.org/html/2608.10626#bib.bib13); Jianget al\.[2025](https://arxiv.org/html/2608.10626#bib.bib14),[2026](https://arxiv.org/html/2608.10626#bib.bib19); Liuet al\.[2026](https://arxiv.org/html/2608.10626#bib.bib17); Liet al\.[2026](https://arxiv.org/html/2608.10626#bib.bib18); Zenget al\.[2026](https://arxiv.org/html/2608.10626#bib.bib20); Nguyenet al\.[2026](https://arxiv.org/html/2608.10626#bib.bib15)\), or evolves tasks, users, evaluators, and harnesses\(Daiet al\.[2026](https://arxiv.org/html/2608.10626#bib.bib7); Wanget al\.[2026a](https://arxiv.org/html/2608.10626#bib.bib8); Yanget al\.[2026](https://arxiv.org/html/2608.10626#bib.bib9); Chenet al\.[2026](https://arxiv.org/html/2608.10626#bib.bib10)\)\. We instead freeze role\-play and evolve a hierarchical distribution over multi\-turn conditions\.
## 3Dual\-Loop Self\-Evolution
Figure 2:Overview of the proposed framework\. A complete scenario remains uniformly sampled and is paired with a simulator\-side interaction state conditioned on its native support intent\. Trajectories within a rollout group share the same environment\. Continuous emotion reward optimizes the dialogue policy, while thresholded group outcomes update hierarchical intent–state statistics and the subsequent interaction distribution\.Figure[2](https://arxiv.org/html/2608.10626#S3.F2)presents the complete training pipeline\. Our framework organizes multi\-turn empathetic\-dialogue training as two*nested*feedback loops driven by the same verified emotion outcomes\. The*inner loop*optimizes the policy within the interaction distribution currently supplied by the controller; the*outer loop*observes completed group outcomes and revises that distribution as the policy changes\. The two loops are logically nested rather than parallel learners\.
The remainder of this section follows one complete feedback cycle\. Section 3\.1 formalizes scenario\-conditioned rollouts and the nested objectives, while Section 3\.2 constructs the controllable interaction space\. Sections 3\.3–3\.4 describe shared\-condition policy learning and verified group feedback\. Section 3\.5 converts sparse outcomes into robust policy\-relative utility, and Section 3\.6 closes the loop by sampling the next interaction distribution\.
### 3\.1Task Setup and Dual\-Loop Framework
Training begins from a set𝒳\\mathcal\{X\}of complete emotional\-support scenarios\. A scenariox∈𝒳x\\in\\mathcal\{X\}specifies the user’s persona, precipitating event, current dilemma, response tendencies, and hidden support need; an intent maphhassociates it with a support intentc=h\(x\)c=h\(x\)\. We preserve this semantic core throughout training\. To control how the same concern unfolds in dialogue, we additionally choose a simulator\-side interaction statezzfrom a finite space𝒵\\mathcal\{Z\}\. The paire=\(x,z\)e=\(x,z\)is one concrete rollout environment:xxdetermines*what*the user is experiencing, whilezzdetermines*how*the user discloses and reacts\.
For a sampled condition\(x,z\)\(x,z\), a user simulator𝒰\\mathcal\{U\}and the dialogue policyπθ\\pi\_\{\\theta\}interact for at mostHHturns\. Their alternating user and assistant utterances form a trajectory
τ=\(u1,a1,…,uT,aT\)∼P\(τ∣x,z,πθ,𝒰\),\\tau=\(u\_\{1\},a\_\{1\},\\ldots,u\_\{T\},a\_\{T\}\)\\sim P\(\\tau\\mid x,z,\\pi\_\{\\theta\},\\mathcal\{U\}\),whereT≤HT\\leq His the realized number of turns\. After the interaction ends, an emotion verifier evaluates the complete trajectory and returns a scalar outcomeR\(τ\)R\(\\tau\)\. The policy observes only the natural\-language conversation: support\-intent labels, interaction\-state labels, controller statistics, and sampling probabilities remain simulator\-side metadata\. Consequently, improved performance must be expressed through dialogue behavior rather than direct access to a difficulty label\.
At training stagett, the outer loop represents its current curriculum by a conditional distributionpt\(z∣c\)p\_\{t\}\(z\\mid c\)over interaction states for each support intent, initialized uniformly before feedback accumulates\. Sampling from this distribution produces the conditions under which the inner loop optimizes the policy:
maxθ\\displaystyle\\max\_\{\\theta\}𝔼x∼Unif\(𝒳\),z∼pt\(⋅∣h\(x\)\),τ\\displaystyle\\mathbb\{E\}\_\{x\\sim\\mathrm\{Unif\}\(\\mathcal\{X\}\),\\,z\\sim p\_\{t\}\(\\cdot\\mid h\(x\)\),\\,\\tau\}\(1\)\[R\(τ\)\]\.\\displaystyle\\left\[R\(\\tau\)\\right\]\.Meanwhile, verified outcomes update controller historyℋt\\mathcal\{H\}\_\{t\}and hence the next interaction distribution:
ℋt\+1\\displaystyle\\mathcal\{H\}\_\{t\+1\}=Update\(ℋt;𝒟t,Rt\),\\displaystyle=\\operatorname\{Update\}\(\\mathcal\{H\}\_\{t\};\\mathcal\{D\}\_\{t\},R\_\{t\}\),\(2\)pt\+1\\displaystyle p\_\{t\+1\}=Controller\(ℋt\+1\)\.\\displaystyle=\\operatorname\{Controller\}\(\\mathcal\{H\}\_\{t\+1\}\)\.Here𝒟t\\mathcal\{D\}\_\{t\}denotes the completed rollout groups collected at stagett,RtR\_\{t\}their verified outcomes, andℋt\\mathcal\{H\}\_\{t\}the controller’s accumulated evidence\. Updatingℋt\\mathcal\{H\}\_\{t\}changespt\+1p\_\{t\+1\}, which in turn changes the experience used by the next policy updates\. Crucially, the outer\-loop target is policy\-relative interaction utility—the value of allocating another rollout group to a condition under the current policy—rather than a permanent difficulty label attached to a scenario\.
### 3\.2A Controllable Interaction Space
The framework requires interaction states to be behaviorally realizable, compatible with the underlying scenario, hidden from the assistant, and reusable within a rollout group\. We preserve each complete scenario and define three task\-operational axes:*disclosure readiness*controls when details and deeper needs are revealed;*emotional activation*controls the intensity of negative\-affect expression; and*relational trust*controls whether questions and support are accepted\. These axes capture information release, affective expression, and relational feedback rather than clinical traits\.
We discretize the axes at a granularity that remains behaviorally distinguishable and statistically estimable: disclosure has three levels \(delayed, conditional, proactive\), activation has two bounded levels \(moderate, high\), and trust has four levels from skepticism to sustained engagement\. Their Cartesian product gives\|𝒵\|=3×2×4=24\|\\mathcal\{Z\}\|=3\\times 2\\times 4=24interaction states\. A statez=\(d,a,r\)z=\(d,a,r\)is rendered by composing one natural\-language behavioral clause for each axis,ϕ\(z\)=ϕd\(d\)⊕ϕa\(a\)⊕ϕr\(r\)\\phi\(z\)=\\phi\_\{d\}\(d\)\\oplus\\phi\_\{a\}\(a\)\\oplus\\phi\_\{r\}\(r\)forz=\(d,a,r\)z=\(d,a,r\)\. These clauses govern disclosure pace, affective expression, and receptivity, but cannot overwrite the persona, event, or hidden need specified byxx\.
Estimating a separate curriculum value for every scenario–state pair would fragment evidence across many nearly related conditions\. We therefore retain individual scenarios for rollout generation but organize controller evidence by support intent\. Combining intentccwith statezzdefines the reusable unitm=\(c,z\)m=\(c,z\)\. This factorization lets different scenarios with the same support objective contribute evidence about a shared interaction state, while preserving intent\-specific differences\. Scenario identities and intent frequencies remain fixed; only the conditional state distributionpt\(z∣c\)p\_\{t\}\(z\\mid c\)evolves\. Section[4\.4](https://arxiv.org/html/2608.10626#S4.SS4)later separates the contribution of this interaction space from that of adaptive allocation\.
#### Behavioral realization and information boundaries\.
Each state is translated into simulator\-side behavioral guidance rather than exposed as a symbolic label\. Delayed, conditional, and proactive disclosure regulate when relevant facts and deeper needs become available\. Moderate and high activation regulate the intensity of negative affect without inventing an unsupported crisis\. The four trust levels range from rejecting generic reassurance to sustained engagement with situation\-specific support\. Every composed instruction additionally requires the simulator to preserve the original persona, event, and hidden support need, and forbids mentioning axis names, levels, state identifiers, or difficulty labels\.
### 3\.3Group\-Shared Multi\-Turn Policy Learning
For each scenario, the controller selects one statezz, and group\-relative policy optimization drawsτ1:K∼i\.i\.d\.P\(τ∣x,z,πθ,𝒰\)\\tau\_\{1:K\}\\overset\{\\mathrm\{i\.i\.d\.\}\}\{\\sim\}P\(\\tau\\mid x,z,\\pi\_\{\\theta\},\\mathcal\{U\}\)under the shared condition\(x,z\)\(x,z\)\. Holding the environment fixed within a group prevents differences in user behavior from confounding group\-relative advantages, so rollout variation primarily reflects policy and generation stochasticity\.
Abstracting from stabilization terms held fixed across all controlled systems, the method\-level policy reward is the continuous verifier outcome,rkpolicy=R\(τk\)r\_\{k\}^\{\\mathrm\{policy\}\}=R\(\\tau\_\{k\}\)\. GRPO computes group\-relative advantages from these rewards and applies the standard clipped policy objective with KL regularization\(Shaoet al\.[2024](https://arxiv.org/html/2608.10626#bib.bib30); Schulmanet al\.[2017](https://arxiv.org/html/2608.10626#bib.bib29)\)\. Building on this inner optimization, our outer allocation loop reuses the same outcomes to organize subsequent training experience\.
### 3\.4Emotion\-Thresholded Group Feedback
The policy and controller consume different transformations of the same verified outcome\. Given an emotion criterionη\\eta, the controller computes
yk=𝕀\[R\(τk\)≥η\],pm=K−1∑k=1Kyk,y\_\{k\}=\\mathbb\{I\}\[R\(\\tau\_\{k\}\)\\geq\\eta\],\\qquad p\_\{m\}=K^\{\-1\}\\sum\_\{k=1\}^\{K\}y\_\{k\},\(3\)whereyky\_\{k\}records whether trajectorykkpasses the criterion, andpmp\_\{m\}is the pass rate of the shared unitm=\(h\(x\),z\)m=\(h\(x\),z\)in that group\. We use the pass rate because it directly reveals whether the current policy produces both successful and unsuccessful behavior under a controlled condition\. Thresholding affects only outer\-loop allocation: the continuous valueR\(τk\)R\(\\tau\_\{k\}\)remains the inner\-loop reward and therefore retains fine\-grained differences between trajectories\. One complete group contributes one pass\-rate observation, preventing correlated trajectories from masquerading as independently sampled environments\.
### 3\.5Robust Policy\-Relative Utility Estimation
Controller evidence is uneven: some intent–state units are observed repeatedly, whereas others have only a few completed groups\. Trusting a sparse empirical rate can turn an early random success or failure into a persistent sampling bias\. For unit\(c,z\)\(c,z\), letSc,zS\_\{c,z\}be the cumulative sum of group pass rates andNc,zN\_\{c,z\}the number of complete groups\. We stabilize its estimate with a leave\-one\-intent prior formed from the same state in the other support intents:
Sz−c\\displaystyle S\_\{z\}^\{\-c\}=∑c′≠cSc′,z,Nz−c=∑c′≠cNc′,z,\\displaystyle=\\sum\_\{c^\{\\prime\}\\neq c\}S\_\{c^\{\\prime\},z\},\\qquad N\_\{z\}^\{\-c\}=\\sum\_\{c^\{\\prime\}\\neq c\}N\_\{c^\{\\prime\},z\},\(4\)p^z−c\\displaystyle\\hat\{p\}\_\{z\}^\{\-c\}=Sz−c\+α/2Nz−c\+α,p^c,z=Sc,z\+λp^z−cNc,z\+λ\.\\displaystyle=\\frac\{S\_\{z\}^\{\-c\}\+\\alpha/2\}\{N\_\{z\}^\{\-c\}\+\\alpha\},\\qquad\\hat\{p\}\_\{c,z\}=\\frac\{S\_\{c,z\}\+\\lambda\\hat\{p\}\_\{z\}^\{\-c\}\}\{N\_\{c,z\}\+\\lambda\}\.Herep^z−c\\hat\{p\}\_\{z\}^\{\-c\}summarizes how statezzbehaves outside intentcc, andp^c,z\\hat\{p\}\_\{c,z\}is the resulting unit estimate\. The coefficientα\\alphasupplies a neutral prior when pooled evidence is scarce, whileλ\\lambdacontrols how strongly the unit borrows that evidence\. Excludingccavoids counting its observations in both terms\. AsNc,zN\_\{c,z\}grows, the unit’s own history naturally dominates\.
The stabilized pass rate is then converted into an allocation score\. We assign maximal boundary value when success and failure coexist, and add an uncertainty bonus that revisits under\-observed units\(Aueret al\.[2002](https://arxiv.org/html/2608.10626#bib.bib31)\):
Bc,z\\displaystyle B\_\{c,z\}=max\(0,1−2\|p^c,z−12\|\),\\displaystyle=\\max\(0,1\-2\|\\hat\{p\}\_\{c,z\}\-\\tfrac\{1\}\{2\}\|\),\(5\)Uc,z\\displaystyle U\_\{c,z\}=log\(N\+1\)/\(Nc,z\+1\),\\displaystyle=\\sqrt\{\\log\(N\+1\)/\(N\_\{c,z\}\+1\)\},Ac,z\\displaystyle A\_\{c,z\}=max\(Amin,Bc,z\+βUc,z\)\.\\displaystyle=\\max\(A\_\{\\min\},B\_\{c,z\}\+\\beta U\_\{c,z\}\)\.HereBc,zB\_\{c,z\}measures proximity to the current success boundary,Uc,zU\_\{c,z\}expresses uncertainty from limited visits, andAc,zA\_\{c,z\}combines the two with exploration weightβ\\betawhile retaining a minimum scoreAminA\_\{\\min\}\. InUc,zU\_\{c,z\},N=∑c′,z′Nc′,z′N=\\sum\_\{c^\{\\prime\},z^\{\\prime\}\}N\_\{c^\{\\prime\},z^\{\\prime\}\}denotes the total number of completed groups across all units, so that under\-observed units receive a larger exploration bonus\. Nearp^c,z=0\.5\\hat\{p\}\_\{c,z\}=0\.5, successful and unsuccessful trajectories coexist, providing contrasting behavior under the same interaction condition\. Unlike raw reward variance, which is scale\-sensitive and can mix learnability with simulator noise or outliers,Bc,zB\_\{c,z\}is bounded and aligned with the verified success criterion\. The controller can therefore estimate policy\-relative interaction utility directly from outcomes already produced by training\.
The controller first samples states uniformly to establish evidence, then activates feedback\-guided allocation\. Incomplete groups and groups with non\-finite outcomes do not update its history\. Thereafter,
pt\(z∣c\)=\(1−ϵ\)Ac,z1/γ∑z′∈𝒵Ac,z′1/γ\+ϵ1\|𝒵\|,p\_\{t\}\(z\\mid c\)=\(1\-\\epsilon\)\\frac\{A\_\{c,z\}^\{1/\\gamma\}\}\{\\sum\_\{z^\{\\prime\}\\in\\mathcal\{Z\}\}A\_\{c,z^\{\\prime\}\}^\{1/\\gamma\}\}\+\\epsilon\\frac\{1\}\{\|\\mathcal\{Z\}\|\},\(6\)whereγ\\gammais a sampling temperature andϵ\\epsilonreserves uniform rehearsal\. Thus, no state is permanently removed, and conditions that were previously misestimated or excessive can re\-enter training as the policy develops\.
### 3\.6Coupled Self\-Evolution
Every complete group thus closes both feedback paths: continuous outcomes update the policy, thresholded group outcomes update the allocation, and the revised distribution generates the next round of experience\. Because the simulator, scenario semantics, and verifier stay frozen, experience evolves without a drifting role\-playing target\. Algorithm[1](https://arxiv.org/html/2608.10626#alg1)summarizes the end\-to\-end process\.
Algorithm 1Dual\-Loop Self\-Evolution with Verifiable Emotion Feedback0:scenarios
𝒳\\mathcal\{X\}with intent map
hh, interaction states
𝒵\\mathcal\{Z\}, policy
πθ\\pi\_\{\\theta\}, group size
KK
1:foreach training batchdo
2:foreach uniformly sampled scenario
xxdo
3:set
c=h\(x\)c=h\(x\)and sample
zzuniformly during initialization, otherwise from
pt\(z∣c\)p\_\{t\}\(z\\mid c\)
4:generate
KKtrajectories under the shared environment
\(x,z\)\(x,z\)
5:score each trajectory with continuous emotion outcome
RkR\_\{k\}
6:update
πθ\\pi\_\{\\theta\}using group\-relative advantages from
R1:KR\_\{1:K\}
7:set
yk=𝕀\[Rk≥η\]y\_\{k\}=\\mathbb\{I\}\[R\_\{k\}\\geq\\eta\]and
pm=K−1∑kykp\_\{m\}=K^\{\-1\}\\sum\_\{k\}y\_\{k\}
8:if the group is complete and finite, accumulate
pmp\_\{m\}for state
zzand unit
\(c,z\)\(c,z\)
9:endfor
10:endfor
## 4Experiments
### 4\.1Experimental Setup
Our experiments answer two questions: whether capability\-aligned allocation improves empathetic dialogue under a fixed rollout budget, and whether this improvement comes from adaptive allocation rather than a larger state space or generic sample prioritization\. The 500 training scenarios, development split, and 100\-scenario SAGE test set are instance\-disjoint\. Each training scenario retains one of eight native support\-intent labels and is composed online with one of 24 interaction states, yielding 192 controller units\. Development data is used for configuration and checkpoint selection; test instances never enter policy training, controller updates, or model selection\.
We use complementary automatic, interactive, and human evaluations\.SAGEmeasures final user emotion and success/failure over 100 multi\-turn scenarios\(Zhanget al\.[2026](https://arxiv.org/html/2608.10626#bib.bib5)\)\.ESC\-Evalevaluates 331 fixed\-role interactions with the official InternLM2 and ESC\-RANK protocol\(Zhaoet al\.[2024](https://arxiv.org/html/2608.10626#bib.bib48)\)\.EIBenchcovers 213 held\-out scenarios spanning Support, Defense, Repair, and Charm\(Zhuet al\.[2026](https://arxiv.org/html/2608.10626#bib.bib16)\)\.ESConvevaluates 2,895 reference responses using strategy accuracy, BLEU\-2, ROUGE\-L, BERTScore, and Distinct\-2 \(all multiplied by 100\)\. To separate training feedback from evaluation, DeepSeek\-V3 is used only for training simulation and verification; transfer is measured independently by InternLM2\+ESC\-RANK in ESC\-Eval, Qwen3\-Max in EIBench, reference\-based ESConv metrics, and blinded human ratings\. RLVER and ours are independently trained three times; remaining fixed checkpoints are evaluated three times under shared seeds\. Table[1](https://arxiv.org/html/2608.10626#S4.T1)reports sample standard deviations for aggregate automatic metrics\(Liuet al\.[2023](https://arxiv.org/html/2608.10626#bib.bib35); Zhanget al\.[2024a](https://arxiv.org/html/2608.10626#bib.bib36)\)\.
We also conduct a blinded human evaluation\. Three annotators independently rate every system dialogue for 100 randomly sampled held\-out scenarios, with dialogue and system order randomized\. Ratings use a 0–4 overall\-support scale covering empathy, contextual relevance, helpfulness, and coherence, and are averaged across annotators and then scenarios\. Ordinal Krippendorff’sα\\alphais 0\.78\.
We compare all systems under a controlled protocol with the same Qwen3\-8B initialization, frozen simulator and verifier, data, rollout budget, group size, and optimization settings\.Qwen3\-8Bapplies no RL;RLVERperforms uniform emotion\-reward RL;Interaction State Onlyadds the 24 states without adaptive allocation;VCRLprioritizes within\-group reward variance\(Jianget al\.[2025](https://arxiv.org/html/2608.10626#bib.bib14)\); andOursuses the complete controller\. All RL systems therefore receive the same GRPO budget, isolating how training interactions are allocated\.
### 4\.2Main Results
Table 1:Protocol\-matched Qwen3\-8B results across complementary automatic, interactive, and human evaluations\. Under SAGE, “Avg\.” is the Overall score \(mean final user emotion\), i\.e\., the value reported as “SAGE Overall” in the text and in Table[5](https://arxiv.org/html/2608.10626#S4.T5)\.Against protocol\-matched controls, dual\-loop self\-evolution raises mean SAGE Overall from 72\.01 under uniform emotion\-reward RL to 79\.24 across three independent training runs\. Its weakest run \(78\.11\) still exceeds the strongest uniform run \(73\.76\)\. Interaction states alone reach 69\.42 and variance\-based allocation reaches 68\.51, showing that neither a larger condition space nor generic reward dispersion explains the gain\.
SAGE directly evaluates the emotional consequence of a complete multi\-turn interaction: Avg\. is final user emotion, while Succ\. and Fail\. summarize trajectories reaching the official high\- and low\-emotion regions\. Our framework improves Avg\. and success while reducing failure to 11\.33%, indicating more reliable long\-horizon emotional recovery rather than an isolated response\-quality gain\. The particularly large SAGE margin is consistent with the method’s target: reallocating complete interaction conditions as the policy’s multi\-turn competence changes\.
The remaining benchmarks test whether this gain transfers beyond the training\-facing emotion outcome\. Ours leads all five ESConv metrics, including a Distinct\-2 increase from 28\.18 to 33\.58 over RLVER\. Better strategy accuracy, reference overlap, semantic similarity, and lexical diversity show that stronger interactive outcomes are accompanied by broader response quality rather than repetitive reassurance\. Ours also attains the highest ESC\-Eval aggregate score \(2\.55\) and improves EIBench from−9\.79\-9\.79to−7\.55\-7\.55, extending the gain to fixed\-role dialogue quality and wider emotional\-intelligence behavior\. Crucially, adding interaction states without adaptive allocation lowers EIBench from the base model’s−22\.40\-22\.40to−40\.60\-40\.60, whereas the complete framework reaches−7\.55\-7\.55\. This contrast shows that state diversity is not inherently beneficial: the controller must convert it into policy\-relevant training experience\. Human overall quality likewise rises from 3\.0 to 3\.5\. Together, the results connect the framework’s central advantage in multi\-turn emotional recovery to complementary improvements in strategy, diversity, and interaction quality\.
### 4\.3Run\-Level Robustness and Transfer Detail
Table[2](https://arxiv.org/html/2608.10626#S4.T2)exposes the individual SAGE results underlying the training\-level statistics in Table[1](https://arxiv.org/html/2608.10626#S4.T1)\. The weakest complete\-framework run \(78\.11\) remains above the strongest RLVER run \(73\.76\), so the aggregate margin is present throughout the observed independent\-run range rather than being determined by one favorable optimization run\.
Table 2:SAGE Overall across three independent training runs\. SD is the sample standard deviation across runs\.The dimension\-level ESC\-Eval results in Table[3](https://arxiv.org/html/2608.10626#S4.T3)further show that the aggregate gain is broad rather than concentrated in one rubric\. The complete framework improves empathy and suggestion quality most strongly over RLVER while also maintaining or improving fluency, diversity, humanness, technical quality, and overall interaction quality\.
Table 3:Dimension\-level results under the official ESC\-Eval protocol\. All dimensions use the original 0–4 scale; Macro is the arithmetic mean of the seven columns\.Table[4](https://arxiv.org/html/2608.10626#S4.T4)localizes EIBench transfer\. Ours obtains the strongest aggregate reward and the best Support, Repair, and Charm values among the matched systems\. Interaction states alone enlarge behavioral variation without deciding where a fixed rollout budget is useful, whereas adaptive allocation converts that diversity into stronger transfer\.
Table 4:EIBench raw rewards by category under one shared 213\-scenario evaluation protocol\. Higher is better\.
### 4\.4Ablations and Training Analysis
We additionally use two diagnostic controllers to isolate allocation design choices\.*Scenario\-only allocation*maintains pass\-rate statistics for individual scenario identities without the intent–state factorization\. The*variance\-gated pass\-rate controller*retains our pass\-rate priority but admits only high\-variance groups when updating controller history; all groups still participate in policy optimization\.
†Last stable checkpoint before KL instability\.
Table 5:Single\-run component ablations on SAGE under a matched training and evaluation protocol\.Δ\\Deltais measured against the complete\-framework run \(78\.11\); Table[1](https://arxiv.org/html/2608.10626#S4.T1)separately reports three\-run statistics for the complete framework\.Table[5](https://arxiv.org/html/2608.10626#S4.T5)maps each intervention to a specific part of the framework\. Interaction State Only removes the adaptive outer loop while retaining the 24\-state environment representation; the 8\.69\-point drop shows that additional diversity can dilute a fixed budget when it is not allocated according to policy feedback\. Scenario\-only allocation removes the intent–state factorization and instead estimates individual scenario identities\. Its 70\.03 score indicates that sparse identity\-level histories cannot effectively reuse evidence across semantically related interactions\.
The utility\-estimation ablations examine how the controller turns outcomes into stable evidence\. Replacing hard group outcomes with a soft mapping reduces Overall by 1\.59 points\. Removing hierarchical sharing costs 3\.04 points because sparsely visited units can no longer borrow evidence from the same state under related support intents\. The variance\-gated controller preserves pass\-rate priority but admits only high\-variance groups into controller history; its KL growth and 61\.45 score highlight the stability advantage of bounded, criterion\-aligned pass\-rate evidence over raw reward dispersion\.
The remaining variants test exploration\. Without the uncertainty bonus, Overall falls by 4\.19 points, showing that exploiting current estimates alone can neglect conditions that remain poorly understood\. Removing uniform rehearsal has a smaller 1\.08\-point effect, consistent with its role as a safety floor that lets previously mastered, excessive, or misestimated conditions re\-enter training\. A matched sensitivity study obtains 77\.46, 78\.11, and 77\.72 forη=40,50,60\\eta=40,50,60, respectively; the controller therefore does not depend on a narrowly tuned success criterion\. Each completed group requires only constant\-time statistic updates and normalization over 24 state scores; the outer loop adds no trajectories, model passes, backward passes, or simulator/verifier calls\.
### 4\.5Validating the Interaction\-State Space
The outer loop relies on two properties of the interaction\-state space: requested states should produce observable differences in user behavior, and those differences should not alter the underlying persona, event, or support need\. We test behavioral separability and semantic preservation directly\.
Figure 3:Behavioral validation of the interaction\-state space: row\-normalized confusion matrices \(%\) of requested vs\. recovered \(a\) disclosure readiness, \(b\) emotional activation, and \(c\) relational trust\.Figure 4:Representative trajectories under the same SAGE profile and initial emotion\. The base policy remains generic, RLVER \(uniform emotion\-reward RL\) echoes escalating hostility, and ours identifies the hidden need for recognition and guides self\-reflection\. User turns diverge as the frozen simulator reacts to each policy—a trajectory\-level rather than turn\-aligned comparison\.For behavioral separability, we audit 120 multi\-turn trajectories, with five examples for each of the 24 states\. A held\-out Qwen3\-Max judge sees only the resulting dialogue and attempts to recover the requested level of each state dimension\. Because the judge never observes the control labels, successful recovery indicates that the requested state is expressed in dialogue behavior rather than remaining as prompt metadata\. Figure[3](https://arxiv.org/html/2608.10626#S4.F3)shows accuracies of 86\.7%, 92\.5%, and 82\.5% for disclosure, activation, and trust, with corresponding macro\-F1 scores of 86\.7%, 92\.5%, and 82\.1%\. All three dimensions are recovered correctly in 68\.3% of trajectories, a substantially stricter test of joint realization\.
We then verify that behavioral control preserves scenario meaning\. Human reviewers inspect 30 cases; their decisions agree with the automatic judge in 90\.0% of cases, with Cohen’sκ=0\.89\\kappa=0\.89\. Across the full audit, 92\.5% of trajectories retain the original persona, event, and hidden need\. The interaction states therefore induce recognizable differences in disclosure, emotion, and trust while leaving the support problem intact\. This is precisely the structure required by the outer loop: the controller reallocates distinct but semantically comparable multi\-turn experiences, not nominal prompt labels\.
### 4\.6Qualitative Analysis
Figure[4](https://arxiv.org/html/2608.10626#S4.F4)localizes the aggregate gain within one complete interaction\. Starting from the same profile and initial emotion, the policies encounter resistance and induce different subsequent user trajectories\. The base policy repeats broad relationship advice and ends at 45, showing limited adaptation to information disclosed across turns\. RLVER endorses the user’s hostile framing; this emotional echoing validates the immediate complaint but intensifies the conflict, and emotion falls to 10\. Our policy instead grounds its response in the user’s creative labor, identifies the hidden need for recognition, and reframes the conflict around effort and artistic identity before guiding self\-reflection\. Emotion consequently rises to 98\.
The case clarifies the capability reflected by the main results\. The benefit is not simply more positive wording: the learned policy integrates disclosures across turns, distinguishes validation from escalation, and moves from surface affect toward the latent support need\. This behavior illustrates the value of repeatedly allocating training to conditions where the current policy exhibits mixed success, rather than spending the same budget uniformly on already mastered or currently uninformative interactions\.
## 5Discussion and Conclusion
Experiments establish that dual\-loop self\-evolution converts a fixed rollout budget into higher\-value multi\-turn experience\. Its largest gain appears on SAGE, which directly measures complete emotional recovery, while consistent improvements on ESConv, ESC\-Eval, EIBench, and human ratings extend this advantage to strategy choice, diversity, and interaction quality\. Ablations and state validation further connect the gain to factorized evidence sharing, uncertainty\-guided exploration, and behaviorally realizable interaction states\.
We introduced a nested framework in which verified emotion feedback serves two coordinated roles: continuous outcomes improve how the dialogue policy responds, while group\-level outcome patterns determine what it should practice next\. By adapting experience rather than increasing rollout count, the framework makes existing feedback an additional source of training efficiency\. More broadly, verifiable feedback can supervise both policy behavior and the organization of the experience from which that behavior is learned\.
## References
- A\. Ahmadian, C\. Cremer, M\. Gallé, M\. Fadaee, J\. Kreutzer, O\. Pietquin, A\. Üstün, and S\. Hooker \(2024\)Back to basics: revisiting REINFORCE\-style optimization for learning from human feedback in LLMs\.InProceedings of ACL,pp\. 12248–12267\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.662),[Link](https://aclanthology.org/2024.acl-long.662/)Cited by:[§2](https://arxiv.org/html/2608.10626#S2.SS0.SSS0.Px2.p1.1)\.
- Finite\-time analysis of the multiarmed bandit problem\.Machine learning47\(2\),pp\. 235–256\.Cited by:[§3\.5](https://arxiv.org/html/2608.10626#S3.SS5.p2.1)\.
- Y\. Bengio, J\. Louradour, R\. Collobert, and J\. Weston \(2009\)Curriculum learning\.InProceedings of ICML,pp\. 41–48\.Cited by:[§1](https://arxiv.org/html/2608.10626#S1.p4.1),[§2](https://arxiv.org/html/2608.10626#S2.SS0.SSS0.Px2.p1.1)\.
- G\. Chen, Y\. Shi, Y\. Li, B\. Li, X\. Xu, H\. Wei, S\. Ni, M\. Yang, and J\. Ye \(2026\)EvoTrainer: co\-evolving llm policies and training harnesses for autonomous agentic reinforcement learning\.External Links:2606\.03108,[Link](https://arxiv.org/abs/2606.03108)Cited by:[§2](https://arxiv.org/html/2608.10626#S2.SS0.SSS0.Px2.p1.1)\.
- Z\. Chen, Y\. Deng, H\. Yuan, K\. Ji, and Q\. Gu \(2024\)Self\-play fine\-tuning converts weak language models to strong language models\.InProceedings of ICML,Vol\.235,pp\. 6621–6642\.External Links:[Link](https://proceedings.mlr.press/v235/chen24j.html)Cited by:[§1](https://arxiv.org/html/2608.10626#S1.p5.1),[§2](https://arxiv.org/html/2608.10626#S2.SS0.SSS0.Px2.p1.1)\.
- Y\. Cheng, W\. Liu, W\. Li, J\. Wang, R\. Zhao, B\. Liu, X\. Liang, and Y\. Zheng \(2022\)Improving multi\-turn emotional support dialogue generation with lookahead strategy planning\.InProceedings of EMNLP,pp\. 3014–3026\.External Links:[Document](https://dx.doi.org/10.18653/v1/2022.emnlp-main.195)Cited by:[§1](https://arxiv.org/html/2608.10626#S1.p1.1),[§1](https://arxiv.org/html/2608.10626#S1.p2.1),[§2](https://arxiv.org/html/2608.10626#S2.SS0.SSS0.Px1.p1.1)\.
- Y\. Dai, N\. Gao, W\. Zhang, J\. Wang, Luozichen, R\. Wu, J\. Wang, and C\. Wang \(2026\)SEAD: self\-evolving agent for multi\-turn service dialogue\.InFindings of ACL,pp\. 3674–3684\.External Links:[Document](https://dx.doi.org/10.18653/v1/2026.findings-acl.180),[Link](https://aclanthology.org/2026.findings-acl.180/)Cited by:[§1](https://arxiv.org/html/2608.10626#S1.p5.1),[§2](https://arxiv.org/html/2608.10626#S2.SS0.SSS0.Px2.p1.1)\.
- Y\. Dou, M\. Galley, B\. Peng, C\. Kedzie, W\. Cai, A\. Ritter, C\. Quirk, W\. Xu, and J\. Gao \(2025\)SimulatorArena: are user simulators reliable proxies for multi\-turn evaluation of AI assistants?\.InProceedings of EMNLP,pp\. 35212–35290\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-main.1786)Cited by:[§1](https://arxiv.org/html/2608.10626#S1.p2.1),[§1](https://arxiv.org/html/2608.10626#S1.p5.1),[§2](https://arxiv.org/html/2608.10626#S2.SS0.SSS0.Px1.p1.1)\.
- A\. Graves, M\. G\. Bellemare, J\. Menick, R\. Munos, and K\. Kavukcuoglu \(2017\)Automated curriculum learning for neural networks\.InProceedings of ICML,Vol\.70,pp\. 1311–1320\.External Links:[Link](https://proceedings.mlr.press/v70/graves17a.html)Cited by:[§1](https://arxiv.org/html/2608.10626#S1.p4.1)\.
- J\. Gromada, A\. Kasicka, E\. Komkowska, L\. Krajewski, N\. Krawczyk, M\. Veyret, B\. Przybyl, L\. M\. Rojas\-Barahona, and M\. K\. Szczerbak \(2025\)Evaluating conversational agents with persona\-driven user simulations based on large language models: a sales bot case study\.InProceedings of the EMNLP Industry Track,pp\. 230–245\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-industry.16),[Link](https://aclanthology.org/2025.emnlp-industry.16/)Cited by:[§1](https://arxiv.org/html/2608.10626#S1.p2.1),[§2](https://arxiv.org/html/2608.10626#S2.SS0.SSS0.Px1.p1.1)\.
- G\. Jiang, W\. Feng, G\. Quan, C\. Hao, Y\. Zhang, G\. Liu, and H\. Wang \(2025\)VCRL: variance\-based curriculum reinforcement learning for large language models\.External Links:2509\.19803,[Link](https://arxiv.org/abs/2509.19803)Cited by:[§1](https://arxiv.org/html/2608.10626#S1.p4.1),[§2](https://arxiv.org/html/2608.10626#S2.SS0.SSS0.Px2.p1.1),[§4\.1](https://arxiv.org/html/2608.10626#S4.SS1.p4.1)\.
- L\. Jiang, D\. Meng, Q\. Zhao, S\. Shan, and A\. G\. Hauptmann \(2015\)Self\-paced curriculum learning\.InProceedings of AAAI,Vol\.29,pp\. 2694–2700\.External Links:[Document](https://dx.doi.org/10.1609/aaai.v29i1.9608),[Link](https://ojs.aaai.org/index.php/AAAI/article/view/9608)Cited by:[§2](https://arxiv.org/html/2608.10626#S2.SS0.SSS0.Px2.p1.1)\.
- M\. Jiang, E\. Grefenstette, and T\. Rocktäschel \(2021\)Prioritized level replay\.InProceedings of ICML,pp\. 4940–4950\.Cited by:[§1](https://arxiv.org/html/2608.10626#S1.p4.1),[§2](https://arxiv.org/html/2608.10626#S2.SS0.SSS0.Px2.p1.1)\.
- Z\. Jiang, J\. Han, T\. Li, X\. Wang, S\. Jiang, X\. Meng, J\. Wei, J\. Liang, and Y\. Xiao \(2026\)Difficulty is not enough: curriculum learning for llms fine\-tuning must consider utility\.InProceedings of AAAI,Vol\.40,pp\. 31365–31373\.External Links:[Document](https://dx.doi.org/10.1609/aaai.v40i37.40400),[Link](https://ojs.aaai.org/index.php/AAAI/article/view/40400)Cited by:[§2](https://arxiv.org/html/2608.10626#S2.SS0.SSS0.Px2.p1.1)\.
- D\. Kang, S\. Kim, T\. Kwon, S\. Moon, H\. Cho, Y\. Yu, D\. Lee, and J\. Yeo \(2024\)Can large language models be good emotional supporter? mitigating preference bias on emotional support conversation\.InProceedings of ACL,pp\. 15232–15261\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.813),[Link](https://aclanthology.org/2024.acl-long.813/)Cited by:[§2](https://arxiv.org/html/2608.10626#S2.SS0.SSS0.Px1.p1.1)\.
- J\. Kim, C\. Mok, J\. Lee, H\. S\. Kim, and Y\. Jo \(2025\)Dialogue systems for emotional support via value reinforcement\.InProceedings of ACL,pp\. 28733–28766\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.1395),[Link](https://aclanthology.org/2025.acl-long.1395/)Cited by:[§2](https://arxiv.org/html/2608.10626#S2.SS0.SSS0.Px1.p1.1)\.
- M\. P\. Kumar, B\. Packer, and D\. Koller \(2010\)Self\-paced learning for latent variable models\.InProceedings of NeurIPS,Vol\.23\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2010/file/e57c6b956a6521b28495f2886ca0977a-Paper.pdf)Cited by:[§1](https://arxiv.org/html/2608.10626#S1.p4.1),[§2](https://arxiv.org/html/2608.10626#S2.SS0.SSS0.Px2.p1.1)\.
- R\. Li, H\. Huang, F\. Wei, F\. Xiong, Y\. Wang, and X\. Chu \(2026\)AdaCuRL: adaptive curriculum reinforcement learning with invalid sample mitigation and historical revisiting\.InProceedings of AAAI,Vol\.40,pp\. 23123–23131\.External Links:[Document](https://dx.doi.org/10.1609/aaai.v40i27.39479)Cited by:[§2](https://arxiv.org/html/2608.10626#S2.SS0.SSS0.Px2.p1.1)\.
- Z\. Lin, A\. Madotto, J\. Shin, P\. Xu, and P\. Fung \(2019\)MoEL: mixture of empathetic listeners\.InProceedings of EMNLP,pp\. 121–132\.External Links:[Document](https://dx.doi.org/10.18653/v1/D19-1012),[Link](https://aclanthology.org/D19-1012/)Cited by:[§2](https://arxiv.org/html/2608.10626#S2.SS0.SSS0.Px1.p1.1)\.
- S\. Liu, C\. Zheng, O\. Demasi, S\. Sabour, Y\. Li, Z\. Yu, Y\. Jiang, and M\. Huang \(2021\)Towards emotional support dialog systems\.InProceedings of ACL,pp\. 3469–3483\.Cited by:[§1](https://arxiv.org/html/2608.10626#S1.p1.1),[§1](https://arxiv.org/html/2608.10626#S1.p2.1),[§2](https://arxiv.org/html/2608.10626#S2.SS0.SSS0.Px1.p1.1)\.
- W\. Liu, L\. Huo, Y\. Jing, X\. Zhang, and J\. Xie \(2026\)MRACL: multi\-reward space guided adaptive curriculum reinforcement learning for llms\.InProceedings of AAAI,Vol\.40,pp\. 37663–37672\.External Links:[Document](https://dx.doi.org/10.1609/aaai.v40i44.41101),[Link](https://ojs.aaai.org/index.php/AAAI/article/view/41101)Cited by:[§2](https://arxiv.org/html/2608.10626#S2.SS0.SSS0.Px2.p1.1)\.
- Y\. Liu, D\. Iter, Y\. Xu, S\. Wang, R\. Xu, and C\. Zhu \(2023\)G\-Eval: nlg evaluation using gpt\-4 with better human alignment\.InProceedings of EMNLP,pp\. 2511–2522\.External Links:[Document](https://dx.doi.org/10.18653/v1/2023.emnlp-main.153),[Link](https://aclanthology.org/2023.emnlp-main.153/)Cited by:[§4\.1](https://arxiv.org/html/2608.10626#S4.SS1.p2.1)\.
- N\. Majumder, P\. Hong, S\. Peng, J\. Lu, D\. Ghosal, A\. Gelbukh, R\. Mihalcea, and S\. Poria \(2020\)MIME: MIMicking emotions for empathetic response generation\.InProceedings of EMNLP,pp\. 8968–8979\.External Links:[Document](https://dx.doi.org/10.18653/v1/2020.emnlp-main.721),[Link](https://aclanthology.org/2020.emnlp-main.721/)Cited by:[§2](https://arxiv.org/html/2608.10626#S2.SS0.SSS0.Px1.p1.1)\.
- H\. T\. Nguyen, B\. Nguyen, W\. Ma, Y\. Zhao, R\. She, and V\. A\. Nguyen \(2026\)Adaptive rollout allocation for online reinforcement learning with verifiable rewards\.InProceedings of ICLR,External Links:[Link](https://openreview.net/forum?id=Z5sWYACAop)Cited by:[§2](https://arxiv.org/html/2608.10626#S2.SS0.SSS0.Px2.p1.1)\.
- H\. Rashkin, E\. M\. Smith, M\. Li, and Y\. Boureau \(2019\)Towards empathetic open\-domain conversation models: a new benchmark and dataset\.InProceedings of ACL,pp\. 5370–5381\.External Links:[Document](https://dx.doi.org/10.18653/v1/P19-1534),[Link](https://aclanthology.org/P19-1534/)Cited by:[§1](https://arxiv.org/html/2608.10626#S1.p1.1),[§2](https://arxiv.org/html/2608.10626#S2.SS0.SSS0.Px1.p1.1)\.
- T\. Schaul, J\. Quan, I\. Antonoglou, and D\. Silver \(2016\)Prioritized experience replay\.InProceedings of ICLR,External Links:[Link](https://arxiv.org/abs/1511.05952)Cited by:[§1](https://arxiv.org/html/2608.10626#S1.p4.1),[§2](https://arxiv.org/html/2608.10626#S2.SS0.SSS0.Px2.p1.1)\.
- J\. Schulman, F\. Wolski, P\. Dhariwal, A\. Radford, and O\. Klimov \(2017\)Proximal policy optimization algorithms\.External Links:1707\.06347,[Link](https://arxiv.org/abs/1707.06347)Cited by:[§3\.3](https://arxiv.org/html/2608.10626#S3.SS3.p2.1)\.
- Z\. Shao, P\. Wang, Q\. Zhu, R\. Xu, J\. Song, X\. Bi, H\. Zhang, M\. Zhang, Y\. K\. Li, Y\. Wu, and D\. Guo \(2024\)DeepSeekMath: pushing the limits of mathematical reasoning in open language models\.External Links:2402\.03300,[Link](https://arxiv.org/abs/2402.03300)Cited by:[§3\.3](https://arxiv.org/html/2608.10626#S3.SS3.p2.1)\.
- T\. Shi, Y\. Wu, L\. Song, T\. Zhou, and J\. Zhao \(2026\)Efficient reinforcement finetuning via adaptive curriculum learning\.Transactions on Machine Learning Research\.External Links:[Link](https://openreview.net/forum?id=UEhpyq41b9)Cited by:[§1](https://arxiv.org/html/2608.10626#S1.p4.1),[§2](https://arxiv.org/html/2608.10626#S2.SS0.SSS0.Px2.p1.1)\.
- Y\. Shi, J\. Hao, and F\. Kong \(2025\)Beyond coarse labels: fine\-grained problem augmentation and multi\-dimensional feedback for emotional support conversation\.InFindings of EMNLP,pp\. 1634–1647\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.findings-emnlp.86),[Link](https://aclanthology.org/2025.findings-emnlp.86/)Cited by:[§2](https://arxiv.org/html/2608.10626#S2.SS0.SSS0.Px1.p1.1)\.
- C\. Wan, M\. Labeau, and C\. Clavel \(2025\)EmoDynamiX: emotional support dialogue strategy prediction by modelling mixed emotions and discourse dynamics\.InProceedings of NAACL,pp\. 1678–1695\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.naacl-long.81),[Link](https://aclanthology.org/2025.naacl-long.81/)Cited by:[§1](https://arxiv.org/html/2608.10626#S1.p1.1),[§2](https://arxiv.org/html/2608.10626#S2.SS0.SSS0.Px1.p1.1)\.
- A\. Wang, Y\. Yan, N\. Zhou, Z\. Lu, W\. Lu, J\. Xiao, Y\. Zhuang, and Y\. Shen \(2026a\)Code\-A1: adversarial evolving of code LLM and test LLM via reinforcement learning\.External Links:2603\.15611,[Link](https://arxiv.org/abs/2603.15611)Cited by:[§2](https://arxiv.org/html/2608.10626#S2.SS0.SSS0.Px2.p1.1)\.
- F\. Wang, X\. Shen, J\. Yu, and R\. Xia \(2025\)Flexible thinking for multimodal emotional support conversation via reinforcement learning\.InFindings of EMNLP,pp\. 1341–1356\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.findings-emnlp.70),[Link](https://aclanthology.org/2025.findings-emnlp.70/)Cited by:[§2](https://arxiv.org/html/2608.10626#S2.SS0.SSS0.Px1.p1.1)\.
- P\. Wang, R\. Ma, B\. Zhang, X\. Chen, Z\. He, K\. Luo, Q\. Lv, Q\. Jiang, Z\. Xie, S\. Wang, C\. Li, Y\. Li, F\. Ye, J\. Li, Y\. Yang, J\. Li, Z\. Tu, and X\. Li \(2026b\)RLVER: reinforcement learning with verifiable emotion rewards for empathetic agents\.InProceedings of ICLR,External Links:[Link](https://openreview.net/forum?id=P7wBg0vPTh)Cited by:[§1](https://arxiv.org/html/2608.10626#S1.p2.1),[§1](https://arxiv.org/html/2608.10626#S1.p3.1),[§2](https://arxiv.org/html/2608.10626#S2.SS0.SSS0.Px1.p1.1)\.
- H\. Wu, Z\. Tian, Z\. Huang, T\. Hu, L\. Qiao, Y\. Gao, F\. Liu, and D\. Li \(2026\)Emotion trajectory\-aware retrieval for markov\-driven emotion anticipation in LLM\-based emotional support conversation\.InFindings of ACL,pp\. 42887–42905\.External Links:[Document](https://dx.doi.org/10.18653/v1/2026.findings-acl.2127),[Link](https://aclanthology.org/2026.findings-acl.2127/)Cited by:[§1](https://arxiv.org/html/2608.10626#S1.p1.1)\.
- S\. Yang, Z\. Ma, T\. Huang, Y\. Hu, Y\. Wang, and X\. Chu \(2026\)CoEvolve: training LLM agents via agent\-data mutual evolution\.InProceedings of ACL,pp\. 23015–23036\.External Links:[Document](https://dx.doi.org/10.18653/v1/2026.acl-long.1055),[Link](https://aclanthology.org/2026.acl-long.1055/)Cited by:[§2](https://arxiv.org/html/2608.10626#S2.SS0.SSS0.Px2.p1.1)\.
- J\. Ye, L\. Xiang, Y\. Zhang, and C\. Zong \(2025\)From generic empathy to personalized emotional support: a self\-evolution framework for user preference alignment\.InFindings of EMNLP,pp\. 18826–18853\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.findings-emnlp.1024),[Link](https://aclanthology.org/2025.findings-emnlp.1024/)Cited by:[§1](https://arxiv.org/html/2608.10626#S1.p2.1),[§2](https://arxiv.org/html/2608.10626#S2.SS0.SSS0.Px1.p1.1)\.
- S\. Yoon, Z\. He, J\. Echterhoff, and J\. McAuley \(2024\)Evaluating large language models as generative user simulators for conversational recommendation\.InProceedings of NAACL,pp\. 1490–1504\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.naacl-long.83)Cited by:[§1](https://arxiv.org/html/2608.10626#S1.p2.1),[§1](https://arxiv.org/html/2608.10626#S1.p5.1),[§2](https://arxiv.org/html/2608.10626#S2.SS0.SSS0.Px1.p1.1)\.
- W\. Yuan, R\. Y\. Pang, K\. Cho, S\. Sukhbaatar, J\. Xu, and J\. Weston \(2024\)Self\-rewarding language models\.InProceedings of ICML,Vol\.235,pp\. 57905–57923\.External Links:[Link](https://proceedings.mlr.press/v235/yuan24d.html)Cited by:[§2](https://arxiv.org/html/2608.10626#S2.SS0.SSS0.Px2.p1.1)\.
- Y\. Zeng, Z\. Sun, B\. Ji, E\. Min, H\. Cai, S\. Wang, D\. Yin, H\. Zhang, X\. Chen, and J\. Wang \(2026\)CurES: from gradient analysis to efficient curriculum learning for reasoning llms\.InProceedings of ICLR,External Links:[Link](https://openreview.net/forum?id=QXrZ0Y3yGJ)Cited by:[§2](https://arxiv.org/html/2608.10626#S2.SS0.SSS0.Px2.p1.1)\.
- B\. Zhang, R\. Ma, Q\. Jiang, P\. Wang, J\. Chen, Z\. Xie, X\. Chen, Y\. Wang, F\. Ye, J\. Li, Y\. Yang, Z\. Tu, and X\. Li \(2026\)Sentient agent as a judge: evaluating higher\-order social cognition in large language models\.InFindings of ACL,pp\. 38196–38224\.External Links:[Document](https://dx.doi.org/10.18653/v1/2026.findings-acl.1905),[Link](https://aclanthology.org/2026.findings-acl.1905/)Cited by:[§1](https://arxiv.org/html/2608.10626#S1.p1.1),[§1](https://arxiv.org/html/2608.10626#S1.p2.1),[§2](https://arxiv.org/html/2608.10626#S2.SS0.SSS0.Px1.p1.1),[§4\.1](https://arxiv.org/html/2608.10626#S4.SS1.p2.1)\.
- C\. Zhang, L\. F\. D’Haro, Y\. Chen, M\. Zhang, and H\. Li \(2024a\)A comprehensive analysis of the effectiveness of large language models as automatic dialogue evaluators\.InProceedings of AAAI,Vol\.38,pp\. 19515–19524\.External Links:[Document](https://dx.doi.org/10.1609/aaai.v38i17.29923),[Link](https://ojs.aaai.org/index.php/AAAI/article/view/29923)Cited by:[§4\.1](https://arxiv.org/html/2608.10626#S4.SS1.p2.1)\.
- T\. Zhang, X\. Zhang, J\. Zhao, L\. Zhou, and Q\. Jin \(2024b\)ESCoT: towards interpretable emotional support dialogue systems\.InProceedings of ACL,pp\. 13395–13412\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.723)Cited by:[§2](https://arxiv.org/html/2608.10626#S2.SS0.SSS0.Px1.p1.1)\.
- H\. Zhao, L\. Li, S\. Chen, S\. Kong, J\. Wang, K\. Huang, T\. Gu, Y\. Wang, J\. Wang, D\. Liang, Z\. Li, Y\. Teng, Y\. Xiao, and Y\. Wang \(2024\)ESC\-Eval: evaluating emotion support conversations in large language models\.InProceedings of EMNLP,pp\. 15785–15810\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.emnlp-main.883),[Link](https://aclanthology.org/2024.emnlp-main.883/)Cited by:[§4\.1](https://arxiv.org/html/2608.10626#S4.SS1.p2.1)\.
- W\. Zhao, X\. Sui, X\. Han, Y\. Deng, Y\. Hu, J\. Guo, L\. Qin, Q\. Du, S\. Wang, Y\. Zhao, B\. Qin, and T\. Liu \(2025\)Chain of strategy optimization makes large language models better emotional supporter\.InFindings of EMNLP,pp\. 15361–15381\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.findings-emnlp.831)Cited by:[§2](https://arxiv.org/html/2608.10626#S2.SS0.SSS0.Px1.p1.1)\.
- C\. Zheng, S\. Sabour, J\. Wen, Z\. Zhang, and M\. Huang \(2023\)AugESC: dialogue augmentation with large language models for emotional support conversation\.InFindings of ACL,pp\. 1552–1568\.External Links:[Document](https://dx.doi.org/10.18653/v1/2023.findings-acl.99)Cited by:[§1](https://arxiv.org/html/2608.10626#S1.p2.1),[§2](https://arxiv.org/html/2608.10626#S2.SS0.SSS0.Px1.p1.1)\.
- J\. Zhou, Z\. Chen, B\. Wang, and M\. Huang \(2023\)Facilitating multi\-turn emotional support conversation with positive emotion elicitation: a reinforcement learning approach\.InProceedings of ACL,pp\. 1714–1729\.External Links:[Document](https://dx.doi.org/10.18653/v1/2023.acl-long.96),[Link](https://aclanthology.org/2023.acl-long.96/)Cited by:[§1](https://arxiv.org/html/2608.10626#S1.p1.1),[§2](https://arxiv.org/html/2608.10626#S2.SS0.SSS0.Px1.p1.1)\.
- R\. Zhu, X\. Huang, Y\. Wu, R\. Wang, Z\. Sun, T\. Ren, W\. Luo, B\. Qiu, J\. Ye, Y\. Li, and W\. Hu \(2026\)EIBench: a simulator\-based benchmark and turn\-credit RL for emotion management\.External Links:2606\.15532,[Link](https://arxiv.org/abs/2606.15532)Cited by:[§4\.1](https://arxiv.org/html/2608.10626#S4.SS1.p2.1)\.Similar Articles
Verifiable Self-Evolution for Open-Ended Dialogue Skills via Future-Feedback Prediction
This paper introduces a method for self-evolution of open-ended dialogue skills using future-feedback prediction, converting conversational feedback into a fixed offline objective to enable reproducible skill optimization without live traffic. The approach achieves over 75% prediction accuracy on a privacy-preserving sales-assistant dataset.
SkillEvo: Self-Renewing Evolution Gradients from Multi-Turn Interaction Feedback
SkillEvo introduces a method to continuously improve AI agent skills by using multi-turn interaction feedback and governance layers to maintain evolution gradients, surpassing self-reflection and single-turn QA-driven approaches.
Breaking the Impasse: Dual-Scale Evolutionary Policy Training for Social Language Agents
This paper proposes Dual-Scale Evolutionary Policy Training (DEPT) to address the evolution impasse in social language agents, using asymmetric advantage reshaping to restore gradient signals during self-play.
CoEvolve: Training LLM Agents via Agent-Data Mutual Evolution
CoEvolve proposes an agent-data mutual evolution framework for training LLM agents through closed-loop, interaction-driven learning that adapts both the agent and its training data distribution. The method extracts feedback signals from rollout trajectories to guide LLM-based task synthesis, demonstrating significant improvements (15-19% absolute gains) across multiple Qwen models on AppWorld and BFCL benchmarks.
Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration
This paper proposes a method to train LLM agents with intrinsic meta-evolution capabilities, enabling spontaneous self-improvement without external rewards at inference time. Applied to Qwen3-30B and Seed-OSS-36B, the approach yields a 20% performance boost on web navigation benchmarks, with a 14B model outperforming Gemini-2.5-Flash.