Tag
This paper introduces Preference Tree Optimization (PTO), a framework that generates preference data via look-ahead simulations to iteratively improve goal-oriented dialogue agents, with experiments showing gains in Motivational Interviewing settings.
This paper proposes UP-NRPA, an online framework that integrates user portraits with nested rollout policy adaptation using large language models to dynamically customize dialogue strategies without offline training, achieving 100% success on multiple dialogue tasks.