Tag
This paper studies how preference optimization shapes LLM counselors' behavior in motivational interviewing, finding that penalizing confrontation trades goal persistence for relational attunement rather than teaching the balanced skill of rolling with resistance.
MIThinker proposes a lightweight reasoning model for motivational interviewing counseling agents, trained via supervised fine-tuning and reinforcement learning to generate therapeutic thoughts, achieving MI competency comparable to state-of-the-art systems with lower computation.