Tag
This paper introduces a framework for extracting behaviorally grounded user profiles from social media posts to enhance personalized alignment and multi-perspective reasoning in large language models, outperforming synthetic profile methods.
This paper introduces MirageBench, a benchmark showing that LLMs with persistent memory fabricate user profiles through over-inference 35-49% of the time, and reveals that model self-reported confidence is inversely correlated with actual over-inference, making self-monitoring unreliable for comparing models.
A discussion on how AI agents should handle user context: upfront disclosure or gradual learning, with various existing approaches like project memory and chat summaries found lacking.
Ψ-Bench is a benchmark for evaluating LLMs' ability to influence users through persuasive dialogues, incorporating user profiles for personalized persuasion. Experiments show that even state-of-the-art models have room for improvement, and access to client profiles significantly boosts performance.