Tag
Introduces TAF-MED, a physician-reviewed benchmark of 500 multi-turn medical safety scenarios, showing that LLMs often collapse from safe initial refusals to unsafe responses when users declare self-treatment intent. Evaluation of eight LLMs across 4,000 conversations finds 71.6% contained unsafe responses and first-turn safety is an insufficient proxy for conversational safety persistence.
This paper introduces LLM-OSDA, a dynamic cost-per-click auction for native advertising in multi-turn LLM conversations, integrating Bellman optimal stopping, winner allocation, and envelope pricing. Experiments show an 11% net revenue improvement over fixed-timing baselines while maintaining user retention.
This paper proposes PUMA, a framework for LLM personalization in multi-turn conversations that models latent user states and uses the Free Energy Principle to select dialogue actions, improving long-horizon outcomes on healthcare counseling benchmarks.
A new paper by Microsoft Research and Salesforce reveals that LLM performance drops significantly in multi-turn conversations due to a 'Lost in Conversation' phenomenon, challenging the reliability of current single-turn benchmarks.