PersonaArena: Dynamic Simulation for Evaluating and Enhancing Persona-Level Role-Playing in Large Language Models
Summary
PersonaArena is a dynamic simulation framework that uses a large corpus of social content and a multi-agent debating judge to evaluate and improve LLMs' ability to maintain coherent and authentic persona-level role-playing in realistic social scenarios.
View Cached Full Text
Cached at: 05/19/26, 06:38 AM
# PersonaArena: Dynamic Simulation for Evaluating and Enhancing Persona-Level Role-Playing in Large Language Models Source: [https://arxiv.org/abs/2605.17044](https://arxiv.org/abs/2605.17044) [View PDF](https://arxiv.org/pdf/2605.17044) > Abstract:Large language models \(LLMs\) increasingly serve as interactive social agents, yet their ability to maintain coherent and authentic persona\-level role\-playing remains limited, particularly in realistic social scenarios\. Existing research predominantly focuses on character\-level settings and relies on static evaluation formats, failing to capture the complexity of everyday social interactions\. In this work, we present PersonaArena, a dynamic simulation framework for evaluating and improving persona\-level role\-playing in LLMs\. PersonaArena leverages a large, filtered corpus of user\-generated social content to construct a nuanced persona bank, and elicits multi\-turn, context\-rich interactions within simulated social environments\. Our framework features a multi\-agent debating judge for holistic and unbiased assessment\. Through extensive experiments, we demonstrate that PersonaArena enables rigorous evaluation and enhancement of LLMs' role\-playing capabilities, advancing the development of more authentic and socially adept AI agents\. ## Submission history From: Wenlong Shi \[[view email](https://arxiv.org/show-email/193ca9d4/2605.17044)\] **\[v1\]**Sat, 16 May 2026 15:23:28 UTC \(5,612 KB\)
Similar Articles
Beyond Cooperative Simulators: Generating Realistic User Personas for Robust Evaluation of LLM Agents
Introduces Persona Policies (PPol), a plug-and-play control layer that uses LLM-driven evolutionary program search to generate diverse, human-like user personas for evaluating LLM agents. Achieves 33–62% fitness gains over baseline, with human-likeness rated at 80.4%, and improves agent robustness with +17% task success.
Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation
Introduces PALATE, a scalable benchmark for evaluating role-playing agents using person-aligned LLM-simulated users and personalized rubrics, addressing limitations of fixed-history evaluation.
Point of Order: Action-Aware LLM Persona Modeling for Data-Grounded Civic Deliberation
A reproducible pipeline converts public Zoom recordings into speaker-attributed transcripts enriched with personas, topics, and action tags, then fine-tunes LLM personas on this data. Action-aware fine-tuning significantly improves persona fidelity, consistency, and deliberative responsiveness, enabling realistic civic deliberation simulations.
How Well Do Large Language Models Capture Human Personality?
This paper systematically evaluates assumptions about LLM persona prompting and identifies 'persona manifold collapse,' where richer persona descriptions reduce behavioral diversity and simulation fidelity. The findings show that simple age-gender personas often outperform more detailed profiles.
SocialPersona: Benchmarking Personalized Profiling and Response with Multimodal Social-Media Context
Introduces SocialPersona, a benchmark for evaluating multimodal large language models on their ability to recover revealed preferences from longitudinal social-media timelines and use them in personalized dialogue.