Tag
This paper introduces SDR-Bench, a benchmark for evaluating the personalization capabilities of large language models in a two-party Bayesian Persuasion framework, finding a consistent plateau across frontier LLMs and validating the framework with a field deployment.
This paper introduces Incognita, a framework for evaluating generative agents in socially distributed task environments where knowledge is partitioned across roles. Experiments show improvements in agent success rates but overall reliability remains low.
Edu-Theater is a data-efficient agent framework that uses LLM-powered generative agents to simulate learner behavior in educational settings. It employs a cohort-aware roll-call paradigm to infer learner states with fewer data and computational resources, achieving higher simulation accuracy.
This paper introduces a validation framework to evaluate whether LLM-based urban simulators reproduce empirical human mobility patterns. Using data from Paris and Shanghai, the authors find a substantial gap between plausible narratives and realistic mobility constraints, and provide open infrastructure for reproducible evaluation.