generative-agents

Tag

Cards List
#generative-agents

Benchmarking the Personalization Capabilities of Large Language Models

arXiv cs.AI · 2026-07-24 Cached

This paper introduces SDR-Bench, a benchmark for evaluating the personalization capabilities of large language models in a two-party Bayesian Persuasion framework, finding a consistent plateau across frontier LLMs and validating the framework with a field deployment.

0 favorites 0 likes
#generative-agents

Evaluating Generative Agents with Actions Grounded in Socially Distributed Task Environments using Incognita

arXiv cs.AI · 2026-07-07 Cached

This paper introduces Incognita, a framework for evaluating generative agents in socially distributed task environments where knowledge is partitioned across roles. Experiments show improvements in agent success rates but overall reliability remains low.

0 favorites 0 likes
#generative-agents

Edu-Theater: A Data-Efficient Agent Framework for Scalable Learner Behavior Simulation through Staging Roll-Call

arXiv cs.LG · 2026-06-16 Cached

Edu-Theater is a data-efficient agent framework that uses LLM-powered generative agents to simulate learner behavior in educational settings. It employs a cohort-aware roll-call paradigm to infer learner states with fewer data and computational resources, achieving higher simulation accuracy.

0 favorites 0 likes
#generative-agents

When Plausible Is Not Realistic: Evaluating Human Mobility in LLM-Based Urban Simulation

arXiv cs.CL · 2026-06-15 Cached

This paper introduces a validation framework to evaluate whether LLM-based urban simulators reproduce empirical human mobility patterns. Using data from Paris and Shanghai, the authors find a substantial gap between plausible narratives and realistic mobility constraints, and provide open infrastructure for reproducible evaluation.

0 favorites 0 likes
← Back to home

Submit Feedback