Tag
This research investigates how LLM agents influence consensus formation in mixed human-AI groups, identifying three regimes where agent proportions alter convergence strength and the semantic nature of consensus.
GroupMemBench is a new benchmark for evaluating LLM agent memory in multi-party conversations, exposing failures in current memory systems with the best achieving only 46% average accuracy.