Tag
This paper reports a pre-registered experiment on small economies of frontier LLM agents (Claude Opus 4.8), testing predictions about information-theoretic capacity regions for wealth growth and mean-field residual-scaling laws for population misalignment. Results confirm a quantitative information law connecting agent knowledge to earnings but reject the mean-field assumption, revealing discrete attractor dynamics instead.
CoffeeBench is a benchmark for evaluating LLM agents in a long-horizon multi-agent economic simulation where firms interact over 90 days to maximize profits, revealing differences in communication patterns and performance among various models.
A technical blog post describing a hackathon project where five different small AI models run a simulated economy, revealing that emergent market behavior differs when using heterogeneous agents compared to a single model, and that the price is a residue of agent decisions rather than a controllable dial.