Economy of Minds: Emerging Multi-Agent Intelligence with Economic Interactions

arXiv cs.CL Papers

Summary

This paper introduces an economic framework for multi-agent AI systems, where agents interact through economic mechanisms to produce emergent collective intelligence, drawing from Harvard and MIT researchers.

arXiv:2606.02859v1 Announce Type: new Abstract: How can a population of agents self-orchestrate and self-adapt into stronger collective intelligence without centralized control? Inspired by Friedrich Hayek's economic theory of decentralized coordination in markets, we study this question through an agent economy in which agents compete via auctions for the right to act, exchange payments, and accumulate wealth from environmental rewards. These simple economic signals induce decentralized credit assignment, driving planning without global orchestration or explicit communication protocols. The population evolves through economic selection: effective agents accumulate wealth and are mutated via exploitation, while ineffective ones go bankrupt and are replaced via exploration. We show that, initialized with weak agents, the economy produces emergent multi-step reasoning strategies and outperforms stronger monolithic baselines across five agentic tasks, including mathematical reasoning, financial research, scientific research, accelerator design, and distributed-system optimization. We further provide theoretical insights into how economic dynamics shape agent behaviors, linking local incentives to long-term global performance. Our results suggest a new path to multi-agent intelligence: rather than engineering coordination, we can design decentralized incentive structures under which it automatically emerges.
Original Article
View Cached Full Text

Cached at: 06/03/26, 09:35 AM

# Economy of Minds: Emerging Multi-Agent Intelligence with Economic Interactions
Source: [https://arxiv.org/html/2606.02859](https://arxiv.org/html/2606.02859)
Zhenting Qi![[Uncaptioned image]](https://arxiv.org/html/2606.02859v1/logos/Harvard-Crimson-Logo-2002.jpg),\{\}^\{\\raisebox\{\-0\.2pt\}\{\\includegraphics\[height=6\.3778pt\]\{logos/Harvard\-Crimson\-Logo\-2002\.jpg\}\},\}Huangyuan Su![[Uncaptioned image]](https://arxiv.org/html/2606.02859v1/logos/Harvard-Crimson-Logo-2002.jpg),![[Uncaptioned image]](https://arxiv.org/html/2606.02859v1/logos/Kempner-logo_Full-Color-Harvard.jpg)\{\}^\{\\raisebox\{\-0\.2pt\}\{\\includegraphics\[height=6\.3778pt\]\{logos/Harvard\-Crimson\-Logo\-2002\.jpg\}\},\\raisebox\{\-0\.2pt\}\{\\includegraphics\[height=6\.3778pt\]\{logos/Kempner\-logo\_Full\-Color\-Harvard\.jpg\}\}\}Ao Qu![[Uncaptioned image]](https://arxiv.org/html/2606.02859v1/logos/MIT_logo_2003-2023.svg.png)\{\}^\{\\raisebox\{\-0\.2pt\}\{~\\includegraphics\[height=4\.78339pt\]\{logos/MIT\_logo\_2003\-2023\.svg\.png\}\}\}Chenyu Wang![[Uncaptioned image]](https://arxiv.org/html/2606.02859v1/logos/Harvard-Crimson-Logo-2002.jpg)\{\}^\{\\raisebox\{\-0\.2pt\}\{\\includegraphics\[height=6\.3778pt\]\{logos/Harvard\-Crimson\-Logo\-2002\.jpg\}\}\} Yu Yao![[Uncaptioned image]](https://arxiv.org/html/2606.02859v1/logos/MIT_logo_2003-2023.svg.png)\{\}^\{\\raisebox\{\-0\.2pt\}\{~\\includegraphics\[height=4\.78339pt\]\{logos/MIT\_logo\_2003\-2023\.svg\.png\}\}\}Han Zheng![[Uncaptioned image]](https://arxiv.org/html/2606.02859v1/logos/MIT_logo_2003-2023.svg.png)\{\}^\{\\raisebox\{\-0\.2pt\}\{~\\includegraphics\[height=4\.78339pt\]\{logos/MIT\_logo\_2003\-2023\.svg\.png\}\}\}Kushal Chattopadhyay![[Uncaptioned image]](https://arxiv.org/html/2606.02859v1/logos/Harvard-Crimson-Logo-2002.jpg)\{\}^\{\\raisebox\{\-0\.2pt\}\{\\includegraphics\[height=6\.3778pt\]\{logos/Harvard\-Crimson\-Logo\-2002\.jpg\}\}\}Guowei Xu![[Uncaptioned image]](https://arxiv.org/html/2606.02859v1/logos/Harvard-Crimson-Logo-2002.jpg)\{\}^\{\\raisebox\{\-0\.2pt\}\{\\includegraphics\[height=6\.3778pt\]\{logos/Harvard\-Crimson\-Logo\-2002\.jpg\}\}\} Zihan Wang![[Uncaptioned image]](https://arxiv.org/html/2606.02859v1/logos/2077.jpeg)\{\}^\{\\raisebox\{\-0\.2pt\}\{~\\includegraphics\[height=3\.98611pt\]\{logos/2077\.jpeg\}\}\}Weirui Ye![[Uncaptioned image]](https://arxiv.org/html/2606.02859v1/logos/MIT_logo_2003-2023.svg.png)\{\}^\{\\raisebox\{\-0\.2pt\}\{~\\includegraphics\[height=4\.78339pt\]\{logos/MIT\_logo\_2003\-2023\.svg\.png\}\}\}Vijay Janapa Reddi![[Uncaptioned image]](https://arxiv.org/html/2606.02859v1/logos/Harvard-Crimson-Logo-2002.jpg)\{\}^\{\\raisebox\{\-0\.2pt\}\{\\includegraphics\[height=6\.3778pt\]\{logos/Harvard\-Crimson\-Logo\-2002\.jpg\}\}\}Ju Li![[Uncaptioned image]](https://arxiv.org/html/2606.02859v1/logos/MIT_logo_2003-2023.svg.png)\{\}^\{\\raisebox\{\-0\.2pt\}\{~\\includegraphics\[height=4\.78339pt\]\{logos/MIT\_logo\_2003\-2023\.svg\.png\}\}\}Paul Pu Liang![[Uncaptioned image]](https://arxiv.org/html/2606.02859v1/logos/MIT_logo_2003-2023.svg.png)\{\}^\{\\raisebox\{\-0\.2pt\}\{~\\includegraphics\[height=4\.78339pt\]\{logos/MIT\_logo\_2003\-2023\.svg\.png\}\}\} Himabindu Lakkaraju![[Uncaptioned image]](https://arxiv.org/html/2606.02859v1/logos/Harvard-Crimson-Logo-2002.jpg)\{\}^\{\\raisebox\{\-0\.2pt\}\{\\includegraphics\[height=6\.3778pt\]\{logos/Harvard\-Crimson\-Logo\-2002\.jpg\}\}\}Sham Kakade![[Uncaptioned image]](https://arxiv.org/html/2606.02859v1/logos/Harvard-Crimson-Logo-2002.jpg),![[Uncaptioned image]](https://arxiv.org/html/2606.02859v1/logos/Kempner-logo_Full-Color-Harvard.jpg)\{\}^\{\\raisebox\{\-0\.2pt\}\{\\includegraphics\[height=6\.3778pt\]\{logos/Harvard\-Crimson\-Logo\-2002\.jpg\}\},\\raisebox\{\-0\.2pt\}\{\\includegraphics\[height=6\.3778pt\]\{logos/Kempner\-logo\_Full\-Color\-Harvard\.jpg\}\}\}Yilun Du![[Uncaptioned image]](https://arxiv.org/html/2606.02859v1/logos/Harvard-Crimson-Logo-2002.jpg),![[Uncaptioned image]](https://arxiv.org/html/2606.02859v1/logos/Kempner-logo_Full-Color-Harvard.jpg),\{\}^\{\\raisebox\{\-0\.2pt\}\{\\includegraphics\[height=6\.3778pt\]\{logos/Harvard\-Crimson\-Logo\-2002\.jpg\}\},\\raisebox\{\-0\.2pt\}\{\\includegraphics\[height=6\.3778pt\]\{logos/Kempner\-logo\_Full\-Color\-Harvard\.jpg\}\},\}11footnotemark:1 ![[Uncaptioned image]](https://arxiv.org/html/2606.02859v1/logos/Harvard-Crimson-Logo-2002.jpg)\{\}^\{\\raisebox\{\-0\.2pt\}\{\\includegraphics\[height=6\.3778pt\]\{logos/Harvard\-Crimson\-Logo\-2002\.jpg\}\}\}Harvard![[Uncaptioned image]](https://arxiv.org/html/2606.02859v1/logos/MIT_logo_2003-2023.svg.png)\{\}^\{\\raisebox\{\-0\.2pt\}\{~\\includegraphics\[height=4\.78339pt\]\{logos/MIT\_logo\_2003\-2023\.svg\.png\}\}\}MIT![[Uncaptioned image]](https://arxiv.org/html/2606.02859v1/logos/2077.jpeg)\{\}^\{\\raisebox\{\-0\.2pt\}\{~\\includegraphics\[height=3\.98611pt\]\{logos/2077\.jpeg\}\}\}2077AI![[Uncaptioned image]](https://arxiv.org/html/2606.02859v1/logos/Kempner-logo_Full-Color-Harvard.jpg)\{\}^\{\\raisebox\{\-0\.2pt\}\{\\includegraphics\[height=6\.3778pt\]\{logos/Kempner\-logo\_Full\-Color\-Harvard\.jpg\}\}\}Kempner Institute [![[Uncaptioned image]](https://arxiv.org/html/2606.02859v1/logos/github.jpg)GitHub](https://github.com/zhentingqi/EoM)[![[Uncaptioned image]](https://arxiv.org/html/2606.02859v1/logos/web.png)Project Page](https://zhentingqi.github.io/internal/projects/EoM)

###### Abstract

How can a population of agents self\-orchestrate and self\-adapt into stronger collective intelligence without centralized control? Inspired by Friedrich Hayek’s economic theory of decentralized coordination in markets, we study this question through anagent economyin which agents compete via auctions for the right to act, exchange payments, and accumulate wealth from environmental rewards\. These simple economic signals induce decentralized credit assignment, driving planning without global orchestration or explicit communication protocols\. The population evolves through economic selection: effective agents accumulate wealth and are mutated via exploitation, while ineffective ones go bankrupt and are replaced via exploration\. We show that, initialized with weak agents, the economy produces emergent multi\-step reasoning strategies and outperforms stronger monolithic baselines across five agentic tasks, including mathematical reasoning, financial research, scientific research, accelerator design, and distributed\-system optimization\. We further provide theoretical insights into how economic dynamics shape agent behaviors, linking local incentives to long\-term global performance\. Our results suggest a new path to multi\-agent intelligence: rather than engineering coordination, we can design decentralized incentive structures under which it automatically emerges\.

![Refer to caption](https://arxiv.org/html/2606.02859v1/x1.png)Figure 1:Evolution of an agent society over a stream of tasks\.Each panel shows the population at a given stage, where agents are continuously created, selected, connected, and eliminated\. As the society encounters more tasks, ineffective agents are removed and corrected, while useful ones persist and diversify, leading to an alive and increasingly structured population\.## 1Introduction

Imagine a world populated by a number of intelligent agents\. Each agent may perform well on certain tasks, but it remains fundamentally limited: each operates with its own priors, partial observations of environments, and bounded computational resources\. When faced with complex tasks that exceed individual capabilities, no single agent can reliably solve problems from start to finish\. How, then, can such a population collectively solve these tasks?

A natural approach is to introduce a central orchestrator that creates agents, assigns specializations, and coordinates actions across agents\[[19](https://arxiv.org/html/2606.02859#bib.bib19),[32](https://arxiv.org/html/2606.02859#bib.bib32),[46](https://arxiv.org/html/2606.02859#bib.bib46)\]\. However, such centralized systems suffer from two fundamental limitations\. First, planning is bottlenecked at a single coordination gate: all information and decision\-making must flow through the orchestrator, creating both a performance bottleneck and a single point of failure\[[6](https://arxiv.org/html/2606.02859#bib.bib6),[49](https://arxiv.org/html/2606.02859#bib.bib49)\]\. Second, learning and adaptation become increasingly inefficient as the system scales\. As the number of agents grows, the orchestrator must reason about and manage an ever\-expanding set of agents, leading to coordination complexity that grows linearly with system size\[[49](https://arxiv.org/html/2606.02859#bib.bib49),[22](https://arxiv.org/html/2606.02859#bib.bib22)\]\. These limitations motivate a shift in perspective: rather than asking how to build better centralized controllers, we ask whether a group of agents can form decentralized intelligence, which organizes itself into strong collective capability and naturally evolves through time\.

These questions are especially compelling today\. Advances in large foundation models such as large language models \(LLMs\) and world action models make it possible to instantiate intelligent agents that can reason, use tools, and operate in open\-ended digital or embodied environments\[[50](https://arxiv.org/html/2606.02859#bib.bib50),[30](https://arxiv.org/html/2606.02859#bib.bib30),[21](https://arxiv.org/html/2606.02859#bib.bib21),[51](https://arxiv.org/html/2606.02859#bib.bib51)\]\. However, such agents are typically carefully engineered and tailored to specific domains, introducing inductive biases that might hinder general utility\. Also, constraints in the underlying models, such as finite context length, partial knowledge, and restricted inference budget, combined with environmental limitations like partial observability and constrained tool access, make it difficult for single agents to solve complex, compositional tasks alone\[[28](https://arxiv.org/html/2606.02859#bib.bib28),[20](https://arxiv.org/html/2606.02859#bib.bib20),[39](https://arxiv.org/html/2606.02859#bib.bib39)\]\. Therefore, beyond improving these individual agents, another axis to scale up agentic intelligence is to organize these inherently constrained agents into an effective learnable system as a whole\[[7](https://arxiv.org/html/2606.02859#bib.bib7),[48](https://arxiv.org/html/2606.02859#bib.bib48),[33](https://arxiv.org/html/2606.02859#bib.bib33)\]\.

The economy in human society provides an intuition for such decentralized intelligence\. In his seminal essayThe Use of Knowledge in Society\[[14](https://arxiv.org/html/2606.02859#bib.bib14)\], Friedrich Hayek argued that the core problem of an economy is not optimization under known information, but the utilization of knowledge that is dispersed across individuals, which cannot be aggregated by any central authority\. Hayek’s key insight is that the free market solves this problem through the price system: prices serve as signals that aggregate and communicate dispersed information, enabling individuals to coordinate their actions without global awareness\. As a result, large\-scale social orders emerge from decentralized interactions through competition, specialization, and exchange\. Building on this view, Baum further argued that economic organization provides a concrete model of intelligence, even though individual routines are simple\[[1](https://arxiv.org/html/2606.02859#bib.bib1),[2](https://arxiv.org/html/2606.02859#bib.bib2)\]\. His Hayek machine shows that economic pressure serves as a natural mechanism for credit assignment, determining which routines act, how they are connected, and how they are renewed without centralized evaluation\.

In this work, we show that simple economic incentives are sufficient for modern intelligent agents to self\-orchestrate and self\-evolve as a society\. We presentEconomy of Minds \(EoM\), a system where a population of agents compete via auctions for the right to act, exchange payments through peer\-to\-peer transactions, and accumulate or lose wealth based on an outcome reward\. Agents that consistently contribute to successful trajectories naturally accumulate wealth, thus they survive and are periodically mutated through exploitation, while ineffective agents are eliminated through economic selection and replaced through exploration\. Importantly, each agent operates purely locally: it only requires \(1\) a wake\-up condition, and \(2\) executes its action accordingly\. There is no need for global awareness, explicit coordination, engineered topology, or prescribed communication protocols\.

We implement this framework using LLM\-based agents and evaluate it on five digital agentic tasks, whereEoMimproves mathematical reasoning from 15\.9% to 57\.0%, raises financial research performance from 45\.0% to 60\.0%, increases scientific research accuracy over the baseline from 5\.0% to 20\.0%, reduces accelerator\-design average energy\-delay product to 39\.3 versus 80\.2 for a strong domain\-specific method, and attains a best distributed\-system optimization cost of 657 versus 930 for the baseline\. Across these domains, we show with detailed cases that the agent society gradually self\-organizes into effective workflows: competent agents persist, mutate, and specialize, while weaker agents are pruned and refined, yielding emergent structure and continual adaptation\.

Our findings suggest that economic organization provides a simple, general, and scalable foundation for decentralized multi\-agent intelligence\. Rather than explicitly designing coordination mechanisms, we can define an incentive structure under which coordination, specialization, and cooperation naturally emerge between agents\. This points toward a broader paradigm for multi\-agent systems in the era of large foundation models: not as centrally engineered pipelines, but as evolving agent societies whose collective intelligence is shaped by the economies they inhabit\.

## 2An Economy of Language Agents

We model a society of language agents interacting through an economic mechanism\. Each agent operates locally, making decisions based only on its own triggering condition and policy, while global coordination emerges from economic interactions\. The system consists of two coupled processes: \(1\)*planning*, which governs how agents act and assigns credits to them within an episode, and \(2\)*adaptation*, which governs how the population evolves across episodes\.

### 2\.1Problem Setup

We consider a task environment modeled as a partially observed Markov decision processℰ=\(𝒮,𝒜,P,r,γ,μ0\),\\mathcal\{E\}=\(\\mathcal\{S\},\\mathcal\{A\},P,r,\\gamma,\\mu\_\{0\}\),where𝒮\\mathcal\{S\}is the state space,𝒜\\mathcal\{A\}is the action space,P​\(s′∣s,a\)P\(s^\{\\prime\}\\mid s,a\)is the transition kernel,r:𝒮×𝒜→ℝr:\\mathcal\{S\}\\times\\mathcal\{A\}\\to\\mathbb\{R\}is the reward function,γ∈\(0,1\]\\gamma\\in\(0,1\]is the discount factor, andμ0\\mu\_\{0\}is the initial\-state distribution\. At steptt, the system observesot∈𝒪o\_\{t\}\\in\\mathcal\{O\}; in fully observed settings,ot=sto\_\{t\}=s\_\{t\}\.

Each agent is implemented by a language model parametrized byθ\\theta\. Here we adopt a simplified setting, where we use a shared frozen model backbone for all agents, with diversity arising entirely through system prompts\. Formally, an agent is a tuplea=\(ϕa,πa,ba,Wa\),a=\(\\phi\_\{a\},\\pi\_\{a\},b\_\{a\},W\_\{a\}\),whereϕa:𝒪→\{0,1\}\\phi\_\{a\}:\\mathcal\{O\}\\to\\\{0,1\\\}is a*triggering predicate*that determines whether the agent is eligible to act,πa:𝒪→Δ​\(𝒜\)\\pi\_\{a\}:\\mathcal\{O\}\\to\\Delta\(\\mathcal\{A\}\)is its action policy,ba∈ℝ≥0b\_\{a\}\\in\\mathbb\{R\}\_\{\\geq 0\}is a fixed bid associated with the agent, andWa∈ℝW\_\{a\}\\in\\mathbb\{R\}is its current wealth\. Bothϕa\\phi\_\{a\}andπa\\pi\_\{a\}are instantiated by the same frozen LLM with agent\-specific promptspa=\(patrig,paact\)p\_\{a\}=\(p\_\{a\}^\{\\mathrm\{trig\}\},p\_\{a\}^\{\\mathrm\{act\}\}\):

ϕa\(o\)=LLMθ\(patrig;o\)∈\{0,1\},πa\(⋅∣o\)=LLMθ\(paact;o\)∈Δ\(𝒜\)\.\\displaystyle\\phi\_\{a\}\(o\)=\\mathrm\{LLM\}\_\{\\theta\}\(p\_\{a\}^\{\\mathrm\{trig\}\};o\)\\in\\\{0,1\\\},\\quad\\pi\_\{a\}\(\\cdot\\mid o\)=\\mathrm\{LLM\}\_\{\\theta\}\(p\_\{a\}^\{\\mathrm\{act\}\};o\)\\in\\Delta\(\\mathcal\{A\}\)\.Thus, an agent is fully specified by its prompt pair together with its fixed bid and current wealth\.111This formulation generalizes Baum’s Hayek machine\[[2](https://arxiv.org/html/2606.02859#bib.bib2),[3](https://arxiv.org/html/2606.02859#bib.bib3)\]from hand\-specified condition\-action rules to general agents: the condition is a model\-based predicate, and the action is drawn from a model\-based policy rather than being a fixed move\.At episodeee, the active population is denoted by𝒫e\\mathcal\{P\}\_\{e\}\.

### 2\.2Planning with Auctions and Transactions

![Refer to caption](https://arxiv.org/html/2606.02859v1/x2.png)Figure 2:Auctions\.Agents whose wake\-up conditions are satisfied become eligible to bid; the highest bidder wins the auction, executes the action, and advances the environment fromsts\_\{t\}tost\+1s\_\{t\+1\}\.![Refer to caption](https://arxiv.org/html/2606.02859v1/x3.png)Figure 3:Transactions\.Credit assignment naturally emerges as profits flow backward through the action sequence, rewarding agents whose actions enable successful downstream outcomes\.At each environment step, agents compete for control through an auction \([Figure˜2](https://arxiv.org/html/2606.02859#S2.F2)\)\. Given the current observationoto\_\{t\}, each agent evaluates its triggering predicate to determine eligibility\. The eligible set isEt=\{a∈𝒫e:ϕa​\(ot\)=1\}\.E\_\{t\}=\\\{a\\in\\mathcal\{P\}\_\{e\}:\\phi\_\{a\}\(o\_\{t\}\)=1\\\}\.IfEt=∅E\_\{t\}=\\emptyset, no agent acts and the episode terminates or defaults to a null transition, depending on the environment\. Otherwise, the auction winner is the highest\-bidding eligible agent,at⋆∈arg⁡maxa∈Et⁡ba,a\_\{t\}^\{\\star\}\\in\\arg\\max\_\{a\\in E\_\{t\}\}b\_\{a\},with ties broken randomly\.

The auction serves as a decentralized action\-selection mechanism\. Rather than relying on a centralized policy, control is allocated to the agent that bids the highest value to acting in the current context, as reflected by its bid\. Once selected, the winning agentat⋆a\_\{t\}^\{\\star\}samples an action from its policyatenv∼πat⋆\(⋅∣ot\),a\_\{t\}^\{\\mathrm\{env\}\}\\sim\\pi\_\{a\_\{t\}^\{\\star\}\}\(\\cdot\\mid o\_\{t\}\),which advances the environment and produces the next observationot\+1o\_\{t\+1\}and rewardrtr\_\{t\}\. Letat−1⋆a\_\{t\-1\}^\{\\star\}denote the previous winning agent in the same episode\. We then apply a bucket\-brigade transfer rule \([Figure˜3](https://arxiv.org/html/2606.02859#S2.F3)\):

Wat⋆\\displaystyle W\_\{a\_\{t\}^\{\\star\}\}←Wat⋆−bat⋆\+rt,Wat−1⋆←Wat−1⋆\+bat⋆\.\\displaystyle\\leftarrow W\_\{a\_\{t\}^\{\\star\}\}\-b\_\{a\_\{t\}^\{\\star\}\}\+r\_\{t\},\\quad W\_\{a\_\{t\-1\}^\{\\star\}\}\\leftarrow W\_\{a\_\{t\-1\}^\{\\star\}\}\+b\_\{a\_\{t\}^\{\\star\}\}\.\(1\)That is, the winner pays its bid to the previously active agent while collecting any environmental rewardrtr\_\{t\}\. For the first winner in an episode, the payment is made to the house\.

This payment rule yields a decentralized form of credit assignment\. An agent profits not only by directly receiving reward, but also by taking actions that place the system in states from which downstream agents are willing to pay highly for control\. As a result, value flows backward along successful trajectories: agents whose actions enable productive continuations accumulate wealth, while agents that lead the system into unproductive states lose it\.

### 2\.3Adaptation with Exploration and Exploitation

Beyond within\-episode planning, the population evolves across episodes through economic selection\. Let𝒢\\mathcal\{G\}denote a prompt\-generation operator that is attached to the agents and proposes new agent prompts from the current one, for example, by mutating a successful agent’s or amending a failed one’s system prompt\. Each newly created agent is initialized with wealthW0≥0W\_\{0\}\\geq 0, while existing agents may also incur a periodic rentρ≥0\\rho\\geq 0\. Examples of adaptations are provided in Appendix[E](https://arxiv.org/html/2606.02859#A5)and[F](https://arxiv.org/html/2606.02859#A6), and theoretical motivations are stated in Appendix[C](https://arxiv.org/html/2606.02859#A3)\.

#### Exploitation\.

Agents that consistently contribute to successful trajectories automatically accumulate wealth and persist in the population\. Wealthy agents are periodically selected as parents and are mutated to produce new agents, allowing their successful patterns to be reused and refined\. This mechanism biases the population toward high\-performing behaviors, reinforces effective strategies, and promotes the emergence of specialized roles\. Concretely, wealthy agents use their prompt generators𝒢\\mathcal\{G\}to propose prompt mutations that preserve their useful wake\-up conditions or action policies while introducing small behavioral variations\.

#### Exploration\.

Agents lose wealth through unhelpful actions or prolonged inactivity\. When their wealth becomes negative, they are removed from the population, and new agents are introduced to replace them, typically by random or complementary variation of these bankrupt agents\. This continual turnover enables learning from failures, discoveries of new behaviors, and prevents premature convergence\. In this case, the prompt generator𝒢\\mathcal\{G\}is used by bankrupt agents to propose amended prompts that correct their failure modes or explore complementary regions of the behavior space\.

Formally, between episodes the population update consists of three stages:

1. 1\.Rent:each agent paysρ\\rho, soWa←Wa−ρW\_\{a\}\\leftarrow W\_\{a\}\-\\rho;
2. 2\.Removal:agents withWa<0W\_\{a\}<0are deleted;
3. 3\.Injection:new agents are added according to the exploitation and exploration mechanisms until the population satisfies the prescribed maximum size constraints\.

This induces a stochastic population process over finite sets of prompted agents, driven jointly by environment randomness, LLM stochasticity, and prompt generation\.

A key design choice is that bids are not learned online; instead, each agent receives a bid when it is introduced, which is then frozen\. For a newly injected agenta′a^\{\\prime\}\(novice\), letttdenote its first step at which it becomes eligible, and letCt=\{a∈𝒫e∖\{a′\}:ϕa​\(ot\)=1\}C\_\{t\}=\\\{a\\in\\mathcal\{P\}\_\{e\}\\setminus\\\{a^\{\\prime\}\\\}:\\phi\_\{a\}\(o\_\{t\}\)=1\\\}be the set of competing eligible agents at that step\. We assign its bid according to the novice rule

ba′=\(maxa∈Ct⁡ba\)\+εa′,εa′∼𝒟ε,b\_\{a^\{\\prime\}\}=\\Bigl\(\\max\_\{a\\in C\_\{t\}\}b\_\{a\}\\Bigr\)\+\\varepsilon\_\{a^\{\\prime\}\},\\qquad\\varepsilon\_\{a^\{\\prime\}\}\\sim\\mathcal\{D\}\_\{\\varepsilon\},\(2\)withmax⁡∅:=0\\max\\emptyset:=0\. In our experiments,𝒟ε\\mathcal\{D\}\_\{\\varepsilon\}is a small positive perturbation distribution\. This rule guarantees that the new agent wins the first auction for which it is eligible, forcing the system to test it at least once before market selection determines whether it should survive\.

Together, exploitation and exploration drive a self\-evolving system\. Exploitation preserves and sharpens useful behaviors, while exploration injects novelty and enables adaptation\. Crucially, evolution is governed entirely by economic signals, i\.e\., wealth accumulation and loss, without centralized supervision or explicit global performance labeling\.

### 2\.4Training and Evaluation

During optimization \([Algorithm˜1](https://arxiv.org/html/2606.02859#alg1)\), each task episode uses the auction\-based planning mechanism from[Section˜2\.2](https://arxiv.org/html/2606.02859#S2.SS2): agents bid for control, winning agents act in the environment, bid transfers propagate wealth along the trajectory, and environment rewards update the final actor\. These interactions provide an implicit evaluation signal, after which the adaptation step in[Section˜2\.3](https://arxiv.org/html/2606.02859#S2.SS3)removes bankrupt agents, retains profitable ones, mutates successful templates, and replenishes the population to preserve diversity\. For evaluation \([Algorithm˜2](https://arxiv.org/html/2606.02859#alg2)\), we reuse the same planning structure but disable the evolutionary and economic updates\. The trained population is frozen, bids are fixed, no payments, rewards, rent, births, or mutations are applied, and each test task runs on a thread\-local snapshot to avoid shared\-state effects\. Thus, optimization uses auctions for both action selection and credit redistribution, while evaluation uses them only as a decentralized decision rule, measuring the behavior of the learned population without further learning or wealth dynamics\.

## 3Experiments

### 3\.1Setup

Partial vs\. complete agents\.We use*partial agent*to denote any agent whose capability is intentionally incomplete to solve the full task\. This partiality can take different forms across domains: an agent may have a restricted action space, access to only one tool, a short generation budget, specialized roles, or only partial observation of the environment\. A*complete agent*, by contrast, has access to the full task interface, thus can use the full action space and attempt to solve the task end\-to\-end\. This distinction lets us test whether economic organization can compensate for, or even outperform, capability concentrated in a single complete agent\.

Tasks\.We instantiateEoMon five domains\. Onmathematical reasoning, we use MATH\[[16](https://arxiv.org/html/2606.02859#bib.bib16)\]to train on an easy\-to\-hard task stream from Level 1 to 5 and evaluate greedy pass@1 accuracy by difficulty level\. The population is initialized with planner, executor, and verifier agents, each with short output budgets of 128 tokens on average\. Onfinancial research, we use Finance\-Agent\-Bench\[[5](https://arxiv.org/html/2606.02859#bib.bib5)\], where agents solve financial questions over company filings using four tools; each partial agent has access to only one tool\. Forscientific research, we use FrontierScience\-Research\[[43](https://arxiv.org/html/2606.02859#bib.bib43)\], where agents are tasked to solve open\-ended scientific questions with literature, planner, executor, and verifier roles\. Foraccelerator design, we use theGemminibenchmark suite\[[13](https://arxiv.org/html/2606.02859#bib.bib13)\]and optimize mappings for 24 ResNet\-50 convolution kernels, with minimizing energy\-delay product \(EDP\) as the objective\. The population is initialized with three role\-specialized partial agents:*Historians*that summarize prior trial outcomes,*Planners*that propose mapping\-level search directions, and*Executors*that carry out local mapping evaluations\. Fordistributed\-system optimization, we use Cloudcast, a task from ADRS\[[8](https://arxiv.org/html/2606.02859#bib.bib8)\], in which agents iteratively improve a program to minimize the total data\-transfer cost\. Full task protocols, splits, agent definitions, tools, and model backbones are provided in Appendix[D](https://arxiv.org/html/2606.02859#A4)\.

Baselines\.Complete\-agent baselines includeReAct\[[50](https://arxiv.org/html/2606.02859#bib.bib50)\], which solves each task end\-to\-end with the same backbone and full task interface;GEA\[[45](https://arxiv.org/html/2606.02859#bib.bib45)\], which performs self\-improvement through experience sharing and agent evolution; andOpenEvolve\[[35](https://arxiv.org/html/2606.02859#bib.bib35),[27](https://arxiv.org/html/2606.02859#bib.bib27)\], an evolutionary monolithic coding\-agent baseline\. As a partial\-agent baseline, we use Multi\-Agent Debate\[[11](https://arxiv.org/html/2606.02859#bib.bib11)\], where several partial agents interact by communication but do not use market\-driven population evolution\. For domain\-specific comparisons, we also includeDOSA\[[18](https://arxiv.org/html/2606.02859#bib.bib18)\]for accelerator design\. Additional baseline details are in Appendix[D\.1](https://arxiv.org/html/2606.02859#A4.SS1)\.

### 3\.2Can Economics Turn Weak Individuals into Stronger Systems?

Across all five domains, we show that economic coordination turns individually partial agents into collective systems that match or outperform complete\-agent baselines\. On MATH,EoMimproves Llama\-3\.1\-8B agents from 15\.9% to 57\.0% and Gemma\-2\-9B agents from 4\.2% to 45\.1%, exceeding the corresponding complete\-agent baselines of 51\.9% and 44\.3% in[Table˜1](https://arxiv.org/html/2606.02859#S3.T1)\(left\)\. This is notable because each individual agent in the population is role\-specialized and restricted to short outputs, while the complete\-agent baseline can attempt the full problem end\-to\-end without restrictions\. On accelerator design,EoMreduces average EDP to 39\.3, compared with 43\.1 for the same\-backbone completeReActagent and 80\.2 for the domain\-specificDOSAbaseline in[Table˜1](https://arxiv.org/html/2606.02859#S3.T1)\(right\)\.

[Figure˜4](https://arxiv.org/html/2606.02859#S3.F4)shows that, on Finance\-Agent\-Bench,EoMrises from 45\.0% at initialization to 60\.0% after 30 training tasks\. This outperforms Multi\-Agent Debate at 50\.0%,ReActat 45\.0%, andGEAat 50\.0%, even though each partial agent inEoMcan access only one tool\. On FrontierScience\-Research,EoMreaches 8\.5% mean accuracy and 20\.0% best\-run accuracy, compared with 1\.8% mean and 5\.0% best\-run accuracy forGEAunder the same Gemini\-3\-Flash backbone\. Finally, on Cloudcast,EoMreaches an average total cost of 673 across three attempts, with the best attempt achieving 657, compared with 930 forOpenEvolve, corresponding to a 28% reduction in best cost while using fewer optimization episodes\.

Together, these results show that the relevant advantage is not merely “many agents” over “one agent\.” Rather,EoMshows that a population of partial agents, when organized by economic interactions, can match or surpass complete agents with greater individual access to the task interface\.

![Refer to caption](https://arxiv.org/html/2606.02859v1/x4.png)

\(a\) Finance\-Agent\-Bench

![Refer to caption](https://arxiv.org/html/2606.02859v1/x5.png)

\(b\) FrontierScience

![Refer to caption](https://arxiv.org/html/2606.02859v1/x6.png)

\(c\) Cloudcast

Figure 4:Performance across domains\.EoMconsistently outperforms baselines, demonstrating the benefits of economic coordination among agents\.Table 1:Performance across domains\.Left: MATH accuracy comparing constrained populations to complete agents \(\* denotes officially reported numbers\)\. Right: Accelerator design results measured by average EDP \(lower is better\)\. In both settings,EoMachieves stronger performance than corresponding baselines\.BackendPartial agentsComplete agentInitialAfter trainingLlama\-3\.1\-8B15\.9 \(1\.37\)57\.0\(3\.36\)51\.9\*Gemma\-2\-9B4\.2 \(0\.52\)45\.1\(4\.12\)44\.3\*
MethodAvg\. EDP \(μ\\muJ⋅\\cdotMcyc\)↓\\downarrowDOSA\[[18](https://arxiv.org/html/2606.02859#bib.bib18)\]80\.2Complete agent \(Gemma\-4\-31B\-it\)43\.1EoM\(Gemma\-4\-31B\-it\)39\.3

### 3\.3Beyond Multiple Agents: The Role of Economic Ingredients

The gains are not explained by merely having multiple agents: they depend on the economic dynamics that allocate control, transfer value, remove unproductive agents, and propagate successful ones\. The ablations in[Table˜2](https://arxiv.org/html/2606.02859#S3.T2)show that weakening these dynamics consistently reduces performance\.

On MATH, the original system achieves the strongest partial\-agent performance, with 43\.9 mean accuracy and 57\.0 best\-run accuracy\. Perturbing the economic parameters lowers performance: increasing rent, decreasing rewards, or increasing rewards reduces the mean to 39\.0–41\.8 and the best run to 44\.0–47\.0\. This indicates that performance depends on the balance between reward inflow, rent pressure, and agent survival\.

On Finance\-Agent\-Bench, the full system again achieves the strongest overall result, with 52\.5 mean accuracy and 65\.0 best\-run accuracy\. Removing exploration causes a large drop to 26\.0 mean and 40\.0 best accuracy, while removing exploitation lowers the mean to 33\.5\. Removing auctions also underperforms the full system, reaching 48\.0 mean and 58\.5 best accuracy\. These results suggest that the population needs both sides of the economic process: exploration introduces new candidate agents with awareness of past failures, while exploitation propagates successful ones\.

Cloudcast provides a complementary comparison\. In[Figure˜4](https://arxiv.org/html/2606.02859#S3.F4)\(c\),EoMreaches a best cost of 673, while the best\-of\-NNmulti\-agent baseline reaches only 999\. Since this baseline uses multiple agents but does not evolve them through market selection, the gap shows that repeated multi\-agent sampling alone is insufficient\. Thus, economic dynamics are not incidental implementation details; they are the mechanism that turns a collection of partial agents into an adaptive society\.

Table 2:Ablations on MATH \(left\) and Finance\-Agent\-Bench \(right\)\.Sensitivity of performance to economic parameters \(left\) and component removal \(right\)\. In both cases, the full/original system achieves the strongest overall results\.Agent ConfigurationMean \(%\)Best \(%\)Complete/51\.951\.9Constrainedlarge rent \(×10\\times 10\)41\.847\.0small reward \(×0\.2\\times 0\.2\)39\.044\.0large reward \(×4\\times 4\)40\.947\.0original43\.957\.0
Agent ConfigurationMean \(%\)Best \(%\)Complete/45\.045\.0Constrainedw/o auction48\.058\.5w/o exploration26\.040\.0w/o exploitation33\.560\.0full52\.565\.0

### 3\.4How Does the Economy Improve Performance?

We show that the economy improves performance by selecting useful action chains, reproducing successful lineages, eliminating unproductive agents, and gradually sharpening both agent prompts and interaction topology\.

What changes inside the society as performance improves?The training dynamics reveal thatEoMimproves by reshaping both the population and the agents themselves\. On Finance\-Agent\-Bench,[Figure˜4](https://arxiv.org/html/2606.02859#S3.F4)\(a\) shows a non\-monotonic but improving trajectory:EoMstarts at 45\.0, dips during early exploration, then recovers and finishes at 60\.0\. This pattern is consistent with a market that first reallocates control and tests alternative specialists before converging to stronger coordination\.

Accelerator design gives a more direct view of the economic mechanism\. In[Figure˜5](https://arxiv.org/html/2606.02859#S3.F5), per\-agent wealth trajectories show how useful lineages survive while ineffective ones disappear\. In panel \(a\), Historian children born during the run quickly lose wealth and are removed, indicating that the inherited mutation could not earn back its rent\. In panel \(b\), a Planner lineage spawns two good\-birth offspring and continues to dominate the auction, while a Historian\-line bad\-birth child eventually bankrupts\. In panel \(c\), a successful Historian lineage and a failing Executor lineage evolve in parallel\. Across these cases, wealth concentrates on agents that repeatedly contribute to record\-breaking submissions, while low\-value agents are pushed toward bankruptcy\. In short, the economy improves performance by turning trajectory\-level success into population\-level selection\.

![Refer to caption](https://arxiv.org/html/2606.02859v1/x7.png)Figure 5:Training dynamics in accelerator design\.Per\-agent wealth on three representative ResNet\-50 kernels\. Wealth flow to agents that produce new EDP records; rent uniformly deducts wealth\. Periodic births spawn*good\-birth*children \(⋆\\star,*exploitation*: mutated from the richest agent\) and*bad\-birth*children \(\+\+,*exploration*: amended from the weakest\); wealth<0<0triggers bankruptcy \(×\\times\)\. Shaded bands are rolling±1​σ\\pm 1\\sigma\.\(a\)Both Historian descendants bankrupt—inherited bias fails market pressure\.\(b\)A Planner lineage reproduces twice while a Historian bad\-birth child eventually fails\.\(c\)A strong Historian and a struggling Executor lineage co\-exist\.Does the society learn reusable structure?We show that the society learns more than task\-specific answer traces\.[Figure˜7](https://arxiv.org/html/2606.02859#S3.F7)\(a\) shows thatEoMachieves a2\.2×2\.2\\timesgeometric\-mean EDP gain overDOSAacross all 24 ResNet\-50 convolution kernels, with much larger gains on the hardest kernels:37\.5×37\.5\\times,26\.3×26\.3\\times,17\.3×17\.3\\times, and12\.0×12\.0\\timesonConv14, 16, 17, and 4\. These gains are structured rather than uniform\. The hardest kernels are the1×11\{\\times\}1convolutions inside ResNet\-50’s bottleneck blocks, which have large input\- and output\-channel counts but small spatial dimensions\. For such shapes, an*output\-stationary*dataflow, which holds each output partial sum in fast on\-chip storage and accumulates contributions along the input\-channel dimension, is a known effective design pattern\. Importantly,EoMis not given this motif as a template\. The auction rewards only EDP record\-breaks, with no per\-kernel labels, dataflow\-specific reward shaping, or hand\-coded output\-stationary preference\. Nevertheless, across the strongest solutions, the population repeatedly converges on the same tiling pattern, recovering a transferable hardware–software co\-design heuristic that DOSA misses\. Thus, market selection can discover reusable domain structure when successful agents are allowed to accumulate wealth and propagate through mutation\.

Beyond this domain\-specific behavior, we observe two broader forms of adaptation\. The population evolves sharper prompt\-level strategies, as illustrated in Appendix[E\.1](https://arxiv.org/html/2606.02859#A5.SS1), enables cross\-domain transfer, as shown in Appendix[E\.2](https://arxiv.org/html/2606.02859#A5.SS2)and reorganizes its interaction topology, as shown in Appendix[E\.3](https://arxiv.org/html/2606.02859#A5.SS3)\. Appendix[F](https://arxiv.org/html/2606.02859#A6)traces both adaptations in a Cloudcast run\. Together, these analyses suggest that economic dynamics shape both internal agent policies and macroscopic social structure\.

### 3\.5Robustness and Generalization

We show that the evolving society is robust in three complementary senses: skills learned on easier tasks transfer to harder ones, performance depends on curriculum but does not collapse under reversed ordering, and adding a complete generalist does not destroy specialization\.

Do learned behaviors transfer from easier tasks to harder ones?On MATH, training on an easy\-to\-hard stream improves not only the easier levels encountered earlier in training, but also harder levels that are initially beyond the agents’ capabilities\. As shown in[Figure˜6](https://arxiv.org/html/2606.02859#S3.F6), both Llama\-3\.1\-8B and Gemma\-2\-9B improve on every difficulty band\. The largest gains appear on Levels 1–3, where Llama\-3\.1\-8B rises to roughly 55–70% and Gemma\-2\-9B to roughly 45–65%\. Importantly, improvement also transfers to Level 5: performance rises from around 10% at the beginning to about 20% by the end for both backbones\. Thus, local reasoning routines learned on simpler problems can be recomposed on harder ones\.

![Refer to caption](https://arxiv.org/html/2606.02859v1/x8.png)Figure 6:Easy\-to\-hard generalization on MATH\.Test accuracy across MATH difficulty levels during training\. The partial agent population improves not only on the easier levels encountered earlier, but also on harder levels that are initially beyond its capability, indicating that behaviors learned on simple problems can be reused on more difficult ones\.How sensitive is the society to curriculum order?We compare the default easy\-to\-hard curriculum with a reversed hard\-to\-easy schedule\.[Figure˜7](https://arxiv.org/html/2606.02859#S3.F7)\(b\) shows that both schedules improve quickly at the beginning, but the easy\-to\-hard curriculum stays ahead for most of training and finishes clearly higher, at roughly 57% versus about 47% for the reversed schedule\. The reversed curriculum plateaus in the low 40s for much of training, while the easy\-to\-hard curriculum continues improving late in training\. This indicates that partial specialists benefit from first mastering reusable local routines before confronting the hardest problems, although the system still improves under the reversed order\.

![Refer to caption](https://arxiv.org/html/2606.02859v1/x9.png)

\(a\) Per\-kernel accelerator EDP

![Refer to caption](https://arxiv.org/html/2606.02859v1/x10.png)

\(b\) Curriculum learning on MATH

![Refer to caption](https://arxiv.org/html/2606.02859v1/x11.png)

\(c\) Finance research with a generalist

Figure 7:Mechanism, robustness, and generalization analyses\.\(a\)Per\-kernel EDP on ResNet\-50\. Best EDP found byDOSA,ReAct, andEoMon log scale; lower is better\.\(b\)Comparison between the default easy\-to\-hard curriculum and a reversed hard\-to\-easy curriculum\.\(c\)Adding a strong generalist agent with access to all tools does not automatically dominate specialized agents\.Can a complete generalist monopolize the economy?A natural concern is that if a single agent is given complete access to the task interface, it may dominate the market and collapse the society into a monolithic system\. We test this in Finance\-Agent\-Bench by adding a generalist agent with access to all tools alongside the partial specialist agents\. As shown in[Figure˜7](https://arxiv.org/html/2606.02859#S3.F7)\(c\), the generalist briefly expands around tasks 11–12, but then contracts and returns to a single agent for the remainder of training\. Meanwhile, specialized populations such as Edgar and Tavily continue to grow, reaching roughly 5–8 agents late in training\. This suggests that broader access does not automatically imply market dominance\. The economy rewards local value: a specialist whose wake\-up condition, tool use, and evidence standard are tuned to a narrow subproblem can outcompete a generalist whose prompt budget is spread across many heterogeneous responsibilities\. We provide a detailed analysis in Appendix[G](https://arxiv.org/html/2606.02859#A7), where we show that specialists evolve sharper local decision rules, while the generalist tends to accumulate broad but diluted procedural instructions\. Thus, the society remains decentralized not because complete agents are forbidden, but because the market continues to favor locally more precise specialists\.

## 4Related Work

Self\-Evolving Agents\.Recent work studies how agents can improve through interaction, adaptation, and accumulated experience rather than relying on fixed centralized control\. In the LLM setting, this includes co\-evolution, iterative refinement, and experience sharing across agent populations\[[7](https://arxiv.org/html/2606.02859#bib.bib7),[45](https://arxiv.org/html/2606.02859#bib.bib45),[9](https://arxiv.org/html/2606.02859#bib.bib9)\]\. Related work on reward\-driven self\-organization shows that structured behaviors can emerge from local feedback alone\[[54](https://arxiv.org/html/2606.02859#bib.bib54)\]\. More broadly, rich ecological environments and evolutionary pressures have been shown to induce increasingly complex adaptive behaviors, suggesting that agent capabilities can emerge through continual interaction with other agents and the environment\[[4](https://arxiv.org/html/2606.02859#bib.bib4)\]\. For more related work, please refer to Appendix[B](https://arxiv.org/html/2606.02859#A2)\.

## 5Limitations, Conclusions, and Future Work

This work shows that decentralized economic interactions can turn a population of agents into adaptive collective intelligence\. Across diverse domains, auctions, transactions, and wealth\-based selection induce specialization and coordination without centralized control, suggesting that designing incentives is a powerful alternative to designing individual agents\.

A key limitation is that adaptation occurs only in prompt space with a frozen backbone, which may restrict capability growth in tasks requiring new skills or representations\. Future work can extend this framework to parameter\-space training, hybrid adaptation, and broader agent backbones, including multimodal or embodied systems\.

## Acknowledgments and Disclosure of Funding

This work was supported in part by a gift from the Chan Zuckerberg Initiative Foundation to establish the Kempner Institute at Harvard University\.

## References

- Baum \[1996\]Eric B\. Baum\.Toward a model of mind as a laissez\-faire economy of idiots\.In*Proceedings of the 13th International Conference on Machine Learning*, pages 20–27, 1996\.
- Baum \[1999\]Eric B Baum\.Toward a model of intelligence as an economy of agents\.*Machine Learning*, 35\(2\):155–185, 1999\.
- Baum and Durdanovic \[2000\]Eric B Baum and Igor Durdanovic\.Evolution of cooperative problem solving in an artificial economy\.*Neural Computation*, 12\(12\):2743–2775, 2000\.
- Bejjani et al\. \[2025\]Joseph Bejjani, Chase Van Amburg, Chengrui Wang, Chloe Huangyuan Su, Sarah M Pratt, Yasin Mazloumi, Naeem Khoshnevis, Sham M Kakade, Kianté Brantley, and Aaron Walsman\.The emergence of complex behavior in large\-scale ecological environments\.*arXiv preprint arXiv:2510\.18221*, 2025\.
- Bigeard et al\. \[2025\]Antoine Bigeard, Langston Nashold, Rayan Krishnan, and Shirley Wu\.Finance agent benchmark: Benchmarking llms on real\-world financial research tasks\.*arXiv preprint arXiv:2508\.00828*, 2025\.
- Cemri et al\. \[2025\]Mert Cemri, Melissa Z Pan, Shuyi Yang, Lakshya A Agrawal, Bhavya Chopra, Rishabh Tiwari, Kurt Keutzer, Aditya Parameswaran, Dan Klein, Kannan Ramchandran, et al\.Why do multi\-agent llm systems fail?*arXiv preprint arXiv:2503\.13657*, 2025\.
- Chen et al\. \[2025\]Yixing Chen, Yiding Wang, Siqi Zhu, Haofei Yu, Tao Feng, Muhan Zhang, Mostofa Patwary, and Jiaxuan You\.Multi\-agent evolve: Llm self\-improve through co\-evolution\.*arXiv preprint arXiv:2510\.23595*, 2025\.
- Cheng et al\. \[2025\]Audrey Cheng, Shu Liu, Melissa Pan, Zhifei Li, Bowen Wang, Alex Krentsel, Tian Xia, Mert Cemri, Jongseok Park, Shuo Yang, et al\.Barbarians at the gate: How ai is upending systems research\.*arXiv preprint arXiv:2510\.06189*, 2025\.
- Dai et al\. \[2026\]X\. Dai, Y\. Zhu, et al\.Geoevolver: A self\-evolving multi\-agent system for earth observation\.*arXiv preprint arXiv:2602\.02559*, 2026\.
- Dias et al\. \[2006\]M Bernardine Dias, Robert Zlot, Nidhi Kalra, and Anthony Stentz\.Market\-based multirobot coordination: A survey and analysis\.*Proceedings of the IEEE*, 94\(7\):1257–1270, 2006\.
- Du et al\. \[2024\]Yilun Du, Shuang Li, Antonio Torralba, Joshua B Tenenbaum, and Igor Mordatch\.Improving factuality and reasoning in language models through multiagent debate\.In*Forty\-first international conference on machine learning*, 2024\.
- Duetting et al\. \[2024\]Paul Duetting, Vahab Mirrokni, Renato Paes Leme, Haifeng Xu, and Song Zuo\.Mechanism design for large language models\.In*Proceedings of the ACM Web Conference 2024*, pages 144–155, 2024\.
- Genc et al\. \[2021\]Hasan Genc, Seah Kim, Alon Amid, Ameer Haj\-Ali, Vighnesh Iyer, Pranav Prakash, Jerry Zhao, Daniel Grubb, Harrison Liew, Howard Mao, et al\.Gemmini: Enabling systematic deep\-learning architecture evaluation via full\-stack integration\.In*2021 58th ACM/IEEE Design Automation Conference \(DAC\)*, pages 769–774\. IEEE, 2021\.
- Hayek \[1945\]Friedrich A\. Hayek\.The use of knowledge in society\.*The American Economic Review*, 35\(4\):519–530, 1945\.
- Hayek \[1973\]Friedrich A\. Hayek\.*Law, Legislation and Liberty*\.University of Chicago Press, 1973\.
- Hendrycks et al\. \[2021\]Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt\.Measuring mathematical problem solving with the math dataset\.*arXiv preprint arXiv:2103\.03874*, 2021\.
- Holland \[1985\]John H\. Holland\.Properties of the bucket brigade algorithm\.In*Proceedings of the 1st International Conference on Genetic Algorithms*, 1985\.
- Hong et al\. \[2023a\]Charles Hong, Qijing Huang, Grace Dinh, Mahesh Subedar, and Yakun Sophia Shao\.Dosa: Differentiable model\-based one\-loop search for dnn accelerators\.In*Proceedings of the 56th Annual IEEE/ACM International Symposium on Microarchitecture*, pages 209–224, 2023a\.
- Hong et al\. \[2023b\]Sirui Hong, Mingchen Zhuge, Jonathan Chen, Xiawu Zheng, Yuheng Cheng, Jinlin Wang, Ceyao Zhang, Zili Wang, Steven Ka Shing Yau, Zijuan Lin, et al\.Metagpt: Meta programming for a multi\-agent collaborative framework\.In*The twelfth international conference on learning representations*, 2023b\.
- Iqbal et al\. \[2022\]Shariq Iqbal, Robby Costales, and Fei Sha\.Alma: Hierarchical learning for composite multi\-agent tasks\.*Advances in neural information processing systems*, 35:7155–7166, 2022\.
- Kim et al\. \[2024\]Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Ashwin Balakrishna, Suraj Nair, Rafael Rafailov, Ethan Foster, Grace Lam, Pannag Sanketi, et al\.Openvla: An open\-source vision\-language\-action model\.*arXiv preprint arXiv:2406\.09246*, 2024\.
- Kim et al\. \[2025\]Yubin Kim, Ken Gu, Chanwoo Park, Chunjong Park, Samuel Schmidgall, A Ali Heydari, Yao Yan, Zhihan Zhang, Yuchen Zhuang, Yun Liu, et al\.Towards a science of scaling agent systems\.*arXiv preprint arXiv:2512\.08296*, 2025\.
- Lifshitz et al\. \[2025\]Shalev Lifshitz, Sheila A McIlraith, and Yilun Du\.Multi\-agent verification: Scaling test\-time compute with multiple verifiers\.*arXiv preprint arXiv:2502\.20379*, 2025\.
- \[24\]Shuo Liu, Tianle Chen, and Christopher Amato\.Improved multi\-agent collaboration with multi\-turn reinforcement learning\.In*First Workshop on Multi\-Turn Interactions in Large Language Models*\.
- Long et al\. \[2026\]Carol Xuan Long, David Simchi\-Levi, Feng Zhu, Huangyuan Su, Andre P\. Calmon, and Flavio P\. Calmon\.Reliability and effectiveness of autonomous ai agents in supply chain management\.*arXiv preprint arXiv:2605\.17036*, 2026\.
- Minsky \[1988\]Marvin Minsky\.*The Society of Mind*\.Simon and Schuster, 1988\.
- Novikov et al\. \[2025\]Alexander Novikov, Ngân Vũ, Marvin Eisenberger, Emilien Dupont, Po\-Sen Huang, Adam Zsolt Wagner, Sergey Shirobokov, Borislav Kozlovskii, Francisco JR Ruiz, Abbas Mehrabian, et al\.Alphaevolve: A coding agent for scientific and algorithmic discovery\.*arXiv preprint arXiv:2506\.13131*, 2025\.
- Omidshafiei et al\. \[2017\]Shayegan Omidshafiei, Jason Pazis, Christopher Amato, Jonathan P How, and John Vian\.Deep decentralized multi\-task multi\-agent reinforcement learning under partial observability\.In*International conference on machine learning*, pages 2681–2690\. PMLR, 2017\.
- Park et al\. \[2025\]Chanwoo Park, Seungju Han, Xingzhi Guo, Asuman E Ozdaglar, Kaiqing Zhang, and Joo\-Kyung Kim\.Maporl: Multi\-agent post\-co\-training for collaborative large language models with reinforcement learning\.In*Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pages 30215–30248, 2025\.
- Patil et al\. \[2024\]Shishir G Patil, Tianjun Zhang, Xin Wang, and Joseph E Gonzalez\.Gorilla: Large language model connected with massive apis\.*Advances in Neural Information Processing Systems*, 37:126544–126565, 2024\.
- Prakash et al\. \[2025\]Shvetank Prakash, Andrew Cheng, Arya Tschand, Mark Mazumder, Varun Gohil, Jeffrey Ma, Jason Yik, Zishen Wan, Jessica Quaye, Elisavet Lydia Alvanaki, et al\.Quarch: A benchmark for evaluating llm reasoning in computer architecture\.*arXiv preprint arXiv:2510\.22087*, 2025\.
- Qian et al\. \[2024\]Chen Qian, Wei Liu, Hongzhang Liu, Nuo Chen, Yufan Dang, Jiahao Li, Cheng Yang, Weize Chen, Yusheng Su, Xin Cong, et al\.Chatdev: Communicative agents for software development\.In*Proceedings of the 62nd annual meeting of the association for computational linguistics \(volume 1: Long papers\)*, pages 15174–15186, 2024\.
- Qu et al\. \[2026\]Ao Qu, Han Zheng, Zijian Zhou, Yihao Yan, Yihong Tang, Shao Yong Ong, Fenglu Hong, Kaichen Zhou, Chonghe Jiang, Minwei Kong, et al\.Coral: Towards autonomous multi\-agent evolution for open\-ended discovery\.*arXiv preprint arXiv:2604\.01658*, 2026\.
- Schmidhuber \[1989\]Jürgen Schmidhuber\.The neural bucket brigade: A local learning algorithm for dynamic feedforward and recurrent networks\.*Connection Science*, 1\(4\):403–412, 1989\.
- Sharma \[2025\]Asankhaya Sharma\.Openevolve: an open\-source evolutionary coding agent, 2025\.URL[https://github\.com/algorithmicsuperintelligence/openevolve](https://github.com/algorithmicsuperintelligence/openevolve)\.
- Subramaniam et al\. \[2025\]Vighnesh Subramaniam, Yilun Du, Joshua B Tenenbaum, Antonio Torralba, Shuang Li, and Igor Mordatch\.Multiagent finetuning: Self improvement with diverse reasoning chains\.*arXiv preprint arXiv:2501\.05707*, 2025\.
- Sudhir and Tran\-Thanh \[2025\]Abhimanyu Pallavi Sudhir and Long Tran\-Thanh\.Market\-based architectures in rl and beyond\.*Accepted to AAMAS 2025 Blue Sky Track*, abs/2503\.05828, 2025\.
- Sun et al\. \[2022\]Hanbo Sun, Chenyu Wang, Zhenhua Zhu, Xuefei Ning, Guohao Dai, Huazhong Yang, and Yu Wang\.Gibbon: Efficient co\-exploration of nn model and processing\-in\-memory architecture\.In*2022 Design, Automation & Test in Europe Conference & Exhibition \(DATE\)*, pages 867–872\. IEEE, 2022\.
- Sun et al\. \[2025\]Weiwei Sun, Miao Lu, Zhan Ling, Kang Liu, Xuesong Yao, Yiming Yang, and Jiecao Chen\.Scaling long\-horizon llm agent via context\-folding\.*arXiv preprint arXiv:2510\.11967*, 2025\.
- Tschand et al\. \[2026\]Arya Tschand, Chenyu Wang, Zishen Wan, Andrew Cheng, Ioana Cristescu, Kevin He, Howard Huang, Alexander Ingare, Akseli Kangaslahti, Sara Kangaslahti, Theo Lebryk, Hongjin Lin, Jeffrey Jian Ma, Alexandru Meterez, Clara Mohri, Depen Morwani, Sunny Qin, Roy Rinberg, Paula Rodriguez\-Diaz, Alyssa Mia Taliotis, Pernille Undrum Fathi, Rosie Zhao, Todd Zhou, and Vijay Janapa Reddi\.Genai for systems: Recurring challenges and design principles from software to silicon, 2026\.URL[https://arxiv\.org/abs/2602\.15241](https://arxiv.org/abs/2602.15241)\.
- Wang et al\. \[2025a\]Chenyu Wang, Zishen Wan, Hao Kang, Emma Chen, Zhiqiang Xie, Tushar Krishna, Vijay Janapa Reddi, and Yilun Du\.Slm\-mux: Orchestrating small language models for reasoning\.*arXiv preprint arXiv:2510\.05077*, 2025a\.
- Wang \[2026\]Jian Sheng Wang\.Aesp: A human\-sovereign economic protocol for ai agents with privacy\-preserving settlement\.*arXiv preprint arXiv:2603\.00318*, 2026\.
- Wang et al\. \[2026\]Miles Wang, Robi Lin, Kat Hu, Joy Jiao, Neil Chowdhury, Ethan Chang, and Tejal Patwardhan\.Frontierscience: Evaluating ai’s ability to perform expert\-level scientific tasks\.*arXiv preprint arXiv:2601\.21165*, 2026\.
- Wang et al\. \[2025b\]Yiping Wang, Shao\-Rong Su, Zhiyuan Zeng, Eva Xu, Liliang Ren, Xinyu Yang, Zeyi Huang, Xuehai He, Luyao Ma, Baolin Peng, et al\.Thetaevolve: Test\-time learning on open problems\.*arXiv preprint arXiv:2511\.23473*, 2025b\.
- Weng et al\. \[2026\]Zhaotian Weng, Antonis Antoniades, Deepak Nathani, Zhen Zhang, Xiao Pu, and Xin Eric Wang\.Group\-evolving agents: Open\-ended self\-improvement via experience sharing\.*arXiv preprint arXiv:2602\.04837*, 2026\.
- Wu et al\. \[2024\]Qingyun Wu, Gagan Bansal, Jieyu Zhang, Yiran Wu, Beibin Li, Erkang Zhu, Li Jiang, Xiaoyun Zhang, Shaokun Zhang, Jiale Liu, et al\.Autogen: Enabling next\-gen llm applications via multi\-agent conversations\.In*First conference on language modeling*, 2024\.
- Xu \[2026\]Minghui Xu\.The agent economy: A blockchain\-based foundation for autonomous ai agents\.2026\.NeurIPS 2026\.
- Xue et al\. \[2025\]Xiangyuan Xue, Yifan Zhou, Guibin Zhang, Zaibin Zhang, Yijiang Li, Chen Zhang, Zhenfei Yin, Philip Torr, Wanli Ouyang, and Lei Bai\.Comas: Co\-evolving multi\-agent systems via interaction rewards\.*arXiv preprint arXiv:2510\.08529*, 2025\.
- Yang et al\. \[2025\]Yingxuan Yang, Huacan Chai, Shuai Shao, Yuanyi Song, Siyuan Qi, Renting Rui, and Weinan Zhang\.Agentnet: Decentralized evolutionary coordination for llm\-based multi\-agent systems, 2025\.URL[https://arxiv\.org/abs/2504\.00587](https://arxiv.org/abs/2504.00587)\.
- Yao et al\. \[2022\]Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R Narasimhan, and Yuan Cao\.React: Synergizing reasoning and acting in language models\.In*The eleventh international conference on learning representations*, 2022\.
- Ye et al\. \[2026\]Seonghyeon Ye, Yunhao Ge, Kaiyuan Zheng, Shenyuan Gao, Sihyun Yu, George Kurian, Suneel Indupuru, You Liang Tan, Chuning Zhu, Jiannan Xiang, et al\.World action models are zero\-shot policies\.*arXiv preprint arXiv:2602\.15922*, 2026\.
- Yuksekgonul et al\. \[2026\]Mert Yuksekgonul, Daniel Koceja, Xinhao Li, Federico Bianchi, Jed McCaleb, Xiaolong Wang, Jan Kautz, Yejin Choi, James Zou, Carlos Guestrin, and Yu Sun\.Learning to discover at test time\.*arXiv preprint*, 2026\.
- Zhang et al\. \[2026\]Miao Zhang, Junsik Kim, Siyuan Xiang, Jian Gao, and Cheng Cao\.Dynamic role assignment for multi\-agent debate\.*arXiv preprint arXiv:2601\.17152*, 2026\.
- Zhou et al\. \[2025\]Heng Zhou, Hejia Geng, Xiangyuan Xue, Li Kang, Yiran Qin, Zhiyong Wang, Zhenfei Yin, and Lei Bai\.Reso: A reward\-driven self\-organizing llm\-based multi\-agent system for reasoning tasks\.In*Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing*, pages 15990–16009, 2025\.
- Zhu and Du \[2025\]Andy Zhu and Yingjun Du\.A role\-aware multi\-agent framework for financial education question answering with llms\.*arXiv preprint arXiv:2509\.09727*, 2025\.

## Appendix \- Table of Contents

## Appendix APseudo Code

Training and evaluation pseudo code forEoMare shown below:

Algorithm 1The Training Loop1:task stream

𝒟\\mathcal\{D\}, initial population

𝒫0\\mathcal\{P\}\_\{0\}, wake\-up oracle

ω\\omega
2:base bid

b0b\_\{0\}, novice premium

ϵ\\epsilon, step cap

SS, trial cap

TT
3:bounds

\[Nmin,Nmax\]\[N\_\{\\min\},N\_\{\\max\}\]; rent

ρ\\rhoevery

KrK\_\{r\}tasks

4:birth probs

pa,pbp\_\{a\},p\_\{b\}\(bankruptcy\),

pgp\_\{g\}\(periodic\); period

KbK\_\{b\}, batch

BB
5:

𝒫←𝒫0\\mathcal\{P\}\\leftarrow\\mathcal\{P\}\_\{0\}
6:forepisode

e=1,2,…e=1,2,\\dotsover

τ∼𝒟\\tau\\sim\\mathcal\{D\}do

7:

σ←\{a↦Snapshot​\(a\):a∈𝒫\}\\sigma\\leftarrow\\\{a\\mapsto\\textsc\{Snapshot\}\(a\):a\\in\\mathcal\{P\}\\\};

ℬe←∅\\mathcal\{B\}\_\{e\}\\leftarrow\\emptyset⊳\\trianglerighte\.g\. wealth, capability, …

8:fortrial

t=1,…,Tt=1,\\dots,Tdo⊳\\trianglerightreplay until no bankruptcy

9:

alast,H←a^\{\\text\{last\}\},H\\leftarrowPlanning\(

τ,𝒫\\tau,\\mathcal\{P\}\);

alast\.wealth\+=env\.reward\(\)a^\{\\text\{last\}\}\.\\text\{wealth\}\\mathrel\{\+\}=\\text\{env\.reward\}\(\)⊳\\trianglerightoutcome reward to final actor

10:

ℬ←\{a:a\.wealth≤0\}\\mathcal\{B\}\\leftarrow\\\{a:a\.\\text\{wealth\}\\leq 0\\\}
11:if

ℬ=∅\\mathcal\{B\}=\\emptysetthenbreak

12:endif

13:

𝒫←𝒫∖ℬ\\mathcal\{P\}\\leftarrow\\mathcal\{P\}\\setminus\\mathcal\{B\};

ℬe∪=ℬ\\mathcal\{B\}\_\{e\}\\mathrel\{\\cup\}=\\mathcal\{B\}
14:for all

a∈𝒫a\\in\\mathcal\{P\}doRestore\(

a,σ​\(a\)a,\\sigma\(a\)\)

15:endfor⊳\\trianglerightrollback survivors

16:endfor

17:if

emodKr=0e\\bmod K\_\{r\}=0then⊳\\trianglerightperiodic rent

18:

a\.wealth\-=ra\.\\text\{wealth\}\\mathrel\{\-\}=rfor all

a∈𝒫a\\in\\mathcal\{P\}
19:

ℬ′←\{a∈𝒫:a\.wealth≤0\}\\mathcal\{B\}^\{\\prime\}\\leftarrow\\\{a\\in\\mathcal\{P\}:a\.\\text\{wealth\}\\leq 0\\\};

𝒫←𝒫∖ℬ′\\mathcal\{P\}\\leftarrow\\mathcal\{P\}\\setminus\\mathcal\{B\}^\{\\prime\};

ℬe∪=ℬ′\\mathcal\{B\}\_\{e\}\\mathrel\{\\cup\}=\\mathcal\{B\}^\{\\prime\}
20:endif

21:Adaptation\(

𝒫,ℬe,e\\mathcal\{P\},\\mathcal\{B\}\_\{e\},e\)⊳\\trianglerightbankruptcy \+ periodic births, replenish toNminN\_\{\\min\}

22:endfor

23:

24:procedurePlanning\(

τ,𝒫\\tau,\\mathcal\{P\}\)

25:init env on

τ\\tau;

H←∅H\\leftarrow\\emptyset;

a−←⊥a^\{\-\}\\leftarrow\\bot
26:for

s=1,…,Ss=1,\\dots,Sdo

27:

ℰ←\{a∈𝒫:ω​\(a,env\)\}\\mathcal\{E\}\\leftarrow\\\{a\\in\\mathcal\{P\}:\\omega\(a,\\text\{env\}\)\\\};if

ℰ=∅\\mathcal\{E\}=\\emptysetbreak

28:SetBids\(

ℰ\\mathcal\{E\}\)⊳\\trianglerightpluggable: e\.g\. novices enter atmaxvet⁡b\+ϵ\\max\_\{\\text\{vet\}\}b\+\\epsilonand lock

29:

a∗←arg⁡maxa∈ℰ⁡baa^\{\*\}\\leftarrow\\arg\\max\_\{a\\in\\mathcal\{E\}\}b\_\{a\}\(random tie\-break\)

30:

a∗\.wealth\-=ba∗a^\{\*\}\.\\text\{wealth\}\\mathrel\{\-\}=b\_\{a^\{\*\}\};if

a−≠⊥a^\{\-\}\\neq\\bot:

a−\.wealth\+=ba∗a^\{\-\}\.\\text\{wealth\}\\mathrel\{\+\}=b\_\{a^\{\*\}\}
31:env\.apply\(

a∗\.act​\(env\)a^\{\*\}\.\\text\{act\}\(\\text\{env\}\)\);

H←H∪\{a∗\}H\\leftarrow H\\cup\\\{a^\{\*\}\\\};

a−←a∗a^\{\-\}\\leftarrow a^\{\*\}
32:ifenv terminatedbreak

33:endfor

34:return

a−,Ha^\{\-\},H
35:endprocedure

36:

37:procedureAdaptation\(

𝒫,ℬe,e\\mathcal\{P\},\\mathcal\{B\}\_\{e\},e\)

38:for all

a∈ℬea\\in\\mathcal\{B\}\_\{e\}with

\|𝒫\|<Nmax\|\\mathcal\{P\}\|<N\_\{\\max\}do

39:

u∼𝒰​\(0,1\)u\\sim\\mathcal\{U\}\(0,1\)
40:if

u<pau<p\_\{a\}:

𝒫←𝒫∪\{Mutate​\(richest​\(𝒫\)\)\}\\mathcal\{P\}\\leftarrow\\mathcal\{P\}\\cup\\\{\\textsc\{Mutate\}\(\\text\{richest\}\(\\mathcal\{P\}\)\)\\\}
41:elif

u<pa\+pbu<p\_\{a\}\+p\_\{b\}:

𝒫←𝒫∪\{Amend​\(a\)\}\\mathcal\{P\}\\leftarrow\\mathcal\{P\}\\cup\\\{\\textsc\{Amend\}\(a\)\\\}
42:endfor

43:if

emodKb=0e\\bmod K\_\{b\}=0then

44:for

i=1,…,Bi=1,\\dots,Bwith

\|𝒫\|<Nmax\|\\mathcal\{P\}\|<N\_\{\\max\}do

45:

a′←Mutate​\(richest​\(𝒫\)\)a^\{\\prime\}\\leftarrow\\textsc\{Mutate\}\(\\text\{richest\}\(\\mathcal\{P\}\)\)with prob\.

pgp\_\{g\}, else

Amend​\(poorest​\(𝒫\)\)\\textsc\{Amend\}\(\\text\{poorest\}\(\\mathcal\{P\}\)\)
46:

𝒫←𝒫∪\{a′\}\\mathcal\{P\}\\leftarrow\\mathcal\{P\}\\cup\\\{a^\{\\prime\}\\\}
47:endfor

48:endif

49:while

\|𝒫\|<Nmin\|\\mathcal\{P\}\|<N\_\{\\min\}do

50:

𝒫←𝒫∪\{Mutate​\(a0\)\}\\mathcal\{P\}\\leftarrow\\mathcal\{P\}\\cup\\\{\\textsc\{Mutate\}\(a\_\{0\}\)\\\}for some template

a0∈𝒫0a\_\{0\}\\in\\mathcal\{P\}\_\{0\}
51:endwhile

52:endprocedure

Algorithm 2The Evaluation Loop1:trained population

𝒫∗\\mathcal\{P\}^\{\*\}, test stream

𝒟test\\mathcal\{D\}\_\{\\text\{test\}\}, wake\-up oracle

ω\\omega
2:step cap

SS; thread\-pool size

WW
3:

𝒫←𝒫∗\\mathcal\{P\}\\leftarrow\\mathcal\{P\}^\{\*\}; freeze

\{wa,ba,statusa\}a∈𝒫\\\{w\_\{a\},b\_\{a\},\\text\{status\}\_\{a\}\\\}\_\{a\\in\\mathcal\{P\}\}⊳\\trianglerightno payments, no rewards, no rent, no births

4:fortask

τ∈𝒟test\\tau\\in\\mathcal\{D\}\_\{\\text\{test\}\}in parallel\(workers

WW\)do

5:

𝒫τ←\\mathcal\{P\}\_\{\\tau\}\\leftarrowSnapshot\(

𝒫\\mathcal\{P\}\)⊳\\trianglerightthread\-local deep copy; no shared mutation

6:

alast,H←a^\{\\text\{last\}\},H\\leftarrowPlanning\(

τ,𝒫τ\\tau,\\mathcal\{P\}\_\{\\tau\}\)⊳\\trianglerightnoalast\.wealtha^\{\\text\{last\}\}\.\\text\{wealth\}update

7:record

\{τ,H,env\.terminated,env\.score​\(\)\}\\\{\\tau,H,\\text\{env\.terminated\},\\text\{env\.score\}\(\)\\\}
8:endfor

9:

10:procedurePlanning\(

τ,𝒫\\tau,\\mathcal\{P\}\)

11:init env on

τ\\tau;

H←∅H\\leftarrow\\emptyset;

a−←⊥a^\{\-\}\\leftarrow\\bot
12:for

s=1,…,Ss=1,\\dots,Sdo

13:

ℰ←\{a∈𝒫:ω​\(a,env\)\}\\mathcal\{E\}\\leftarrow\\\{a\\in\\mathcal\{P\}:\\omega\(a,\\text\{env\}\)\\\};if

ℰ=∅\\mathcal\{E\}=\\emptysetbreak

14:\(skipSetBids\(ℰ\\mathcal\{E\}\): bidsbab\_\{a\}are frozen from training; no novice→\\toveteran promotion\)

15:

a∗←arg⁡maxa∈ℰ⁡baa^\{\*\}\\leftarrow\\arg\\max\_\{a\\in\\mathcal\{E\}\}b\_\{a\}\(random tie\-break\)

16:\(skip wealth transfer:a∗\.wealth\-=ba∗a^\{\*\}\.\\text\{wealth\}\\mathrel\{\-\}=b\_\{a^\{\*\}\}anda−\.wealth\+=ba∗a^\{\-\}\.\\text\{wealth\}\\mathrel\{\+\}=b\_\{a^\{\*\}\}are no\-ops\)

17:env\.apply\(

a∗\.act​\(env\)a^\{\*\}\.\\text\{act\}\(\\text\{env\}\)\);

H←H∪\{a∗\}H\\leftarrow H\\cup\\\{a^\{\*\}\\\};

a−←a∗a^\{\-\}\\leftarrow a^\{\*\}⊳\\trianglerightno reward credited toa∗a^\{\*\}

18:ifenv terminatedbreak

19:endfor

20:return

a−,Ha^\{\-\},H
21:endprocedure

22:

## Appendix BExtended Related Works

Multi\-Agent Systems\.A parallel line of work examines how multiple agents coordinate, specialize, and solve tasks collectively\. Recent work in the LLM era frames multi\-agent orchestration as an “society of minds,” where specialization and coordination emerge through local interactions rather than centralized planning\[[10](https://arxiv.org/html/2606.02859#bib.bib10),[37](https://arxiv.org/html/2606.02859#bib.bib37),[47](https://arxiv.org/html/2606.02859#bib.bib47)\]\. Existing approaches include reinforcement\-learning\-based collaboration\[[24](https://arxiv.org/html/2606.02859#bib.bib24),[29](https://arxiv.org/html/2606.02859#bib.bib29)\], dynamic role assignment and role\-aware specialization\[[53](https://arxiv.org/html/2606.02859#bib.bib53),[55](https://arxiv.org/html/2606.02859#bib.bib55),[25](https://arxiv.org/html/2606.02859#bib.bib25)\], and decentralized evolutionary coordination, where agents dynamically adapt their graph connectivity, specialize through retrieval\-augmented memory, and route tasks without a central orchestrator\[[49](https://arxiv.org/html/2606.02859#bib.bib49)\]\. Other directions include multi\-agent debate and verification for improving reasoning and test\-time scaling\[[11](https://arxiv.org/html/2606.02859#bib.bib11),[23](https://arxiv.org/html/2606.02859#bib.bib23),[41](https://arxiv.org/html/2606.02859#bib.bib41)\], and multi\-agent finetuning with diverse reasoning chains for self\-improvement\[[36](https://arxiv.org/html/2606.02859#bib.bib36)\], alongside broader studies of collaborative and self\-organizing multi\-agent systems\. This perspective also connects to mechanism design for LLMs, which studies how incentives and allocation rules shape agent behavior in strategic environments\[[12](https://arxiv.org/html/2606.02859#bib.bib12)\]\.

Economics and Applications to Agents\.The motivation for decentralized agent systems draws on a long tradition in economics and distributed intelligence\. Smith’s “invisible hand” and Hayek’s “spontaneous order” emphasize how coordinated outcomes can emerge from decentralized actors with only local information\[[15](https://arxiv.org/html/2606.02859#bib.bib15)\]\. Related ideas appear in Minsky’s*Society of Mind*, which views intelligence as arising from the interaction of many simple agents rather than a single centralized reasoner\[[26](https://arxiv.org/html/2606.02859#bib.bib26)\]\. These ideas influenced computational models of coordination and learning, including Holland’s bucket brigade for distributed credit assignment\[[17](https://arxiv.org/html/2606.02859#bib.bib17)\], Schmidhuber’s neural bucket brigade\[[34](https://arxiv.org/html/2606.02859#bib.bib34)\], and Baum’s Hayek Machine, which introduced property rights, wealth accumulation, and competitive selection among simple agents\[[1](https://arxiv.org/html/2606.02859#bib.bib1),[2](https://arxiv.org/html/2606.02859#bib.bib2),[3](https://arxiv.org/html/2606.02859#bib.bib3)\]\. Recent proposals for economic protocols and settlement layers for autonomous AI agents continue this line by formalizing exchange, sovereignty, and privacy\-preserving settlement in modern agent ecosystems\[[42](https://arxiv.org/html/2606.02859#bib.bib42)\]\.

## Appendix CTheoretical Motivations

We use the notation and auction dynamics from[Section˜2](https://arxiv.org/html/2606.02859#S2)\. Agents are prompted LLM policies with fixed bids and wealths, and wealth evolves through the bucket\-brigade transfer rule \([1](https://arxiv.org/html/2606.02859#S2.E1)\)\. This section summarizes why the resulting market dynamics select useful specialists, why outcome\-only reward can be sufficient once the population is strong, and how bucket\-brigade payments provide a structured credit\-assignment signal\.

### C\.1Market selection drives bids toward value

The first result formalizes the basic market\-selection mechanism\. If an agent repeatedly wins at a recurrent context and its bid is below its expected resale\-adjusted payoff, then its wealth has positive drift and it can survive\. Conversely, agents whose bids exceed their expected payoff eventually become insolvent\. Thus bankruptcy and replacement push the surviving bid frontier toward the value of the best available specialist\.

For a recurrent contextxx, letV​\(a,x\)V\(a,x\)denote the expected resale\-adjusted payoff of agentaa, including the immediate reward and the downstream bucket\-brigade payment\. Let

V⋆​\(x\):=supa:ϕa​\(x\)=1V​\(a,x\)V^\{\\star\}\(x\):=\\sup\_\{a:\\phi\_\{a\}\(x\)=1\}V\(a,x\)be the value of the best eligible specialist, and letβ∞​\(x\)\\beta\_\{\\infty\}\(x\)denote the largest bid among agents that win atxxinfinitely often and are never removed\.

###### Theorem 1\(Selection of bids toward expected payoff\)\.

Suppose contextxxis recurrent, payoffs are stationary and bounded, and newly injected agents use the novice bidding rule with resolutionεmax\\varepsilon\_\{\\max\}\. Suppose also that whenever the current solvent frontier atxxis more thanεmax\\varepsilon\_\{\\max\}belowV⋆​\(x\)V^\{\\star\}\(x\), there is a positive probability of injecting an eligible near\-optimal specialist\. Then, almost surely,

V⋆​\(x\)−εmax≤β∞​\(x\)≤V⋆​\(x\)\.V^\{\\star\}\(x\)\-\\varepsilon\_\{\\max\}\\leq\\beta\_\{\\infty\}\(x\)\\leq V^\{\\star\}\(x\)\.Thus the long\-run solvent bid frontier tracks the value of the best prompted specialist, up to the novice\-resolution error\.

This result explains the role of wealth and bankruptcy: the mechanism does not require explicitly fitting value functions\. Instead, agents with overpriced bids lose wealth and are removed, while underpriced useful agents have a chance to accumulate wealth and persist\.

### C\.2Outcome\-only reward is sufficient for a strong population

A natural question is whether the system requires dense process rewards, or whether final outcome reward alone is enough\. The next result shows that outcome\-only reward is sufficient whenever the auction population is already organized so that the winning agent is approximately the best eligible specialist at each reachable history\.

LetJout​\(π\)J^\{\\mathrm\{out\}\}\(\\pi\)denote the expected discounted outcome return of a policyπ\\pi, using only the task\-aligned rewardrtoutr\_\{t\}^\{\\mathrm\{out\}\}\. LetQt⋆​\(ht,u\)Q\_\{t\}^\{\\star\}\(h\_\{t\},u\)be the optimal outcome\-value obtained by selecting agentuuat historyhth\_\{t\}and acting optimally thereafter\.

###### Theorem 2\(Outcome reward suffices for strong agents\)\.

Suppose there existsε≥0\\varepsilon\\geq 0such that for every reachable historyhth\_\{t\}with nonempty eligible setEtE\_\{t\}, the auction winnerat⋆a\_\{t\}^\{\\star\}satisfies

Qt⋆​\(ht,at⋆\)≥maxu∈Et⁡Qt⋆​\(ht,u\)−ε\.Q\_\{t\}^\{\\star\}\(h\_\{t\},a\_\{t\}^\{\\star\}\)\\geq\\max\_\{u\\in E\_\{t\}\}Q\_\{t\}^\{\\star\}\(h\_\{t\},u\)\-\\varepsilon\.Then with the usual interpretation that the sum equalsHHwhenγ=1\\gamma=1, the induced auction policyπauc\\pi^\{\\mathrm\{auc\}\}satisfies

Jout​\(πauc\)≥supπJout​\(π\)−ε​1−γH1−γ,J^\{\\mathrm\{out\}\}\(\\pi^\{\\mathrm\{auc\}\}\)\\geq\\sup\_\{\\pi\}J^\{\\mathrm\{out\}\}\(\\pi\)\-\\varepsilon\\,\\frac\{1\-\\gamma^\{H\}\}\{1\-\\gamma\},

The theorem separates two questions\. Process rewards may help train or shape the population, but they are not required for correctness once the market reliably selects near\-best specialists under the outcome objective\.

### C\.3Regret to an oracle coordinator

We next compare the decentralized auction to a centralized oracle that directly selects the best eligible specialist at every history\. The result shows that if market bids approximate the oracle action\-values, then the auction policy has vanishing regret relative to this coordinator\.

###### Theorem 3\(Vanishing regret of the market auction policy\)\.

Suppose there exists a nonincreasing sequence\{βe\}e≥1\\\{\\beta\_\{e\}\\\}\_\{e\\geq 1\}such that for every episodeee, reachable historyhth\_\{t\}, and eligible agentu∈Ee,t​\(ht\)u\\in E\_\{e,t\}\(h\_\{t\}\),

\|be,t​\(ht,u\)−Qe,t⋆​\(ht,u\)\|≤βe\.\\left\|b\_\{e,t\}\(h\_\{t\},u\)\-Q^\{\\star\}\_\{e,t\}\(h\_\{t\},u\)\\right\|\\leq\\beta\_\{e\}\.Letrer\_\{e\}be the episode regret of the market policy relative to the oracle coordinator\. Thenre≤2​βe​∑t=0H−1γt\.r\_\{e\}\\leq 2\\beta\_\{e\}\\sum\_\{t=0\}^\{H\-1\}\\gamma^\{t\}\.Consequently,Reg​\(E\):=∑e=1Ere≤2​\(∑t=0H−1γt\)​∑e=1Eβe\.\\mathrm\{Reg\}\(E\):=\\sum\_\{e=1\}^\{E\}r\_\{e\}\\leq 2\\Bigl\(\\sum\_\{t=0\}^\{H\-1\}\\gamma^\{t\}\\Bigr\)\\sum\_\{e=1\}^\{E\}\\beta\_\{e\}\.In particular, ifβe≤B​e−1/2\\beta\_\{e\}\\leq Be^\{\-1/2\}, then

Reg​\(E\)E=O​\(E−1/2\)\.\\frac\{\\mathrm\{Reg\}\(E\)\}\{E\}=O\(E^\{\-1/2\}\)\.

This bound makes explicit that the market need not be globally centralized: as long as bids become calibrated to specialist value, the auction tracks the oracle coordinator\.

### C\.4Bucket\-brigade payments recover structured credit assignment

Bucket\-brigade payments also provide a mechanism for assigning credit to earlier agents\. The next two results show that, under additional structure, the payment received from the next winner can be interpreted as either a Bellman continuation value or an ordered marginal contribution in a workflow\.

###### Proposition 1\(Bellman\-like credit under continuation pricing\)\.

If the expected payment received from the next winner equals the discounted continuation value,

𝔼​\[Yt∣ht,at⋆\]=γ​𝔼​\[Vbb​\(ht\+1\)∣ht,at⋆\],\\mathbb\{E\}\[Y\_\{t\}\\mid h\_\{t\},a\_\{t\}^\{\\star\}\]=\\gamma\\mathbb\{E\}\[V^\{\\mathrm\{bb\}\}\(h\_\{t\+1\}\)\\mid h\_\{t\},a\_\{t\}^\{\\star\}\],then the bucket\-brigade target satisfies

𝔼​\[rt\+Yt∣ht,at⋆\]=𝔼​\[rt\+γ​Vbb​\(ht\+1\)∣ht,at⋆\]\.\\mathbb\{E\}\[r\_\{t\}\+Y\_\{t\}\\mid h\_\{t\},a\_\{t\}^\{\\star\}\]=\\mathbb\{E\}\[r\_\{t\}\+\\gamma V^\{\\mathrm\{bb\}\}\(h\_\{t\+1\}\)\\mid h\_\{t\},a\_\{t\}^\{\\star\}\]\.Thus the winner’s wealth update uses a Bellman\-like one\-step target\.

###### Theorem 4\(Shapley\-like credit in DAG workflows\)\.

Consider an acyclic workflow in which subtasks are executed in topological order, and suppose the next winner’s bid prices the continuation value of the remaining feasible subgraph\. Then the bucket\-brigade payment to the current subtask agent coincides, up to a downstream baseline, with that agent’s ordered marginal contribution to the final outcome value\.

Hence bucket\-brigade transfers do more than redistribute wealth: they can serve as a local credit\-assignment signal that propagates downstream value backward through the executed chain of agents\.

Throughout the proofs, we writexxfor a decision context\. In a fully observed Markov environment,x=ot=stx=o\_\{t\}=s\_\{t\}; in partially observed settings,xxmay denote the history

ht=\(o0,a0env,r0,…,ot\)\.h\_\{t\}=\(o\_\{0\},a\_\{0\}^\{\\mathrm\{env\}\},r\_\{0\},\\ldots,o\_\{t\}\)\.For an agentaa, letbab\_\{a\}denote its fixed bid andWaW\_\{a\}its wealth\. Whenaawins at contextxx, define its resale\-adjusted economic payoff as

Ga​\(x\):=rt\+𝟏​\{t\+1<H\}​bat\+1⋆,G\_\{a\}\(x\):=r\_\{t\}\+\\mathbf\{1\}\\\{t\+1<H\\\}b\_\{a\_\{t\+1\}^\{\\star\}\},wherertr\_\{t\}is the immediate environment reward andbat\+1⋆b\_\{a\_\{t\+1\}^\{\\star\}\}is the next winning bid paid back toaaunder the bucket\-brigade transfer rule\. At a terminal step, the resale term is zero\. Thus the net wealth increment of a winning agent, ignoring rent, is

Ga​\(x\)−ba\.G\_\{a\}\(x\)\-b\_\{a\}\.Let

V​\(a,x\):=𝔼​\[Ga​\(x\)\],V⋆​\(x\):=supa:ϕa​\(x\)=1V​\(a,x\)\.V\(a,x\):=\\mathbb\{E\}\[G\_\{a\}\(x\)\],\\qquad V^\{\\star\}\(x\):=\\sup\_\{a:\\phi\_\{a\}\(x\)=1\}V\(a,x\)\.Let𝖲𝗎𝗋𝗏∞​\(x\)\\mathsf\{Surv\}\_\{\\infty\}\(x\)denote the set of agents that win at contextxxinfinitely often and are never removed, and define the limiting surviving bid frontier

β∞​\(x\):=supa∈𝖲𝗎𝗋𝗏∞​\(x\)ba,\\beta\_\{\\infty\}\(x\):=\\sup\_\{a\\in\\mathsf\{Surv\}\_\{\\infty\}\(x\)\}b\_\{a\},withsup∅:=0\\sup\\varnothing:=0\.

### C\.5Malicious agents and collusion

Finally, the bankruptcy rule discourages agents that bid aggressively but do not create commensurate value\. A malicious agent whose bid exceeds its expected resale\-adjusted payoff has negative wealth drift and is removed in finite time\.

For colluding agents, the relevant object is not individual wealth but coalition wealth, because internal transfers cancel\. A cartel can survive only if some surviving sub\-coalition has nonnegative aggregate external wealth drift\.

Thus collusion changes the unit of selection from an individual agent to a surviving cartel\. Bankruptcy eliminates colluding agents only when every possible surviving cartel is externally loss\-making; otherwise a cartel with nonnegative aggregate drift may persist or even form a monopoly\.

### C\.6Proofs for[Section˜C\.1](https://arxiv.org/html/2606.02859#A3.SS1)

###### Lemma 1\(Survival of underpriced agents\)\.

Suppose contextxxis recurrent and the payoff sequence of agentaaatxxis stationary and bounded\. If

V​\(a,x\)−ba≥δ\>0,V\(a,x\)\-b\_\{a\}\\geq\\delta\>0,then agentaasurvives forever after its first win atxxwith positive probability\.

###### Proof of Lemma[1](https://arxiv.org/html/2606.02859#Thmlemma1)\.

Fix an agentaaand contextxx\. LetGa,k​\(x\)G\_\{a,k\}\(x\)be the resale\-adjusted payoff obtained on thekk\-th win of agentaaatxx\. Define

Xk:=ba−Ga,k​\(x\)\.X\_\{k\}:=b\_\{a\}\-G\_\{a,k\}\(x\)\.By stationarity and boundedness,\(Xk\)k≥1\(X\_\{k\}\)\_\{k\\geq 1\}is i\.i\.d\. and bounded\. Moreover,

𝔼​\[Xk\]=ba−V​\(a,x\)≤−δ<0\.\\mathbb\{E\}\[X\_\{k\}\]=b\_\{a\}\-V\(a,x\)\\leq\-\\delta<0\.SinceXkX\_\{k\}is bounded and has strictly negative mean, there existsλ\>0\\lambda\>0small enough such that

mλ:=𝔼​\[eλ​Xk\]<1\.m\_\{\\lambda\}:=\\mathbb\{E\}\\\!\\left\[e^\{\\lambda X\_\{k\}\}\\right\]<1\.Then

Mn:=exp⁡\(λ​∑k=1nXk\)M\_\{n\}:=\\exp\\\!\\left\(\\lambda\\sum\_\{k=1\}^\{n\}X\_\{k\}\\right\)is a nonnegative supermartingale\.

LetTTbe the bankruptcy time measured in wins ofaaat contextxx:

T:=inf\{n≥1:∑k=1nXk\>Wa​\(0\)\},T:=\\inf\\left\\\{n\\geq 1:\\sum\_\{k=1\}^\{n\}X\_\{k\}\>W\_\{a\}\(0\)\\right\\\},whereWa​\(0\)W\_\{a\}\(0\)is the wealth immediately before its first win atxx\. On\{T≤n\}\\\{T\\leq n\\\},

MT≥eλ​Wa​\(0\)\.M\_\{T\}\\geq e^\{\\lambda W\_\{a\}\(0\)\}\.By optional stopping applied toT∧nT\\wedge n,

1=M0≥𝔼​\[MT∧n\]≥eλ​Wa​\(0\)​ℙ​\(T≤n\)\.1=M\_\{0\}\\geq\\mathbb\{E\}\[M\_\{T\\wedge n\}\]\\geq e^\{\\lambda W\_\{a\}\(0\)\}\\mathbb\{P\}\(T\\leq n\)\.Thus

ℙ​\(T≤n\)≤e−λ​Wa​\(0\)\.\\mathbb\{P\}\(T\\leq n\)\\leq e^\{\-\\lambda W\_\{a\}\(0\)\}\.Lettingn→∞n\\to\\inftygives

ℙ​\(T<∞\)≤e−λ​Wa​\(0\)<1\.\\mathbb\{P\}\(T<\\infty\)\\leq e^\{\-\\lambda W\_\{a\}\(0\)\}<1\.Therefore

ℙ\(T=∞\)≥1−e−λ​Wa​\(0\)=:qδ\>0\.\\mathbb\{P\}\(T=\\infty\)\\geq 1\-e^\{\-\\lambda W\_\{a\}\(0\)\}=:q\_\{\\delta\}\>0\.∎

###### Proof of Theorem[1](https://arxiv.org/html/2606.02859#Thmtheorem1)\.

We prove the upper and lower bounds separately\.

#### Upper bound\.

Fix anya∈𝖲𝗎𝗋𝗏∞​\(x\)a\\in\\mathsf\{Surv\}\_\{\\infty\}\(x\)\. After its firstnnwins atxx, its wealth is bounded above by

Wa​\(n\)≤Wa​\(0\)\+∑k=1n\(Ga,k​\(x\)−ba\),W\_\{a\}\(n\)\\leq W\_\{a\}\(0\)\+\\sum\_\{k=1\}^\{n\}\\bigl\(G\_\{a,k\}\(x\)\-b\_\{a\}\\bigr\),ignoring nonpositive rent terms\. By the strong law of large numbers,

1n​∑k=1n\(Ga,k​\(x\)−ba\)→a\.s\.V​\(a,x\)−ba\.\\frac\{1\}\{n\}\\sum\_\{k=1\}^\{n\}\\bigl\(G\_\{a,k\}\(x\)\-b\_\{a\}\\bigr\)\\xrightarrow\{\\mathrm\{a\.s\.\}\}V\(a,x\)\-b\_\{a\}\.Ifba\>V​\(a,x\)b\_\{a\}\>V\(a,x\), thenWa​\(n\)→−∞W\_\{a\}\(n\)\\to\-\\inftyalmost surely, soaawould eventually be removed\. This contradictsa∈𝖲𝗎𝗋𝗏∞​\(x\)a\\in\\mathsf\{Surv\}\_\{\\infty\}\(x\)\. Hence every forever\-surviving recurrent winner satisfies

ba≤V​\(a,x\)≤V⋆​\(x\)\.b\_\{a\}\\leq V\(a,x\)\\leq V^\{\\star\}\(x\)\.Taking the supremum gives

β∞​\(x\)≤V⋆​\(x\)\.\\beta\_\{\\infty\}\(x\)\\leq V^\{\\star\}\(x\)\.

#### Lower bound\.

Fixδ\>0\\delta\>0\. Suppose, toward contradiction, that with positive probability

β∞​\(x\)<V⋆​\(x\)−εmax−2​δ\.\\beta\_\{\\infty\}\(x\)<V^\{\\star\}\(x\)\-\\varepsilon\_\{\\max\}\-2\\delta\.On this event, the recurrent solvent frontier atxxis eventually belowV⋆​\(x\)−εmax−2​δV^\{\\star\}\(x\)\-\\varepsilon\_\{\\max\}\-2\\delta\. Sincexxis recurrent, there are infinitely many later opportunities for injection\. By the entry condition, with probability at leastλδ\>0\\lambda\_\{\\delta\}\>0, a new eligible agenta′a^\{\\prime\}is injected with

V​\(a′,x\)≥V⋆​\(x\)−δ\.V\(a^\{\\prime\},x\)\\geq V^\{\\star\}\(x\)\-\\delta\.Under the novice rule,

ba′=maxa∈Ct⁡ba\+εa′,εa′∈\(0,εmax\],b\_\{a^\{\\prime\}\}=\\max\_\{a\\in C\_\{t\}\}b\_\{a\}\+\\varepsilon\_\{a^\{\\prime\}\},\\qquad\\varepsilon\_\{a^\{\\prime\}\}\\in\(0,\\varepsilon\_\{\\max\}\],whereCtC\_\{t\}is the competing eligible set whena′a^\{\\prime\}first competes\. Since the competing frontier is belowV⋆​\(x\)−εmax−2​δV^\{\\star\}\(x\)\-\\varepsilon\_\{\\max\}\-2\\delta, we have

ba′≤V⋆​\(x\)−2​δ\.b\_\{a^\{\\prime\}\}\\leq V^\{\\star\}\(x\)\-2\\delta\.Therefore

V​\(a′,x\)−ba′≥δ\.V\(a^\{\\prime\},x\)\-b\_\{a^\{\\prime\}\}\\geq\\delta\.By Lemma[1](https://arxiv.org/html/2606.02859#Thmlemma1), each such entrant survives forever after its first win atxxwith probability at leastqδ\>0q\_\{\\delta\}\>0\. Across infinitely many independent injection opportunities, the probability that no such profitable entrant survives is zero\. Hence, almost surely, one such entrant eventually survives, contradicting the assumed upper bound onβ∞​\(x\)\\beta\_\{\\infty\}\(x\)\. Therefore

β∞​\(x\)≥V⋆​\(x\)−εmax−2​δ\.\\beta\_\{\\infty\}\(x\)\\geq V^\{\\star\}\(x\)\-\\varepsilon\_\{\\max\}\-2\\delta\.Lettingδ↓0\\delta\\downarrow 0gives

β∞​\(x\)≥V⋆​\(x\)−εmax\.\\beta\_\{\\infty\}\(x\)\\geq V^\{\\star\}\(x\)\-\\varepsilon\_\{\\max\}\.Combining the two bounds proves the theorem\. ∎

### C\.7Proofs for[Section˜C\.2](https://arxiv.org/html/2606.02859#A3.SS2)

###### Proof of Theorem[2](https://arxiv.org/html/2606.02859#Thmtheorem2)\.

Let

Vtauc\(ht\):=𝔼πauc\[∑u=tH−1γu−truout\|ht\],V\_\{t\}^\{\\mathrm\{auc\}\}\(h\_\{t\}\):=\\mathbb\{E\}\_\{\\pi^\{\\mathrm\{auc\}\}\}\\\!\\left\[\\sum\_\{u=t\}^\{H\-1\}\\gamma^\{u\-t\}r\_\{u\}^\{\\mathrm\{out\}\}\\,\\middle\|\\,h\_\{t\}\\right\],and define

Δt​\(ht\):=Vt⋆​\(ht\)−Vtauc​\(ht\)\.\\Delta\_\{t\}\(h\_\{t\}\):=V\_\{t\}^\{\\star\}\(h\_\{t\}\)\-V\_\{t\}^\{\\mathrm\{auc\}\}\(h\_\{t\}\)\.At any reachable history withEt≠∅E\_\{t\}\\neq\\varnothing, the assumption gives

Qt⋆​\(ht,at⋆\)≥maxu∈Et⁡Qt⋆​\(ht,u\)−ε\.Q\_\{t\}^\{\\star\}\(h\_\{t\},a\_\{t\}^\{\\star\}\)\\geq\\max\_\{u\\in E\_\{t\}\}Q\_\{t\}^\{\\star\}\(h\_\{t\},u\)\-\\varepsilon\.After the auction winner acts, the remaining loss is the future auction suboptimality\. Therefore

Δt\(ht\)≤ε\+γ𝔼\[Δt\+1\(ht\+1\)\|ht,at⋆\]\.\\Delta\_\{t\}\(h\_\{t\}\)\\leq\\varepsilon\+\\gamma\\mathbb\{E\}\\\!\\left\[\\Delta\_\{t\+1\}\(h\_\{t\+1\}\)\\,\\middle\|\\,h\_\{t\},a\_\{t\}^\{\\star\}\\right\]\.SinceΔH​\(hH\)=0\\Delta\_\{H\}\(h\_\{H\}\)=0, backward induction yields

Δ0​\(h0\)≤ε​∑t=0H−1γt\.\\Delta\_\{0\}\(h\_\{0\}\)\\leq\\varepsilon\\sum\_\{t=0\}^\{H\-1\}\\gamma^\{t\}\.Taking expectation over the initial history proves the theorem\. ∎

### C\.8Proofs for[Section˜C\.3](https://arxiv.org/html/2606.02859#A3.SS3)

###### Proof of Theorem[3](https://arxiv.org/html/2606.02859#Thmtheorem3)\.

Fix episodeeeand reachable historyhth\_\{t\}\. Letue,torcu^\{\\mathrm\{orc\}\}\_\{e,t\}be the oracle\-selected eligible agent, and letue,taucu^\{\\mathrm\{auc\}\}\_\{e,t\}be the auction winner\. By bid concentration,

Qe,t⋆​\(ht,ue,torc\)≤be,t​\(ht,ue,torc\)\+βe\.Q^\{\\star\}\_\{e,t\}\(h\_\{t\},u^\{\\mathrm\{orc\}\}\_\{e,t\}\)\\leq b\_\{e,t\}\(h\_\{t\},u^\{\\mathrm\{orc\}\}\_\{e,t\}\)\+\\beta\_\{e\}\.Because the auction maximizes bids,

be,t​\(ht,ue,torc\)≤be,t​\(ht,ue,tauc\)\.b\_\{e,t\}\(h\_\{t\},u^\{\\mathrm\{orc\}\}\_\{e,t\}\)\\leq b\_\{e,t\}\(h\_\{t\},u^\{\\mathrm\{auc\}\}\_\{e,t\}\)\.Applying bid concentration again,

be,t​\(ht,ue,tauc\)≤Qe,t⋆​\(ht,ue,tauc\)\+βe\.b\_\{e,t\}\(h\_\{t\},u^\{\\mathrm\{auc\}\}\_\{e,t\}\)\\leq Q^\{\\star\}\_\{e,t\}\(h\_\{t\},u^\{\\mathrm\{auc\}\}\_\{e,t\}\)\+\\beta\_\{e\}\.Combining the three inequalities gives

Qe,t⋆​\(ht,ue,tauc\)≥maxu∈Ee,t​\(ht\)⁡Qe,t⋆​\(ht,u\)−2​βe\.Q^\{\\star\}\_\{e,t\}\(h\_\{t\},u^\{\\mathrm\{auc\}\}\_\{e,t\}\)\\geq\\max\_\{u\\in E\_\{e,t\}\(h\_\{t\}\)\}Q^\{\\star\}\_\{e,t\}\(h\_\{t\},u\)\-2\\beta\_\{e\}\.Thus, in episodeee, the auction winner is2​βe2\\beta\_\{e\}\-optimal at every reachable history\. Applying[Theorem˜2](https://arxiv.org/html/2606.02859#Thmtheorem2)withε=2​βe\\varepsilon=2\\beta\_\{e\}gives

re≤2​βe​∑t=0H−1γt\.r\_\{e\}\\leq 2\\beta\_\{e\}\\sum\_\{t=0\}^\{H\-1\}\\gamma^\{t\}\.Summing overe=1,…,Ee=1,\\ldots,Eproves the cumulative regret bound\. Ifβe≤B​e−1/2\\beta\_\{e\}\\leq Be^\{\-1/2\}, then

∑e=1Eβe≤B​∑e=1Ee−1/2≤2​B​E,\\sum\_\{e=1\}^\{E\}\\beta\_\{e\}\\leq B\\sum\_\{e=1\}^\{E\}e^\{\-1/2\}\\leq 2B\\sqrt\{E\},which gives

Reg​\(E\)≤4​B​\(∑t=0H−1γt\)​E\.\\mathrm\{Reg\}\(E\)\\leq 4B\\Bigl\(\\sum\_\{t=0\}^\{H\-1\}\\gamma^\{t\}\\Bigr\)\\sqrt\{E\}\.Dividing byEEyields the average\-regret rate\. ∎

### C\.9Proofs for[Section˜C\.4](https://arxiv.org/html/2606.02859#A3.SS4)

###### Proof of Proposition[1](https://arxiv.org/html/2606.02859#Thmproposition1)\.

By assumption,

𝔼​\[Yt∣ht,at⋆\]=γ​𝔼​\[Vbb​\(ht\+1\)∣ht,at⋆\]\.\\mathbb\{E\}\[Y\_\{t\}\\mid h\_\{t\},a\_\{t\}^\{\\star\}\]=\\gamma\\mathbb\{E\}\[V^\{\\mathrm\{bb\}\}\(h\_\{t\+1\}\)\\mid h\_\{t\},a\_\{t\}^\{\\star\}\]\.Substituting this identity into

Δ​Wat⋆,twin=rt\+Yt−bat⋆\\Delta W^\{\\mathrm\{win\}\}\_\{a\_\{t\}^\{\\star\},t\}=r\_\{t\}\+Y\_\{t\}\-b\_\{a\_\{t\}^\{\\star\}\}and taking conditional expectations gives

𝔼​\[Δ​Wat⋆,twin∣ht\]=𝔼​\[rt\+γ​Vbb​\(ht\+1\)∣ht,at⋆\]−bat⋆\.\\mathbb\{E\}\[\\Delta W^\{\\mathrm\{win\}\}\_\{a\_\{t\}^\{\\star\},t\}\\mid h\_\{t\}\]=\\mathbb\{E\}\[r\_\{t\}\+\\gamma V^\{\\mathrm\{bb\}\}\(h\_\{t\+1\}\)\\mid h\_\{t\},a\_\{t\}^\{\\star\}\]\-b\_\{a\_\{t\}^\{\\star\}\}\.∎

###### Proof sketch of Theorem[4](https://arxiv.org/html/2606.02859#Thmtheorem4)\.

Because the workflow is acyclic, any execution induces a topological order\. Completing subtaskuuenlarges the completed prefix fromSStoS∪\{u\}S\\cup\\\{u\\\}\. Under the continuation\-pricing assumption, the successor’s willingness to pay equals the continuation value of the expanded feasible subgraph, up to a downstream baseline\. Therefore the payment flowing back to the current agent measures the additional downstream value unlocked by executinguu\. This is precisely the ordered marginal contribution

v​\(S∪\{u\}\)−v​\(S\),v\(S\\cup\\\{u\\\}\)\-v\(S\),up to the baseline term\. Summing over the realized topological order yields a telescoping decomposition of the final outcome value\. ∎

### C\.10Proofs for[Section˜C\.5](https://arxiv.org/html/2606.02859#A3.SS5)

###### Lemma 2\(Extinction of malicious agents\)\.

Suppose payoffs are recurrent, stationary, and bounded, and suppose rentρ\>0\\rho\>0is charged infinitely often\. If an agentaais malicious in the sense that there existsη\>0\\eta\>0such that

V​\(a,x\)≤ba−ηV\(a,x\)\\leq b\_\{a\}\-\\etaat every recurrent contextxxwhereaais eligible, thenaais removed in finite time almost surely\.

###### Proof of Lemma[2](https://arxiv.org/html/2606.02859#Thmlemma2)\.

Fix an agentaa\. Suppose first thataawins infinitely often\. LetGa,kG\_\{a,k\}be the resale\-adjusted payoff on itskk\-th win\. Its wealth afternnwins satisfies

Wa​\(n\)=Wa​\(0\)\+∑k=1n\(Ga,k−ba\)−ρ​Naep​\(n\),W\_\{a\}\(n\)=W\_\{a\}\(0\)\+\\sum\_\{k=1\}^\{n\}\\bigl\(G\_\{a,k\}\-b\_\{a\}\\bigr\)\-\\rho N\_\{a\}^\{\\mathrm\{ep\}\}\(n\),whereNaep​\(n\)N\_\{a\}^\{\\mathrm\{ep\}\}\(n\)is the number of elapsed rent\-charging episodes by the time of itsnn\-th win\. For a malicious agent,

𝔼​\[Ga,k−ba\]≤−η\.\\mathbb\{E\}\[G\_\{a,k\}\-b\_\{a\}\]\\leq\-\\eta\.By the strong law of large numbers,

1n​∑k=1n\(Ga,k−ba\)→𝔼​\[Ga,1−ba\]≤−ηa\.s\.\\frac\{1\}\{n\}\\sum\_\{k=1\}^\{n\}\\bigl\(G\_\{a,k\}\-b\_\{a\}\\bigr\)\\to\\mathbb\{E\}\[G\_\{a,1\}\-b\_\{a\}\]\\leq\-\\eta\\qquad\\text\{a\.s\.\}The rent term is nonpositive, soWa​\(n\)→−∞W\_\{a\}\(n\)\\to\-\\inftyalmost surely\. Thus bankruptcy occurs in finite time\.

Ifaawins only finitely many times, then after its final win it no longer receives downstream payments\. Sinceρ\>0\\rho\>0and rent is charged infinitely often, rent alone eventually drivesWaW\_\{a\}below zero\. Henceaais removed in finite time almost surely\. ∎

Before proving the collusion result, we record the coalition wealth decomposition\. For a nonempty set of agentsSS, define

WS​\(t\):=∑a∈S∩𝒜tWa​\(t\)\.W\_\{S\}\(t\):=\\sum\_\{a\\in S\\cap\\mathcal\{A\}\_\{t\}\}W\_\{a\}\(t\)\.Bucket\-brigade transfers between two agents inSScancel inWS​\(t\)W\_\{S\}\(t\)\. Therefore, over an episode segmentτ\\tau, the only changes toWSW\_\{S\}come from environment rewards, payments enteringSS, and payments leavingSS\. Ifptp\_\{t\}denotes the predecessor of the winner at steptt, withpt=∅p\_\{t\}=\\varnothingat the first winning step of an episode, then

Δ​WS​\(τ\)=\\displaystyle\\Delta W\_\{S\}\(\\tau\)=\{\}∑t∈τ𝟏​\{at⋆∈S\}​rt\\displaystyle\\sum\_\{t\\in\\tau\}\\mathbf\{1\}\\\{a\_\{t\}^\{\\star\}\\in S\\\}\\,r\_\{t\}\+∑t∈τ𝟏​\{at⋆∉S,pt∈S\}​bat⋆\\displaystyle\+\\sum\_\{t\\in\\tau\}\\mathbf\{1\}\\\{a\_\{t\}^\{\\star\}\\notin S,\\ p\_\{t\}\\in S\\\}\\,b\_\{a\_\{t\}^\{\\star\}\}−∑t∈τ𝟏​\{at⋆∈S,pt∉S∪\{∅\}\}​bat⋆\\displaystyle\-\\sum\_\{t\\in\\tau\}\\mathbf\{1\}\\\{a\_\{t\}^\{\\star\}\\in S,\\ p\_\{t\}\\notin S\\cup\\\{\\varnothing\\\}\\\}\\,b\_\{a\_\{t\}^\{\\star\}\}−∑t∈τ𝟏​\{at⋆∈S,pt=∅\}​bat⋆\.\\displaystyle\-\\sum\_\{t\\in\\tau\}\\mathbf\{1\}\\\{a\_\{t\}^\{\\star\}\\in S,\\ p\_\{t\}=\\varnothing\\\}\\,b\_\{a\_\{t\}^\{\\star\}\}\.
###### Lemma 3\(Extinction of externally loss\-making collusions\)\.

LetSSbe a coalition of agents that may coordinate bids, prompts, and internal transfers\. If every nonempty surviving sub\-coalitionT⊆ST\\subseteq Shas strictly negative expected aggregate external wealth drift on recurrent episode segments, then all agents inSSare removed in finite time almost surely\.

###### Proof of Lemma[3](https://arxiv.org/html/2606.02859#Thmlemma3)\.

Consider the currently surviving colluding setT⊆ST\\subseteq S\. By assumption, wheneverTTis active on recurrent episode segments,

𝔼​\[Δ​WT​\(τ\)∣ℱτ−\]≤−ηT\\mathbb\{E\}\[\\Delta W\_\{T\}\(\\tau\)\\mid\\mathcal\{F\}\_\{\\tau^\{\-\}\}\]\\leq\-\\eta\_\{T\}for someηT\>0\\eta\_\{T\}\>0\. The increments are bounded, so the coalition wealthWTW\_\{T\}has strictly negative drift on recurrent segments\. By the strong law of large numbers for the corresponding bounded adapted increments,WTW\_\{T\}eventually falls below zero unless some member ofTTis removed earlier\. Once a member is removed, the argument is restarted with the smaller surviving sub\-coalition\. The assumption holds for every nonempty sub\-coalition that can remain active after bankruptcies, so this induction continues until no member ofSSsurvives\. Hence all agents inSSare removed in finite time almost surely\. ∎

## Appendix DAdditional Experiment Details

This appendix provides task\-level and baseline\-level details omitted from the main experiment section\. In the main text, we use the term*partial agent*for any agent whose capability is incomplete relative to the full task interface\. Partiality can arise from explicit constraints, such as limited tools or output length, or from specialization, such as role\-specific prompting\. This section specifies how partiality is instantiated in each domain\.

We compareEoMwith complete\-agent, partial\-agent, and domain\-specific baselines\. The experiments were conducted primarily for inference\-time evaluation\. We use NVIDIA H200 GPUs for running local models, with at most a single H200 GPU required per experiment\. In addition to local computation, we leveraged several commercial APIs, including the official APIs from OpenAI, Google Gemini, and Anthropic Claude, to evaluate model performance across different providers\.

### D\.1Baseline Details

#### Complete\-agent baselines\.

ReAct\[[50](https://arxiv.org/html/2606.02859#bib.bib50)\]is the primary complete\-agent baseline\. In each domain where it is used, a single agent with the same backbone model is allowed to solve the task end\-to\-end with access to the full action or tool space\. This baseline tests whether economic organization among partial agents can compensate for capability that is concentrated in a single complete agent\.

GEA\[[45](https://arxiv.org/html/2606.02859#bib.bib45)\]is also instantiated as a complete\-agent baseline\. LikeReAct, it uses complete agents with access to the full task interface, but additionally provides a self\-improvement mechanism based on experience sharing and agent evolution\. This baseline tests whether the gains ofEoMcome specifically from market\-based execution\-time coordination and wealth\-driven population evolution, rather than from self\-improvement alone\.

OpenEvolve\[[35](https://arxiv.org/html/2606.02859#bib.bib35),[27](https://arxiv.org/html/2606.02859#bib.bib27)\]is used for distributed\-system optimization\. It performs iterative program improvement with a monolithic evolutionary coding agent\. In Cloudcast,OpenEvolvefollows its standard setup with 300 iterations, whileEoMruns for 30 episodes\.

#### Partial\-agent baselines\.

For partial\-agent comparison, we use Multi\-Agent Debate\[[11](https://arxiv.org/html/2606.02859#bib.bib11)\]\. This baseline uses multiple agents that interact with one another, but it does not include auctions, bid\-based control allocation, peer\-to\-peer payments, rent, bankruptcy, or wealth\-based exploration and exploitation\. It isolates the effect of economic coordination from the effect of merely using multiple partial agents\.

We also compare against a best\-of\-NNmulti\-agent baseline on Cloudcast, whereNNis set to the number of episodes used byEoM\. This baseline tests whether improvements can be explained by repeated multi\-agent sampling without market\-driven evolution\.

#### Domain\-specific baselines\.

For accelerator design, we compare againstDOSA\[[18](https://arxiv.org/html/2606.02859#bib.bib18)\], a strong non\-LLM method for differentiable model\-based one\-loop search in DNN accelerator optimization\.

### D\.2Mathematical Reasoning

We evaluate on MATH\[[16](https://arxiv.org/html/2606.02859#bib.bib16)\], a competition\-level mathematics benchmark whose problems are annotated from Level 1 to Level 5 by difficulty\. We randomly sample 20 training problems from each difficulty level and train on an easy\-to\-hard stream from Level 1 to Level 5\. Evaluation reports greedy pass@1 accuracy separately by difficulty level\.

The population is initialized with a planner–executor–verifier decomposition\. The planner proposes the immediate next step, the executor carries out the proposed step, and the verifier judges whether the executed step is correct\. These agents are partial in two ways\. First, each agent is role\-specific and therefore sees the task through a limited functional interface\. Second, each agent is given a short maximum output budget of 64–256 tokens, preventing any individual agent from solving the entire problem with a long chain\-of\-thought style response\.

We intentionally use relatively weak LLM backbones, Llama\-3\.1\-8B and Gemma\-2\-9B, to test whether economic organization can amplify partial individual capabilities\. The complete\-agent baseline is a single end\-to\-end agent with the corresponding backbone and full output/action capability\.

### D\.3Financial Research

We evaluate on Finance\-Agent\-Bench\[[5](https://arxiv.org/html/2606.02859#bib.bib5)\], a benchmark of real\-world financial research tasks over company filings\. The environment provides four tools\. We randomly select 30 tasks for training and use the remaining 20 for testing\.

Each partial agent is restricted to exactly one tool, so no individual partial agent can complete the full research task alone\. The population therefore must coordinate across tool\-specialized agents, including agents responsible for filing search, web search, HTML parsing, retrieval, and answer submission\. We compare against Multi\-Agent Debate under the same partial\-agent setting, as well as complete\-agentReActandGEAbaselines that can access all tools\.

For the generalist\-v\.s\.\-specialist study in[Figure˜7](https://arxiv.org/html/2606.02859#S3.F7), we introduce a complete generalist agent with access to all four tools, alongside the ordinary partial specialist agents\. Appendix[G](https://arxiv.org/html/2606.02859#A7)analyzes why this generalist does not monopolize the market\.

### D\.4Scientific Research

We evaluate on the*Research*subset of FrontierScience\[[43](https://arxiv.org/html/2606.02859#bib.bib43)\], which assesses open\-ended scientific problem\-solving\. We adopt a 40/20 train–test split\.

The society is initialized with four role\-specialized partial agents: literature, planner, executor, and verifier\. The literature agent lists possible background knowledge required for the scientific question\. The planner decomposes the problem into subgoals\. The executor advances one unresolved sub\-question with algebraic or scientific reasoning\. The verifier checks whether intermediate reasoning is valid and flags errors\. These agents are partial because each is prompted to operate only within one stage of the problem\-solving pipeline\.

We use Gemini\-3\-Flash as the backbone model and compare againstGEAunder the same backbone\. We report both mean accuracy and best\-run accuracy across evaluated checkpoints during the evolution process\. Appendix[E\.1](https://arxiv.org/html/2606.02859#A5.SS1)and Appendix[E\.3](https://arxiv.org/html/2606.02859#A5.SS3)provide case studies showing how scientific reasoning prompts and auction\-selected topologies evolve over training\.

### D\.5Accelerator Design

We evaluate EOM on accelerator design, following the task description in Section 3\.1\. The task is a discrete search problem over hardware mappings for deep learning accelerators\. We use the standardized GEMMINI benchmark suite\[[13](https://arxiv.org/html/2606.02859#bib.bib13)\], which provides a hardware simulation environment for 24 ResNet\-50 convolution kernels\. For each kernel, the goal is to find a hardware mapping that minimizes energy\-delay product \(EDP\), with lower EDP indicating a better design\. Broader context on generative models applied across the systems stack, including DSE, can be found in\[[31](https://arxiv.org/html/2606.02859#bib.bib31),[40](https://arxiv.org/html/2606.02859#bib.bib40),[38](https://arxiv.org/html/2606.02859#bib.bib38)\]\.

The population contains three role\-specialized partial agent types\.*Historians*summarize previous trajectory trials and maintain memory of promising or failed design directions\.*Planners*propose architectural or mapping\-level search directions\.*Executors*carry out fine\-grained local evaluations\. These roles are partial because each agent is responsible for only one component of the long\-horizon design\-space exploration loop\.

We compare against the non\-LLMDOSAbaseline\[[18](https://arxiv.org/html/2606.02859#bib.bib18)\]and a completeReActagent\. Both LLM\-based methods use the Gemma\-4\-31B backbone\. The main metric is average EDP across the 24 kernels, with lower values indicating better designs\.[Figure˜5](https://arxiv.org/html/2606.02859#S3.F5)analyzes wealth dynamics during search, and[Figure˜7](https://arxiv.org/html/2606.02859#S3.F7)reports per\-kernel EDP improvements\.

Table 3:Representative discovered mappings\.Example accelerator mappings discovered byEoMon representative kernels, with EDP reduction reported*relative toDOSA*\. The learned solutions repeatedly recover effective output\-stationary motifs\.Notation:L0–L3index memory\-hierarchy levels \(L0=innermost registers,L3=outermost DRAM\);K,Care output\- and input\-channel loops;P,Qare output spatial dimensions;R,Sare filter spatial dimensions;Nis the batch loop; suffixXmarks a spatially parallelized loop;\[O\]marks the output\-stationary level\.
### D\.6Distributed\-System Optimization

We evaluate on the Cloudcast task from the ADRS benchmark\[[8](https://arxiv.org/html/2606.02859#bib.bib8)\], which minimizes total data\-transfer cost in a multi\-region, multi\-cloud environment\. We formulate Cloudcast as iterative code optimization\. Each episode starts from the last successfully verified program, regardless of its score, and receives a reward at the end\. This resembles test\-time reinforcement learning\[[44](https://arxiv.org/html/2606.02859#bib.bib44),[52](https://arxiv.org/html/2606.02859#bib.bib52)\], where improvement attempts are made from a parent solution and updated using reward signals\.

The population uses role\-specialized partial agents commonly used in coding tasks\. The Planner sets subgoals; the Reader inspects the codebase and summarizes findings; the Implementer edits code; the Builder compiles and tests; the Evaluator calls the verifier and records score changes; and the Finalizer submits the run\. These agents are partial because each is responsible for only one stage of the coding and verification workflow\.

We compareEoMwithOpenEvolve\[[35](https://arxiv.org/html/2606.02859#bib.bib35),[27](https://arxiv.org/html/2606.02859#bib.bib27)\]and a best\-of\-NNmulti\-agent baseline\. All methods use GPT\-5\-mini\.EoMruns for 30 episodes, whileOpenEvolveuses its standard 300\-iteration setup, which takes longer in wall\-clock time\. Appendix[F](https://arxiv.org/html/2606.02859#A6)traces how prompt\-level action discipline and auction\-selected topology evolve over a representative Cloudcast run\.

### D\.7Evaluation Protocol

During training,EoMalternates between within\-episode planning and across\-episode population adaptation\. Within an episode, agents wake up according to their local predicates, compete via auctions, execute actions, and exchange payments through the bucket\-brigade transaction rule\. Across episodes, agents pay rent, agents with negative wealth are removed, and new agents are injected through exploration and exploitation\.

During evaluation, the trained population is frozen\. Bids, prompts, wealth values, and agent status are fixed; no payments, rewards, rent, births, or bankruptcies are applied\. Each test task is evaluated using a thread\-local copy of the trained population, and the auction mechanism is used only for action selection\. Thus, reported test results measure the learned society at a fixed checkpoint rather than continued adaptation during testing\.

## Appendix EPrompt and Topology Evolution in Scientific Research

This appendix analyzes the scientific research run at two coupled levels\. At the micro level, economic selection evolves reusable reasoning strategies inside agents’ prompts\. At the macro level, these local strategies reshape the auction\-selected collaboration topology\. The central observation is that topology evolution is not independent of prompt evolution: once anExecuterinternalizes checks that previously required separateVerifierorLiteratureintervention, the wakeup landscape changes, and the same auction rules select a different execution path\.

We focus on FrontierScience\-Research, where theExecuteras an example, advances one unresolved sub\-question with explicit algebra or scientific reasoning\. Across the 40\-episode training run, nine of eleven successful episodes are carried by descendants of a single evolvingExecuteragent\. Since each child inherits its parent’s trainable prompt and receives only a small mutation, this agent family provides a traceable view of prompt\-level evolution\. The evidence is mechanistic rather than a controlled single\-sentence intervention: we do not claim that a specific prompt line deterministically causes a specific answer\. Instead, we show that economic selection accumulates reusable, falsifiable reasoning operators, and that these operators change both local behavior and system\-level routing\.

### E\.1Prompt\-Level Evolution: From Generic Execution to Reusable Scientific Reasoning

The initial prompt is a generic execution policy\. It asks the agent to complete the earliest unfinished sub\-question, show intermediate algebra, track signs and dimensions, and repair verifier\-flagged errors\. This is a reasonable starting point, but it does not yet specify how to expose abstract relations as auditable scalar systems, how to check whether a problem is well\-posed, or how to falsify intermediate results against the original equations\.

Identifytheearliestsub\-questionintheplanthatisnotyetcompleted,

thencarryoutitsderivationcleanly:

\-Statethesub\-partlabel\(e\.g\.’Part\(b\):’\)atthetopofyour<step\>\.

\-Showtheintermediatealgebraorphysicalreasoning;keepequations

compactbutincludejustificationfornon\-trivialsubstitutions

\(limits,approximations,linearisations\)\.

\-Tracksigns,unit/dimensionconsistency,andsmall\-parameterregimes\.

\-IfapriorExecuterturnhasanerrorflaggedbyaVerifier,fixonly

thaterrorinthisturninsteadofstartinganewsub\-part\.

Table[4](https://arxiv.org/html/2606.02859#A5.T4)summarizes the main prompt\-level changes\. The prompt does not merely become longer\. It becomes more operational: each added motif specifies a concrete check or decomposition procedure that can be reused on future tasks\.

Table 4:Reusable reasoning motifs learned by the evolvingExecuteragent\. Each mutation adds a falsifiable operation rather than a task\-specific fact\.The first major mutation replaces generic algebraic exposition with coordinate\-based and component\-wise execution\. This change is distilled from a CMB trajectory in which treatingCℓT​TC\_\{\\ell\}^\{TT\},CℓE​EC\_\{\\ell\}^\{EE\}, andCℓT​EC\_\{\\ell\}^\{TE\}as separate scalar relations proved more reliable than manipulating a single high\-level covariance object\. The resulting prompt requires the agent to expose all relevant variables and scalar equations, making intermediate reasoning easier forVerifieragents to audit\.

Subsequent mutations add increasingly explicit self\-auditing\. The agent learns to state the governing principle before algebraic manipulation, test limiting cases or symmetries, verify boundary and feasibility constraints, count independent equations against unknown degrees of freedom, and, when possible, substitute the derived expression back into the governing equations\. These edits turn theExecuterfrom a generic algebraic operator into a compact scientific reasoning module\.

By generation 5, the prompt has grown from four generic bullets to eight operational checks:

Identifytheearliestsub\-questionintheplanthatisnotyetcompleted,

thencarryoutitsderivationcleanly:

\-Statethesub\-partlabel\(e\.g\.’Part\(b\):’\)andthecorelogicbridge

orgoverningprincipleconnectingtheknownvariablestothetarget\.

\-Identifyanysymmetries,invariants,orsimplifiedregimesthatcan

beusedtostreamlinethederivationorcross\-checktheresult\.

\-Verifythatthesystemiswell\-definedbycomparingthenumberof

independentequationstothedegreesoffreedombeforeproceeding

withthemath\.

\-Prioritizeexplicitcoordinate\-basedorcomponent\-wisederivations;

defineallbasiscoefficientsandvariablesclearly\.

\-Showfullalgebraicexpansionsandprovidethecompletesystemof

scalarequationsforalldegreesoffreedom\.

\-Tracksigns,unit/dimensionconsistency,andsmall\-parameterregimes;

sanity\-checkresultsvialimitingcases,symmetries,orfunctional

scaling\.

\-Confirmthefinalexpressionsatisfiesallrelevantconstraints,

boundaryconditions,andlogicalbounds;whenfeasible,verifyby

substitutingtheresultbackintothegoverningequations\.

\-IfapriorExecuterturnhasanerrorflaggedbyaVerifier,fixonly

thaterrorinthisturninsteadofstartinganewsub\-part\.

This final prompt is not simply a more verbose instruction to “execute carefully\.” It specifies a reusable scientific procedure: identify the principle, check symmetries, verify well\-posedness, expand explicitly, enforce constraints, and falsify the result by substitution\.

### E\.2Cross\-Domain Transfer of the Evolved Reasoning Routine

Table[5](https://arxiv.org/html/2606.02859#A5.T5)shows that the evolved routine transfers across domains\. The agent is shaped by physics tasks, refined on chemistry and pharmacology, and later reused in spectroscopy, biology, and synthesis\.

Table 5:Successful scientific\-research episodes associated with the evolvingExecuteragent\. The transferable object is not factual content, but a reusable reasoning routine\.The clearest transfer occurs in episode 32, the195Pt NMR spectroscopy task\. The prompt had been shaped by CMB cosmology, dark\-matter detection, electrolyte chemistry, and nAChR pharmacology, none of which directly concerns hyperfine NMR shifts\. Nevertheless, the inherited routine applies directly: identify the governing principle, count equations versus unknowns, check symmetry, expand the scalar relation, and substitute the result back into the observation\. This supports the interpretation that the economy evolves content\-independent reasoning strategies rather than memorized task solutions\.

### E\.3From Local Reasoning Evolution to Topology Evolution

The prompt\-level changes above have a system\-level consequence\. As theExecuterbegins to perform dimensional checks, equation\-vs\-DOF counting, constraint verification, and substitute\-back falsification internally, the marginal value of separateVerifierturns decreases in some states\. Similarly, once theExecuterprompt begins by naming the governing principle, separateLiteratureturns become less necessary for some tasks\. Since wakeup decisions are local LLM judgments over the current workspace and each agent’s evolved prompt, these internalized routines directly change who wakes up, who bids, and who acts\.

We define an episode’s topology as the ordered sequence of auction winners, represented at both the role level and the agent\-identity level\. This topology is not a pre\-specified execution graph\. It is reselected at every step through local wakeup decisions and bid competition\.

### E\.4Early Regime: Explicit Multi\-Role Auditing

Early in training, successful trajectories often rely on explicit multi\-role auditing\. Episode 11, the Josephson\-junction task, is the longest successful trajectory in the run\. It obtains a score of 0\.75 with ten auction steps and all five roles:

Lit\.→Plan→Exe\.→Ver\.→Exe\.→Ver\.→Plan→Exe\.→Ver\.→Ans\.\\textsc\{Lit\.\}\\rightarrow\\textsc\{Plan\}\\rightarrow\\textsc\{Exe\.\}\\rightarrow\\textsc\{Ver\.\}\\rightarrow\\textsc\{Exe\.\}\\rightarrow\\textsc\{Ver\.\}\\rightarrow\\textsc\{Plan\}\\rightarrow\\textsc\{Exe\.\}\\rightarrow\\textsc\{Ver\.\}\\rightarrow\\textsc\{Ans\.\}
This trajectory is long, but it is structured\. Literature supplies background, Planner decomposes the problem, Executer advances sub\-parts, Verifier audits intermediate executions, Planner re\-scopes after the audits, and Answer synthesizes the final response\. No agent designs this workflow globally\. It emerges from local wakeup judgments, novice bidding, same\-role blocking, and the terminal answer restriction\.

This early topology compensates for immature local reasoning\. Because theExecuterhas not yet internalized enough verification, the society externalizes uncertainty through repeated execute–verify loops\.

### E\.5Late Regime: Compact Specialist Execution

Later in training, the same auction mechanism can select a much shorter specialist workflow\. Episode 33, the protein\-purification task, receives a score of 1\.0 with only three steps:

Plan→Exe\.→Ans\.\\textsc\{Plan\}\\rightarrow\\textsc\{Exe\.\}\\rightarrow\\textsc\{Ans\.\}
This contraction is not caused by a smaller population\. The population contains 14 agents, including livingLiteratureandVerifieragents\. These agents still run their wakeup judges, but return NO\. The reason is not a change in the auction code; it is the evolved prompt state\.Executeralready performs dimensional checks, equation\-vs\-DOF counting, constraint verification, and substitute\-back falsification\. As a result, separate verifier intervention has lower marginal value for this episode\.

### E\.6Training\-Wide Topology Trend

Table[6](https://arxiv.org/html/2606.02859#A5.T6)summarizes the successful trajectories\. Later episodes often use compact 3–4 step paths despite larger populations\. Long paths remain when the task still benefits from explicit literature or verifier intervention, showing that topology is adaptive rather than simply pruned\.

Table 6:Auction\-derived topology for successful scientific\-research episodes\. Compact paths become common later in training, even though the population remains large\.The key point is that shorter paths are not a consequence of having fewer available agents\. The market often has more agents to choose from, but selects fewer active roles because competence has concentrated in load\-bearing agents\. Conversely, when residual uncertainty remains high, as in episode 36, the auction still selects longer literature and verification chains\. Thus, topology evolution is conditional and task\-sensitive\.

### E\.7Population\-Level Selection

The population graph evolves in parallel with the execution graph\. Births are not uniformly distributed across roles\. TheExecuteragent deepens through descendants because it repeatedly appears on reward\-bearing trajectories\. TheAnsweragent family also broadens as final synthesis becomes valuable once upstream reasoning becomes reliable\. In contrast,LiteratureandPlannerremain comparatively shallow in this run\.

This reflects the economic feedback loop\. Agents that create downstream value accumulate wealth, survive longer, and become more likely to seed future mutations\. Therefore, the economy selects not only who acts in the current episode, but also where future evolutionary capacity is allocated\. The central evidence of evolution is therefore not merely that later agents solve more tasks\. It is that the market converts repeated successes and failures into reusable reasoning routines, and these routines reorganize the society’s workflow without centralized orchestration\.

## Appendix FPrompt and Topology Evolution in CloudCast

This appendix provides evidence that the same economic mechanism also works in a code\-optimization environment\. In Scientific Research, the main learned object is a reusable reasoning routine\. In CloudCast, the main learned object is an action discipline: agents learn when*not*to spend an expensive action\. The topology adapts accordingly\. When the workspace is close to a new high score, the auction selects a short read–edit–evaluate–commit path; when the workspace still contains uncertain regressions or unresolved implementation choices, the auction selects longer edit–build–evaluate loops\. Thus, the economy does not simply shorten workflows\. It selects a workflow whose shape matches the current workspace state\.

CloudCast asks the society to evolve a single Python file,initial\_program\.py, implementing a multi\-cloud broadcast routing algorithm\. Given source and destination regions and a network graph with edge costs and throughputs, the program must produce aBroadCastTopology\. The simulator scores the program by total egress cost across five inter\- and intra\-cloud scenarios\. The seed single\-path Dijkstra implementation costs approximately$​1035\\mathdollar 1035, and the reported score is the fractional reductionmax⁡\(0,1−cost/1035\)\\max\(0,1\-\\mathrm\{cost\}/1035\)\. The workspace persists across episodes, so each episode continues from the previous program rather than starting from scratch\.

### F\.1Auction\-Derived Topology

The CloudCast implementation contains six roles:Reader,Planner,Implementer,Builder,Evaluator, andFinalizer\. TheReaderinspects files, thePlannerproposes sub\-goals, theImplementeredits code, theBuilderruns a build or import check, theEvaluatorcallsrequest\_eval\(\), and theFinalizersubmits the program viafinal\_answer\. At each step, every living agent runs a wakeup judge; eligible agents submit fixed bids, and the highest bidder acts\. CloudCast does not use same\-role blocking, so consecutive same\-role turns are emergent rather than forbidden\. The only structural restriction is terminal: onlyFinalizer\-tagged agents may bid near the end\.

Table[7](https://arxiv.org/html/2606.02859#A6.T7)summarizes the completed checkpoint episodes\. Unlike the scientific\-research case, path length does not monotonically decrease\. This is expected: because the codebase persists, some episodes are one edit away from improvement, while others require multi\-edit search or regression repair\.

Table 7:Auction\-derived topology for completed CloudCast checkpoint episodes\. Superscripts denote consecutive same\-role turns\. Short paths occur when the persisted workspace is close to a new high; long paths occur during multi\-edit or regression\-repair phases\.The table shows two patterns\. First, topology length tracks residual uncertainty rather than training time\. Episode 9 is a clean one\-edit improvement and collapses to four steps\. Episode 15 is a partition\-sweep phase and stretches to the full step budget\. Second, the role mix becomes more selective\. Late episodes often rely on aReader/Implementer/Evaluator/Finalizercore, whilePlannerandBuilderappear when the workspace state calls for them\. The auction is unchanged; the local wakeup landscape changes as prompts and the workspace evolve\.

Episode 27 is the peak checkpoint, with score0\.3650\.365, or36\.5%36\.5\\%cost reduction relative to the Dijkstra seed \($657vs\.$1035\)\. Its topology is:

R→I→B→E→I2→E→I→B→I→E→F\\textsc\{R\}\\rightarrow\\textsc\{I\}\\rightarrow\\textsc\{B\}\\rightarrow\\textsc\{E\}\\rightarrow\\textsc\{I$\{\}^\{2\}$\}\\rightarrow\\textsc\{E\}\\rightarrow\\textsc\{I\}\\rightarrow\\textsc\{B\}\\rightarrow\\textsc\{I\}\\rightarrow\\textsc\{E\}\\rightarrow\\textsc\{F\}This path is not manually scripted\. TheReaderstops after summarizing the relevant file, theImplementercontinues while verifier feedback names unresolved scenarios, theBuilderwakes up when static validity is uncertain, theEvaluatorwakes up when a score\-changing edit is plausible, and theFinalizercommits only when further iteration appears unlikely to help\.

### F\.2Prompt Evolution: Checks Before Costly Actions

The topology changes because the agents’ prompts change\. Across the run, mutations occur in four roles:Evaluator,Builder,Implementer, andFinalizer\. The unifying pattern is simple: each role learns a cheap structural check before its expensive action\. TheEvaluatorchecks markers beforerequest\_eval\(\), theBuilderchecks symbols before running a build command, theImplementerstates intent and re\-reads afterwrite\_file, and theFinalizerrechecks invariants beforefinal\_answer\.

Table 8:Prompt mutations in CloudCast\. Each mutation adds an operational check at the point where the previous failure surfaced\. The learned content is not a new routing heuristic, but a better discipline for deciding when to act\.The initialEvaluatorprompt is mostly a wakeup rule:

WakeupwheneverImplementerhaswrittennewcodesincethelasteval\.

Ifthetaskusespure\-Pythoncode,youcanevaldirectlywithoutwaiting

foraBuildergreen\.Donotevaltwiceinarowwithoutnewcodein

between,anddonotevalonabuild/importfailure\.

After earlyΔ=0\\Delta=0evaluations, the first mutation adds a pre\-evaluation validation layer:

Do:runquicklocalchecks\(syntax/import,lightweighttests,anda

schema/validatorsmoketest\),useinexpensiveheuristicstoestimate

whetherthechangeislikelytoyieldameaningfulimprovement,and

onlycallrequest\_evalwhenthosechecksandtheestimateindicatea

realchanceofimprovement\.

A later mutation makes this check more structural:

Beforecallingrequest\_eval\(\):

\-Do:runaquickcompile/importcheck,runlightweightunit/smoke

testsorstaticsanitychecks,verifyrequiredinterfaces/markers

remainintact,andrunadeterministiclocalheuristic/estimate

thatjustifiesanexpectedimprovement;documenttherationalefor

whyanevalislikelytochangethescore\.

This illustrates the general mechanism\. TheEvaluatordoes not learn a new optimization algorithm\. It learns to avoid wasting expensive evaluation calls\. The same motif then appears in other roles:Builderlearns static checks before build commands,Implementerlearns intent\-and\-verify editing, andFinalizerlearns to recheck invariants before committing\. Thus, prompt evolution changes the wakeup decisions, and the changed wakeup decisions reshape the auction\-selected topology\.

## Appendix GWhy the Generalist Does Not Monopolize

This appendix studies a potential failure mode of the economy: if one agent is given access to all tools, will it monopolize the market and eliminate specialization? In Finance\-Agent\-Bench, we introduce a*full\-tool generalist*agent with access to all four tools \(edgar\_search,web\_search,parse\_html\_page,retrieve\_information\), alongside the ordinary specialized agents\. At first glance, one might expect such an agent to dominate the society\. It has the broadest action space, the least tool\-level restriction, and in principle can subsume filing search, web search, document parsing, retrieval, and answer submission\. Yet empirically it does not monopolize control\. Specialized agents remain persistently active and frequently win auctions\.

This failure to monopolize is informative: in our economy, being more general is not the same as being more competitive\. The generalist must spread its prompt budget across heterogeneous requirements: tool formatting, decomposition, temporal coverage, accounting consistency, numerical verification, cross\-source reconciliation, and final answer submission\. As a result, its evolved prompt becomes broad and procedurally cautious rather than sharply discriminative\. This is useful for coverage, but economically costly\. A generalist that tries to be prepared for everything is often outbid by a specialist whose wakeup condition, search pattern, and evidence standard are tuned to one narrower subproblem\. The generalist does not fail because it is weak; it fails to monopolize because it is too general\. The society rewards agents that are locally precise, not globally omnibus\.

### G\.1Initial Generalist Prompt

The initial generalist prompt is intentionally broad\. It enforces tool\-call correctness and a basic search\-first workflow, but it does not prescribe domain\-specific decomposition, numerical discipline, or an evidence hierarchy beyond a generic sequence of actions\.

TOOL\-CALLFORMATRULES\(donotviolate\):

\-edgar\_search:passCIKnumbersasstrings,e\.g\.ciks=\[’0000002488’\]\.

\-web\_search:callweb\_search\(search\_query="\.\.\."\),NOTweb\_search\(query="\.\.\."\)\.

\-parse\_html\_page:passavalidURLstring\.

\-retrieve\_information:querystoreddocumentsbykey\.

\-final\_answer:callfinal\_answer\(answer="\.\.\."\)tosubmit\.

Workmethodically:searchfirst,retrievedetails,thenanswer\.DoONE

toolcallperturn\.Whenevidenceissufficient,submitimmediately\.

This initial version is a broad operator’s manual\. It ensures syntactic correctness and a sensible high\-level order of operations, but it does not yet tell the agent how to reason under financial query structure\. In particular, it does not force explicit decomposition of the question, distinguish intermediate from final evidence, align fiscal periods, or guard against subtle financial errors such as mixing segment\-level values with consolidated totals\.

### G\.2Evolved Generalist Prompt

After evolution, the generalist prompt becomes substantially more elaborate\. The added instructions do not specialize it around one micro\-skill\. Instead, they accumulate global requirements for acting as a competent all\-purpose financial analyst\.

TOOL\-CALLFORMATRULES\(donotviolate\):

\-edgar\_search:passCIKnumbersasstrings,e\.g\.ciks=\[’0000002488’\]\.

\-web\_search:callweb\_search\(search\_query="\.\.\."\),NOTweb\_search\(query="\.\.\."\)\.

\-parse\_html\_page:passavalidURLstring\.

\-retrieve\_information:querystoreddocumentsbykey\.

\-final\_answer:callfinal\_answer\(answer="\.\.\."\)tosubmit\.

Workmethodically:Decomposethequeryintodiscreterequirements\.For

multi\-yeartrends,reportdataforeveryintermediateintervaland

prioritizeconsolidated/totalfiguresfortheprimaryentityover

segmentedorregionaldata\.Traceeverynumericalvaluedirectlytoraw

tooloutputs;donotadoptfiguresfromdialoguesummarieswithout

verification\.Ensureconsistentaccountingdefinitionsacrossallperiods\.

Cross\-referencesourcestoresolvediscrepancies\.Performasanitycheck

onthemagnitudeandtrendbeforesubmitting\.DoONEtoolcallperturn\.

Submitimmediatelyoncethefulltemporalandcategoricalscopeis

satisfied\.

Relative to generation 0, the mutation improves rigor but does not make the agent sharper about one task family\. It adds a broad bundle of global cautionary rules: decompose the query into discrete requirements, cover every relevant time interval, prefer consolidated totals over segment or regional figures, trace numbers back to raw tool outputs, enforce accounting consistency, cross\-check source discrepancies, and sanity\-check trends and magnitudes\.

These additions make the generalist safer, but they also explain why it does not monopolize\. The prompt becomes a general compliance layer rather than a highly tuned specialist heuristic\. The evolved generalist is taught to be careful everywhere, rather than exceptional in one recurring local decision\.

### G\.3Comparison to Specialized Agent Evolution

The contrast is clearest when compared with the specializedEdgarandTavilyagents\. Their initial prompts are narrow: theEdgaragent is only asked to find relevant filings, and theTavilyagent is only asked to perform one targeted web search\. After evolution, both remain narrow, but become much more exacting about the specific failure modes of their own tool domains\.

#### Edgar agent\.

The evolvedEdgarprompt does not try to become a universal analyst\. It sharpens around filing\-specific correctness: identifying the exact entity, filing type, and fiscal period; distinguishing aggregate totals from plan\- or segment\-specific values; checking filing date and fiscal year; locating future\-dated projections inside the latest official filing; and matching the numerical answer to the exact qualifiers of the query\. This is a specialization trajectory: the prompt is refined to avoid a small set of high\-value, filing\-specific mistakes\.

#### Tavily agent\.

Likewise, the evolvedTavilyprompt becomes more disciplined about source reliability and arithmetic\. It learns to identify the knowledge gap before search, trace numerical claims back to primary evidence, avoid substituting non\-matching periods or metrics, cross\-reference sources, and recompute arithmetic from raw values\. Again, this is not a move toward generality\. It is a move toward narrower epistemic discipline within one tool domain\.

By contrast, the full\-tool generalist must absorb all of these concerns at once\. Its prompt therefore grows by accumulating heterogeneous global obligations rather than repeatedly repairing one local competence\. Economically, this matters because auctions reward agents that are locally high\-value under a particular state\. A specialist can be the best agent precisely when its narrow prompt applies\. The generalist is broad, but its local edge is diluted\.

### G\.4Generalist Failure as Evidence for Specialization

This case study illustrates a general principle of the economy: prompt evolution is most effective when repeated feedback can be compressed into a small number of role\-specific decision rules\. When an agent has a narrow tool interface and a narrow evidentiary responsibility, the mutator can improve it by writing concrete instructions that target the same class of mistakes repeatedly\. The resulting prompt becomes sharper, more falsifiable, and more economically competitive\.

The generalist does not enjoy this advantage\. Because it is responsible for too many heterogeneous subtasks, each mutation tends to add another layer of generic caution rather than strengthening one specific inference pattern\. Its prompt becomes longer, broader, and more procedurally safe, but not more economically dominant\. Thus, the generalist does not monopolize: over\-generality is not specialization, and in this market, specialization is what wins control\.

Similar Articles

Economy of Minds: Emerging Multi-Agent Intelligence with Economic Interactions

Hugging Face Daily Papers

This paper proposes an 'agent economy' framework inspired by Hayek's economic theory, where agents self-organize through auction-based competition and economic selection to produce emergent multi-step reasoning and collective intelligence without centralized control. The system outperforms stronger monolithic baselines across five agentic tasks including mathematical reasoning, financial research, and scientific research.

Agent Bazaar: Enabling Economic Alignment in Multi-Agent Marketplaces

Hugging Face Daily Papers

Introduces Agent Bazaar, a multi-agent simulation framework for evaluating economic alignment of LLMs, identifying failure modes like algorithmic instability and Sybil deception, and training a 9B model that outperforms frontier models using targeted reinforcement learning.

AI Science & Economy: Systems Map

Reddit r/artificial

This article argues that while AI excels at pattern recognition and hypothesis generation, scientific and economic progress requires grounded interaction with reality and institutional execution, emphasizing the need for human-AI collaboration.

Can urban economics help model and improve agentic AI systems?

Reddit r/ArtificialInteligence

This article explores how concepts from urban economics, such as traffic, zoning, and pollution, can model externalities in agentic AI systems. It introduces a Behavioral Externality Multiplier (BEM) and proposes a layered framework involving architecture, substrate, and governance to measure and mitigate costly consequences of cheap AI actions.