A toy framework for single and multi-agent human-AI curiosity ecosystems

arXiv cs.AI Papers

Summary

This paper introduces a toy framework that models curiosity as an ecosystem in single and multi-agent settings, exploring how agents weigh immediate uncertainty reduction, costs, delayed returns, and the value of keeping questions open. It aims to inform future multi-agent AI systems for discovery.

arXiv:2607.06214v1 Announce Type: new Abstract: This paper offers a toy framework for considering curiosity as an ecosystem. First, it suggests that a single agent's inquiry policy (how, when, and why an agent asks a question) depends on how the agent values immediate uncertainty reduction, costs, delayed return, and the value of keeping the question open. A key concept in the framework is that the weights on these decision-related terms can change with experience. For example, a period of cheap, quickly answered questions may change the cost of inquiry on a short timescale and change which kinds of questions the agent is drawn to answer over a longer timescale. Second, these ideas are extended to many agents exploring a shared knowledge landscape, and there the framework tracks inquiry volume, topic diversity, frontier-directed inquiry, redundancy, and reusable knowledge. The result is a conceptual toy framework for studying curiosity ecology and for future efforts towards designing multi-agent AI systems for discovery. It serves as a companion piece for a paper currently under review in Trends in Neurosciences.
Original Article
View Cached Full Text

Cached at: 07/08/26, 04:39 AM

# A toy framework for single and multi-agent human-AI curiosity ecosystems
Source: [https://arxiv.org/html/2607.06214](https://arxiv.org/html/2607.06214)
Ilya E\. Monosov Solomon H\. Snyder Department of Neuroscience, Johns Hopkins University, Baltimore, MD, USA Departments of Biomedical Engineering, Electrical and Computer Engineering, and Psychiatry, Johns Hopkins University, Baltimore, MD, USA Zanvyl Krieger Mind/Brain Institute, Johns Hopkins University, Baltimore, MD, USA Data Science and Artificial Intelligence Institute and the Kavli Neuroscience Discovery Institute, Johns Hopkins University, Baltimore, MD, USA ilya\.monosov@gmail\.com

###### Abstract

This paper offers a toy framework for considering curiosity as an ecosystem\.First, it suggests that a single agent’s inquiry policy \(how, when, and why an agent asks a question\) depends on how the agent values immediate uncertainty reduction, costs, delayed return, and the value of keeping the question open\. A key concept in the framework is that the weights on these decision\-related terms can change with experience\. For example, a period of cheap, quickly answered questions may change the cost of inquiry on a short timescale and change which kinds of questions the agent is drawn to answer over a longer timescale\.Second, these ideas are extended to many agents exploring a shared knowledge landscape, and there the framework tracks inquiry volume, topic diversity, frontier\-directed inquiry, redundancy, and reusable knowledge\. The result is a conceptual toy framework for studying curiosity ecology and for future efforts towards designing multi\-agent AI systems for discovery\. It serves as a companion piece for a paper currently under review in Trends in Neurosciences\.

## Highlights

- •Asking a question is a choice among competing values in immediate uncertainty reduction, effort, delayed return, and the value of leaving the question open\.
- •The weights on those values can change or drift\. Repeated exposure to fast cheap answers may gradually encourage fast resolution to be more attractive and long\-horizon inquiry less attractive\.
- •The framework can help to generate future theories that can be applied to groups of people or AI agents that share and explore a knowledge landscape\. Over time, it can also inform multi\-agent AI development\.
- •This work supports future studies of how inquiry becomes adaptive and generative, or maladaptive\.

## 1Introduction

Why do we ask certain questions and avoid asking others? A question can tempt us because the answer will remove lots of uncertainty, because the answer may pay off substantially, even if only later, or because it is easy to pursue, or simply because everyone around us is already asking it\. Models of curiosity and information\-seeking suggest that the subjective value we place on an answer rises and falls with how uncertain we are and how much uncertainty we think the answer will resolve\(for review see Monosov,[2024](https://arxiv.org/html/2607.06214#bib.bib24); Bromberg\-Martin and Monosov,[2020](https://arxiv.org/html/2607.06214#bib.bib6)\)\. This subjective value of information can change on a moment\-by\-moment basis\.

A lasting change in context can change how our preferences are expressed, or change the preferences themselves\(Tversky and Simonson,[1993](https://arxiv.org/html/2607.06214#bib.bib36)\)\. One neurobiologically inspired way to formalize this is to assume that the brain has weights for different features or attributes of a question \(e\.g\., amount of uncertainty it would reduce, time to ask, cost, delayed payoff, and so on\) and to allow the weights to move, drift, or be updated systematically in response to context and experience\. Then, the questions a person has already been asking \(their history\), the cost of the tools available, the questions being asked by their peers, and the value attached to answers in the surrounding environment shape what questions feel like they are worth pursuing next\. The same question can look trivial in one ecology or one agent and valuable in another ecology \(to the same or to other agents\)\.

This paper attempts to consider this complexity by linking three processes that are often treated separately: the process through which a single agent chooses to pursue or not to pursue a question, the drift in that agent’s inquiry preferences that occurs as a function of their own experience and context, and the process through which many agents build up a shared stock of knowledge and express curiosity collectively\. The resulting toy framework can be helpful in considering the adaptivity of a curiosity ecosystem and for future work designing adaptive multi\-agent AI systems\.

## 2Single agent’s inquiry

Consider a single agentiideciding whether to pursue questionqqat timett\. The agent’s private value for that question is

Vi​\(q,t\)=ψi​\(t\)​𝔼i,t​\[Ii​\(q,t\)\]−λi​\(t\)​𝔼i,t​\[Ci​\(q,t\)\]\+τi​\(t\)​𝔼i,t​\[Li​\(q,t\)\]−ϕi​\(t\)​𝔼i,t​\[Oi​\(q,t\)\]\.V\_\{i\}\(q,t\)=\\psi\_\{i\}\(t\)\\,\\mathbb\{E\}\_\{i,t\}\[I\_\{i\}\(q,t\)\]\-\\lambda\_\{i\}\(t\)\\,\\mathbb\{E\}\_\{i,t\}\[C\_\{i\}\(q,t\)\]\+\\tau\_\{i\}\(t\)\\,\\mathbb\{E\}\_\{i,t\}\[L\_\{i\}\(q,t\)\]\-\\phi\_\{i\}\(t\)\\,\\mathbb\{E\}\_\{i,t\}\[O\_\{i\}\(q,t\)\]\.\(1\)HereIiI\_\{i\}is the immediate value of reducing uncertainty\.CiC\_\{i\}is the cost of asking \(e\.g\., effort, time, attention\)\.LiL\_\{i\}is the long\-horizon return, meaning what the answer may become useful for later\.OiO\_\{i\}is the value of not answering yet\. That is the value of keeping the question open, either for flexibility\(McDonald and Siegel,[1986](https://arxiv.org/html/2607.06214#bib.bib23); Dixit and Pindyck,[1994](https://arxiv.org/html/2607.06214#bib.bib12)\)or because some information is actively not desired by the agent\(Golmanet al\.,[2017](https://arxiv.org/html/2607.06214#bib.bib17); Gigerenzer and Garcia\-Retamero,[2017](https://arxiv.org/html/2607.06214#bib.bib16)\)\. This is further inspired by neural data that suggest that information about rewards and punishments can be valued differently, with partly distinct circuits for information seeking and information avoidance\(Jezziniet al\.,[2021](https://arxiv.org/html/2607.06214#bib.bib19); Charpentieret al\.,[2018](https://arxiv.org/html/2607.06214#bib.bib9)\)\.OiO\_\{i\}is meant to capture cases in which an agent would decline to resolveqqright now even when it could be tempting\. A scientist may keep a promising hypothesis deliberately open to preserve the flexibility of running a more informative experiment, or a person may prefer not to learn a medical or genetic result\. In these examples, keeping the question open is distinct from, and can even act against, the long\-horizon returnLiL\_\{i\}\.

The weightsθi​\(t\)=\(ψi​\(t\),λi​\(t\),τi​\(t\),ϕi​\(t\)\)\\theta\_\{i\}\(t\)=\(\\psi\_\{i\}\(t\),\\lambda\_\{i\}\(t\),\\tau\_\{i\}\(t\),\\phi\_\{i\}\(t\)\)reflect how much the agent cares about each term and define its*curiosity policy*\. All four are written as expectations at timett\(𝔼i,t​\[⋅\]\\mathbb\{E\}\_\{i,t\}\[\\cdot\]\), prior to askingqq\. Looking up who won a chess tournament last night is mostly driven byIiI\_\{i\}: it is expected to resolve uncertainty immediately and costs almost nothing\. Working through a hard homework problem is costly now, soCiC\_\{i\}is high, but the delayed returnLiL\_\{i\}may be high as well\.

IiI\_\{i\}andLiL\_\{i\}are defined by how many questions must be resolved for their value to be realized\.IiI\_\{i\}is the myopic value of resolvingqqon its own as if no further question were asked; this value can arrive far in the future and still count asIiI\_\{i\}so long as it depends only onqq\.LiL\_\{i\}is instead the value that requires the follow\-up questions opened byqqto be resolved\. The long\-horizon term can be split into private and social returns:

Li​\(q,t\)=Liown​\(q,t\)\+κi​Lisocial​\(q,t\)\.L\_\{i\}\(q,t\)=L\_\{i\}^\{\\mathrm\{own\}\}\(q,t\)\+\\kappa\_\{i\}L\_\{i\}^\{\\mathrm\{social\}\}\(q,t\)\.The first component is the future payoff that is collected by the agent itself\. The second is the payoff that spills over to other agents, discounted byκi∈\[0,1\]\\kappa\_\{i\}\\in\[0,1\]\. A self\-interested agent hasκi=0\\kappa\_\{i\}=0; an agent that cares about downstream benefits to others hasκi\>0\\kappa\_\{i\}\>0\.

A second important property of a question is its generativitygi​\(q,t\)g\_\{i\}\(q,t\): the number of new knowledge gaps that resolvingqqopens, relative to what the agent already knows\. This is a gross count of the gaps thatqqopens \(the one gap thatqqcloses by being answered is not deducted from it, so a question that opens even a single new gap hasgi\>0g\_\{i\}\>0\)\. The simple intuition for this definition is that learning can make a person aware of what they don’t know\(Loewenstein,[1994](https://arxiv.org/html/2607.06214#bib.bib21); Murayama,[2022](https://arxiv.org/html/2607.06214#bib.bib25)\)\. For example, a trivia question usually closes its own knowledge gap and often opens little opportunity for future uncertainty reduction, sogi≈0g\_\{i\}\\approx 0\. A question embedded in a sparse region of the knowledge landscape can do the opposite\. For example, new scientific data can answer an immediate question and open an entirely new research program\.

For the agent’s own return, a simple specification is

Liown​\(q,t\)=ℓ0\+ρi⋅gi​\(q,t\),L\_\{i\}^\{\\mathrm\{own\}\}\(q,t\)=\\ell\_\{0\}\+\\rho\_\{i\}\\cdot g\_\{i\}\(q,t\),whereℓ0\\ell\_\{0\}is a baseline return that is constant across questions, andρi\\rho\_\{i\}converts newly opened knowledge gaps into expected future payoff\.

Note that this is not the same asOiO\_\{i\}in Equation[1](https://arxiv.org/html/2607.06214#S2.E1)\.OiO\_\{i\}is the value of not closing a question right now\. In contrast,gig\_\{i\}counts the new questions that appear once a questionisclosed\.

Also note that a cheap question high inIiI\_\{i\}can be near zero ingig\_\{i\}\. Such a question could feel good to answer now \(due to large uncertainty reduction\), but it would not create future opportunities for more uncertainty reduction or learning\. A frontier question \(that is one that is difficult but may produce discoveries\) may be less satisfying in the moment and more costly to pursue, but can be high ingig\_\{i\}and therefore high inLiL\_\{i\}\. Which one is chosen depends on the relative weight placed on immediate resolution and long\-horizon return, especiallyψi\\psi\_\{i\}andτi\\tau\_\{i\}\.

The social componentLisocialL\_\{i\}^\{\\mathrm\{social\}\}can be specified the same way, with the group in place of the individual\. It depends on the question’s*population generativity*gs​\(q,t\)g\_\{s\}\(q,t\): the number of new knowledge gaps it opens relative to what the group already knows rather than to one agent’s knowledge \(treated in Section[5](https://arxiv.org/html/2607.06214#S5)\)\. There are two ways to specify it\. At the micro level, the social return is the spillover thatii’s resolution ofqqcreates for everyone else,

Lisocial​\(q,t\)=∑j≠iΔ​Ljown​\(q,t\),L\_\{i\}^\{\\mathrm\{social\}\}\(q,t\)=\\sum\_\{j\\neq i\}\\Delta L^\{\\mathrm\{own\}\}\_\{j\}\(q,t\),whereΔ​Ljown​\(q,t\)\\Delta L^\{\\mathrm\{own\}\}\_\{j\}\(q,t\)is the long\-horizon own\-return thatii’s answer confers on agentjj\. But this requires knowing every pairwise spillover\. At the macro level, when those pairwise terms are not observable, the same concept is summarized through population generativity,

Lisocial​\(q,t\)=ℓ0soc\+ρisoc​gs​\(q,t\),L\_\{i\}^\{\\mathrm\{social\}\}\(q,t\)=\\ell\_\{0\}^\{\\mathrm\{soc\}\}\+\\rho\_\{i\}^\{\\mathrm\{soc\}\}\\,g\_\{s\}\(q,t\),whereρisoc\\rho\_\{i\}^\{\\mathrm\{soc\}\}converts group\-level gaps opened into expected downstream payoff to others \(the social counterpart ofρi\\rho\_\{i\}, and likewise agent\-specific because agents differ in how far their answers propagate\)\.

The agent is more likely to pursue higher\-value questions\. In the framework, this can be implemented with a standard pursuit probability\. LetPi​\(q,t\)P\_\{i\}\(q,t\)be the probability that agentiipursuesqqat timett, rising withVi​\(q,t\)V\_\{i\}\(q,t\)under a threshold rule, a logistic rule, or a softmax with an outside option \(a fixed no\-pursuit alternative included in the choice set, so that the agent may also choose to ask nothing\)\.Pi​\(q,t\)P\_\{i\}\(q,t\)is then a per\-question pursuit probability\. An agent can pursue several questions, or none, and the total volume of inquiryQtQ\_\{t\}is free to rise or fall\.

## 3Preference drift

A new tool or a change in social settings can change Equation[1](https://arxiv.org/html/2607.06214#S2.E1)by changing the terms themselves, especially by lowering the costCiC\_\{i\}\. However, over longer timescales, repeated inquiry may also change the weights themselves\. A person repeatedly rewarded by fast answers may gradually become biased toward that kind of question\.

Recent inquiry is summarized by an experience stateXi​\(t\)X\_\{i\}\(t\), a smoothed version of the profilexi​\(t\)x\_\{i\}\(t\)of questions the agent has resolved:

xi​\(t\)=\[cost,speed,uncertainty reduction,generativity,reusability\]⊤\.x\_\{i\}\(t\)=\[\\,\\text\{cost\},\\ \\text\{speed\},\\ \\text\{uncertainty reduction\},\\ \\text\{generativity\},\\ \\text\{reusability\}\\,\]^\{\\top\}\.PersistenceγX∈\[0,1\]\\gamma\_\{X\}\\in\[0,1\]controls how slowly past experience fades, inspired by economic models in which current preferences depend on past consumption\(Ryder and Heal,[1973](https://arxiv.org/html/2607.06214#bib.bib29); Becker and Murphy,[1988](https://arxiv.org/html/2607.06214#bib.bib4)\):

Xi​\(t\)=γX​Xi​\(t−1\)\+\(1−γX\)​xi​\(t\)\.X\_\{i\}\(t\)=\\gamma\_\{X\}\\,X\_\{i\}\(t\-1\)\+\(1\-\\gamma\_\{X\}\)\\,x\_\{i\}\(t\)\.\(2\)If no question is resolved, setXi​\(t\)=Xi​\(t−1\)X\_\{i\}\(t\)=X\_\{i\}\(t\-1\)\. The agent updates only from the questions it actually pursued and resolved\. Question selection shapes experience, and experience then feeds back onto selection\. And, the weights update through a drift operator:

θi​\(t\+1\)=θi​\(t\)\+Δ​θi​\(Xi​\(t\),Mt\)\.\\theta\_\{i\}\(t\+1\)=\\theta\_\{i\}\(t\)\+\\Delta\\theta\_\{i\}\\big\(X\_\{i\}\(t\),M\_\{t\}\\big\)\.\(3\)HereMtM\_\{t\}is the surrounding inquiry ecology\. It includes the cost and tool landscape, the questions and values circulating in the population, and the shared knowledge stockStS\_\{t\}already available \(Section[4](https://arxiv.org/html/2607.06214#S4)\)\. The drift termΔ​θi\\Delta\\theta\_\{i\}describes how a given region of this ecology nudges the weights\(ψi,λi,τi,ϕi\)\(\\psi\_\{i\},\\lambda\_\{i\},\\tau\_\{i\},\\phi\_\{i\}\)while keeping them in their natural nonnegative range\. A curiosity regime is a region of the joint space ofXi​\(t\)X\_\{i\}\(t\)andMtM\_\{t\}\. For example, in a cheap\-inquiry regime, fast and low\-cost resolution is abundant\. The same historyXi​\(t\)X\_\{i\}\(t\)can have different effects under different ecologies\. One plausible implementation of the drift is reinforcement learning\. Quickly answered questions can act as small, frequent subjective rewards\. Over longer timescales, such rewards may reshape the value function that guides inquiry, for example through direct changes in valuation circuits\(Sutton and Barto,[2018](https://arxiv.org/html/2607.06214#bib.bib35); Schultzet al\.,[1997](https://arxiv.org/html/2607.06214#bib.bib30)\)\.

Different forms ofΔ​θi\\Delta\\theta\_\{i\}could be useful for future simulations\. One example is a linear update with decay,

Δ​θi​\(Xi​\(t\),Mt\)=ε​\(W​Xi​\(t\)\+wM​m​\(Mt\)−θi​\(t\)\),\\Delta\\theta\_\{i\}\\big\(X\_\{i\}\(t\),M\_\{t\}\\big\)=\\varepsilon\\,\\big\(W\\,X\_\{i\}\(t\)\+w\_\{M\}\\,m\(M\_\{t\}\)\-\\theta\_\{i\}\(t\)\\big\),\(4\)whereε\>0\\varepsilon\>0is a learning rate,WWmaps the experience profileXi​\(t\)X\_\{i\}\(t\)onto target weights,m​\(Mt\)m\(M\_\{t\}\)is a vector summary of the ecology, andwMw\_\{M\}scales its influence\. The weights are clipped at zero after each update to keep\(ψi,λi,τi,ϕi\)\(\\psi\_\{i\},\\lambda\_\{i\},\\tau\_\{i\},\\phi\_\{i\}\)in their natural nonnegative range\. Hereθi\\theta\_\{i\}relaxes toward a target set jointly by recent experience and the surrounding ecology, and the regime\-dependent sign patterns \(discussed below\) correspond to the signs of the relevant entries ofWWandwMw\_\{M\}, and can flip across ecologies through the ecology summarym​\(Mt\)m\(M\_\{t\}\)\.

Other update rules, such as reinforcement learning, can be substituted without changing the rest of the framework\.

Whether repeated cheap inquiry makes the agent*more*or*less*effort\-averse—the sign ofΔ​λi\\Delta\\lambda\_\{i\}—depends on the ecologyMtM\_\{t\}\. In terms of Equation[4](https://arxiv.org/html/2607.06214#S3.E4), it depends on which channel dominates the drift: recent experience,W​Xi​\(t\)W\\,X\_\{i\}\(t\), or the surrounding ecology,wM​m​\(Mt\)w\_\{M\}\\,m\(M\_\{t\}\)\.

Consider first a*habituation regime*, in which experience has a particularly strong influence on the agent’s policy\. The agent keeps receiving cheap answers that are low in generativity and reusability\. Quick, low\-cost resolution gets rewarded, so effort tolerance erodes and the agent becomes*more*cost\-sensitive:

Δ​ψi\>0,Δ​λi\>0,Δ​τi<0,Δ​ϕi<0\.\\Delta\\psi\_\{i\}\>0,\\quad\\Delta\\lambda\_\{i\}\>0,\\quad\\Delta\\tau\_\{i\}<0,\\quad\\Delta\\phi\_\{i\}<0\.Each cheap answer makes the next one more attractive\. Now consider an*abundance regime*, in which the broad ecology—rather than the agent’s own recent experience—dominates the agent’s preference drift\. The agent has only a limited budget of effort to spend, but sits in an ecology where almost every question is fast and cheap and the costly, high\-value questions are rare\. Because those cheap questions consume so little of the budget, effort is no longer the binding constraint on inquiry: the agent can afford to take on each expensive frontier question it occasionally encounters and collect its high return\. Ifλi\\lambda\_\{i\}tracks how scarce effort feels, this abundance of cheap answers lowers it, and the agent becomes*less*cost\-sensitive:

Δ​ψi\>0,Δ​λi<0,Δ​τi<0,Δ​ϕi<0\.\\Delta\\psi\_\{i\}\>0,\\quad\\Delta\\lambda\_\{i\}<0,\\quad\\Delta\\tau\_\{i\}<0,\\quad\\Delta\\phi\_\{i\}<0\.The long\-horizon weightτi\\tau\_\{i\}still falls, because the experience state remains dominated by quick rewards\. But the cost channel now works the other way\. A lowerλi\\lambda\_\{i\}makes the costly frontier question easier to accept, even as the risingψi\\psi\_\{i\}pulls toward the cheap one\. The two channels are in tension, and the signs alone no longer settle the outcome\.

To see which way the balance tips, let the agent choose between a cheap questionccand a frontier questionff\. WriteΔ​I=Ic−If\\Delta I=I\_\{c\}\-I\_\{f\},Δ​C=Cc−Cf\\Delta C=C\_\{c\}\-C\_\{f\}, andΔ​L=Lc−Lf\\Delta L=L\_\{c\}\-L\_\{f\}; I ignore theOiO\_\{i\}term for simplicity\. The cheap question resolves more uncertainty now \(Δ​I\>0\\Delta I\>0\), costs less \(Δ​C<0\\Delta C<0\), and pays less later \(Δ​L<0\\Delta L<0\)\. The drift strengthens the pull toward the cheap question only when

\(Δ​I\)​Δ​ψi\+\(Δ​L\)​Δ​τi\>\(Δ​C\)​Δ​λi\.\(\\Delta I\)\\,\\Delta\\psi\_\{i\}\+\(\\Delta L\)\\,\\Delta\\tau\_\{i\}\\;\>\\;\(\\Delta C\)\\,\\Delta\\lambda\_\{i\}\.\(5\)The left side is the gain in attraction to quick resolution, plus the weakened pull of future payoff\. The right side is the change in the cost term\. In the habituation regime,Δ​λi\>0\\Delta\\lambda\_\{i\}\>0makes the right side negative, so the inequality holds whenever the drifts take the signs above\. In the*abundance*regime,Δ​λi<0\\Delta\\lambda\_\{i\}<0makes the right side positive, so the cost term works*against*the pull toward cheap: shallow inquiry wins only if the immediacy channel outweighs it\. The same history of cheap answers can therefore push an agent toward shallow inquiry in one ecology but not in another\.

The model does not predict that cheaper access to answers always degrades curiosity\. For example, if tools free up our time to pursue harder questions, or if cheap answers themselves open new knowledge gaps, the weight on long\-horizon returnτi\\tau\_\{i\}may not drop\. The drift in Equation[3](https://arxiv.org/html/2607.06214#S3.E3)should not automatically be read as a change in deep preference\(Stigler and Becker,[1977](https://arxiv.org/html/2607.06214#bib.bib34)\)\. It may also reflect an altered environment or altered expression of a policy\. Also, the drift is regulated by the kind of questions asked, not just by the number of questions\. Two people can ask the same number of questions in a day and drift in opposite directions\. Finally, the current examples emphasize cost and speed, but the same update rule can be applied to generativity, that is, to whether an answer opened new questions and created reusable by\-products rather than only resolving a single question\.

## 4Multiple agents and shared knowledge

Discoveries are not solitary\. People and machines explore a shared knowledge landscape and borrow from and duplicate one another, and one agent’s work can make another agent’s next question cheaper\. This is a version of the exploration–exploitation problem\(Nelson and Winter,[1982](https://arxiv.org/html/2607.06214#bib.bib26); March,[1991](https://arxiv.org/html/2607.06214#bib.bib22); Averbeck,[2015](https://arxiv.org/html/2607.06214#bib.bib3); Costa and Averbeck,[2020](https://arxiv.org/html/2607.06214#bib.bib11)\), except that here one agent’s answer can change both the value and the cost of another agent’s inquiry\.

To include this, each agent values a question along two dimensions: its private value, as in Equation[1](https://arxiv.org/html/2607.06214#S2.E1), and what it may contribute to a shared stock of knowledge\. Formally, the shared stock has two layers: a set𝒦t\\mathcal\{K\}\_\{t\}of resolved, deposited results, against which redundancyd​\(q,𝒦t\)d\(q,\\mathcal\{K\}\_\{t\}\)and reusable valueU​\(q,𝒦t\)U\(q,\\mathcal\{K\}\_\{t\}\)are evaluated, and a scalar indexStS\_\{t\}of total reusable value, which evolves according to Equation[7](https://arxiv.org/html/2607.06214#S4.E7)\. The effective value of a question is

Vieff​\(q,t\)=Vi​\(q,t\)\+ηi​𝔼i,t​\[U​\(q,𝒦t\)∣q​resolved\]−ζi​Rred​\(q,t\)\.V\_\{i\}^\{\\mathrm\{eff\}\}\(q,t\)=V\_\{i\}\(q,t\)\+\\eta\_\{i\}\\,\\mathbb\{E\}\_\{i,t\}\[U\(q,\\mathcal\{K\}\_\{t\}\)\\mid q\\ \\text\{resolved\}\]\-\\zeta\_\{i\}R\_\{\\mathrm\{red\}\}\(q,t\)\.\(6\)The middle term rewards a reusable contribution to𝒦t\\mathcal\{K\}\_\{t\}, scaled byηi\\eta\_\{i\}\. HereU​\(q,𝒦t\)≥0U\(q,\\mathcal\{K\}\_\{t\}\)\\geq 0is the reusable, non\-redundant value that resolvingqqadds to the shared knowledge stock: value that other agents can later build on or reuse rather than having to reproduce, such as a result, tool, or dataset that lowers the cost of, or opens, their subsequent questions \(a result whose content is already in𝒦t\\mathcal\{K\}\_\{t\}contributesU=0U=0\)\. It is specified in Equation[7](https://arxiv.org/html/2607.06214#S4.E7)\. The last term penalizes redundancy, scaled byζi\\zeta\_\{i\}\. These two terms act at a different stage than the social benefits insideViV\_\{i\}\. The long\-horizon social returnLisocialL\_\{i\}^\{\\mathrm\{social\}\}is a delayed spillover to others, discounted byκi\\kappa\_\{i\}\. In contrast,ηi\\eta\_\{i\}is an immediate reward for depositing a reusable result, andζi\\zeta\_\{i\}is a penalty for duplicating existing or concurrent work\.

A redundancy score is

Rred​\(q,t\)=ωd​d​\(q,𝒦t\)\+ωp​∑j≠iPj​\(q,t\),R\_\{\\mathrm\{red\}\}\(q,t\)=\\omega\_\{d\}\\,d\(q,\\mathcal\{K\}\_\{t\}\)\+\\omega\_\{p\}\\sum\_\{j\\neq i\}P\_\{j\}\(q,t\),whered​\(q,𝒦t\)d\(q,\\mathcal\{K\}\_\{t\}\)measures overlap with what is already known, and the sum measures how much other agents are already pursuing the same question\. The nonnegative weightsωd,ωp≥0\\omega\_\{d\},\\omega\_\{p\}\\geq 0convert these two distinct sources of overlap onto a common value scale; Section[6](https://arxiv.org/html/2607.06214#S6)discusses setting them separately, includingωp<0\\omega\_\{p\}<0when parallel replication is desirable\. These terms can affect the choice of a question, the decision to preserve a result in𝒦t\\mathcal\{K\}\_\{t\}, or both\.

The scalar index of shared knowledge evolves as an expected index of reusable knowledge:

St\+1=\(1−δ\)​St\+∑q\[1−∏i\(1−Pi​\(q,t\)​ri​\(q,t\)\)\]​U​\(q,𝒦t\),S\_\{t\+1\}=\(1\-\\delta\)\\,S\_\{t\}\+\\sum\_\{q\}\\Big\[\\,1\-\\prod\_\{i\}\\big\(1\-P\_\{i\}\(q,t\)\\,r\_\{i\}\(q,t\)\\big\)\\Big\]\\;U\(q,\\mathcal\{K\}\_\{t\}\),\(7\)whereδ∈\[0,1\]\\delta\\in\[0,1\]is a staleness rate,ri​\(q,t\)r\_\{i\}\(q,t\)is the probability that pursuit succeeds, andU​\(q,𝒦t\)≥0U\(q,\\mathcal\{K\}\_\{t\}\)\\geq 0is the reusable non\-redundant value of resolvingqq\(results whose content is already in𝒦t\\mathcal\{K\}\_\{t\}produceU​\(q,𝒦t\)=0U\(q,\\mathcal\{K\}\_\{t\}\)=0\)\. Each resolved question also deposits its result into the set𝒦t\\mathcal\{K\}\_\{t\}, so the two layers advance together\.

The bracketed term is the probability that at least one agent resolvesqq, under the assumption that agents’ pursuit and success events are independent conditional on the current state; correlated strategies would require replacing it with the joint probability that at least one agent resolvesqq\. Each contribution is therefore credited once, even if several agents pursue the same question\. The sameU​\(q,𝒦t\)U\(q,\\mathcal\{K\}\_\{t\}\)also enters the agent’s own value in Equation[6](https://arxiv.org/html/2607.06214#S4.E6), but only for the questions that agent resolves\. The shared stock grows from everyone’s contributions, but the reward for depositing knowledge goes only to the agent who contributed it\.

In this framework, agents can be coupled in several ways\. First, the drift operator in Equation[3](https://arxiv.org/html/2607.06214#S3.E3)can read a peer summaryX¯−i​\(t\)\\bar\{X\}\_\{\-i\}\(t\), so an agent’s weights are shaped by its own history together with what others are asking\. That coupling can coordinate a population, but it can also make its questions more alike\. Second, one agent’s tools, explanations, or partial answers can lower the costCiC\_\{i\}for another\. Third, the shared stockStS\_\{t\}grows only from contributions that are reusable and non\-redundant, and it decays as knowledge becomes stale\. Thus the same amount of activity can leave behind very different bodies of knowledge\.

Simply adding more agents may not be enough to grow the knowledge “stock”\. A group with the same goal, the same evidence, and the same starting point can generate many parallel inquiry trajectories that converge on effectively the same knowledge rather than genuinely diverse results\. This is a concern for large language models trained on similar data and decoded in similar ways, where more samples do not necessarily mean more production of functional diversity\(Doshi and Hauser,[2024](https://arxiv.org/html/2607.06214#bib.bib13); Padmakumar and He,[2024](https://arxiv.org/html/2607.06214#bib.bib27); Shumailovet al\.,[2024](https://arxiv.org/html/2607.06214#bib.bib33)\)\. To make discoveries, a multi\-agent system needs three things: division of labor, a shared state that agents can inspect, and outputs passed to others in a reusable format\. The parametersηi\\eta\_\{i\}andζi\\zeta\_\{i\}help with discovery but they cannot replace or outpace this structure\.ηi\\eta\_\{i\}andζi\\zeta\_\{i\}are only meaningful if the agent can evaluate the quantities they act on,𝔼i,t​\[U​\(q,𝒦t\)\]\\mathbb\{E\}\_\{i,t\}\[U\(q,\\mathcal\{K\}\_\{t\}\)\]andRred​\(q,t\)R\_\{\\mathrm\{red\}\}\(q,t\)\. Computing those requires access to the shared knowledge: what is already known and what peers are pursuing\. These parameters therefore presuppose that structure rather than substituting for it\. They are most critical when agents are scarce relative to the knowledge landscape \(e\.g\., there are more open questions than agents or when the landscape is too large or even dangerous for uncoordinated search\)\.

## 5Population\-level quantities

At the population level, the analog of a single agent’s generativitygig\_\{i\}is a question’s ability to open knowledge gaps relative to what the group already knows\. In a simplified setup, this isgs​\(q,t\)g\_\{s\}\(q,t\), evaluated against the shared knowledge stockStS\_\{t\}; it is the driver of the social long\-horizon returnLisocialL\_\{i\}^\{\\mathrm\{social\}\}specified in Section[2](https://arxiv.org/html/2607.06214#S2)\. Because a generative question often benefits \(or should benefit\) the group more than the individual who asks it\(Arrow,[1962](https://arxiv.org/html/2607.06214#bib.bib1)\), agents with lowκi\\kappa\_\{i\}tend to under\-pursue such questions\.

Generative questions tend to add reusable value to the shared stock through𝔼i,t​\[U​\(q,𝒦t\)∣q​resolved\]\\mathbb\{E\}\_\{i,t\}\[U\(q,\\mathcal\{K\}\_\{t\}\)\\mid q\\ \\text\{resolved\}\]\. Questions whose generativitygsg\_\{s\}exceeds a thresholdg¯\\bar\{g\}make up the frontier setℱt\\mathcal\{F\}\_\{t\}:

ℱt=\{q:gs​\(q,t\)≥g¯\}\.\\mathcal\{F\}\_\{t\}=\\\{q:g\_\{s\}\(q,t\)\\geq\\bar\{g\}\\\}\.
To study a population’s contribution to knowledge, I use three quantities\. First, query volume is the total inquiry pursuit mass,

Qt=∑i∑qPi​\(q,t\)\.Q\_\{t\}=\\sum\_\{i\}\\sum\_\{q\}P\_\{i\}\(q,t\)\.Second, topic diversity measures how evenly that mass is spread acrossKKtopicsω\\omega, where each topicω\\omegais a mutually exclusive set of questions and theKKtopics together partition the space of questions\.

pt​\(ω\)=Qt−1​∑i∑q∈ωPi​\(q,t\),p\_\{t\}\(\\omega\)=Q\_\{t\}^\{\-1\}\\sum\_\{i\}\\sum\_\{q\\in\\omega\}P\_\{i\}\(q,t\),which I quantify using the normalized Shannon entropy\(Shannon,[1948](https://arxiv.org/html/2607.06214#bib.bib31)\)

Dt=−1log⁡K​∑ω=1Kpt​\(ω\)​log⁡pt​\(ω\)\.D\_\{t\}=\-\\frac\{1\}\{\\log K\}\\sum\_\{\\omega=1\}^\{K\}p\_\{t\}\(\\omega\)\\,\\log p\_\{t\}\(\\omega\)\.\(8\)This equals0when all inquiry sits on one topic and11when inquiry is evenly distributed, using the convention0​log⁡0=00\\log 0=0and assumingK\>1K\>1\.DtD\_\{t\}measures spread across topics but not semantic distance between them\.

The third quantity is the frontier share defined as the fraction of inquiry aimed atℱt\\mathcal\{F\}\_\{t\}:

Ft=∑i∑q∈ℱtPi​\(q,t\)∑i∑qPi​\(q,t\)\.F\_\{t\}=\\frac\{\\sum\_\{i\}\\sum\_\{q\\in\\mathcal\{F\}\_\{t\}\}P\_\{i\}\(q,t\)\}\{\\sum\_\{i\}\\sum\_\{q\}P\_\{i\}\(q,t\)\}\.\(9\)Returns to inquiry are often heavy\-tailed, so a small number of questions can account for much of the eventual impact\(Uzziet al\.,[2013](https://arxiv.org/html/2607.06214#bib.bib37); Fosteret al\.,[2015](https://arxiv.org/html/2607.06214#bib.bib15)\)\. IfQt=0Q\_\{t\}=0, I setDt=Ft=0D\_\{t\}=F\_\{t\}=0and read the population\-performance indexRtSR\_\{t\}^\{S\}\(defined in Equation[10](https://arxiv.org/html/2607.06214#S5.E10)below\) asRtS=0R\_\{t\}^\{S\}=0by convention\. This index summarizes performance at the level of the whole population and is distinct from the per\-question redundancy scoreRredR\_\{\\mathrm\{red\}\}of Section[4](https://arxiv.org/html/2607.06214#S4)\.

The next step summarizes, in a single number, how effectively a population’s inquiry produces valuable knowledge\. The three quantities are combined into a population\-performance indexRtSR\_\{t\}^\{S\}that is high only when inquiry is simultaneously substantial in volume, spread across topics, and aimed at the frontier, rather than maximizing one of these at the expense of the others\. The index is inspired by a Cobb–Douglas form\(Cobb and Douglas,[1928](https://arxiv.org/html/2607.06214#bib.bib10)\):

RtS=AS​QtaQ​\(ϵD\+Dt\)aD​\(ϵF\+Ft\)aF,R\_\{t\}^\{S\}=A^\{S\}\\,Q\_\{t\}^\{a\_\{Q\}\}\\,\(\\epsilon\_\{D\}\+D\_\{t\}\)^\{a\_\{D\}\}\\,\(\\epsilon\_\{F\}\+F\_\{t\}\)^\{a\_\{F\}\},\(10\)whereAS\>0A^\{S\}\>0is a scale constant \(a total\-factor\-productivity term that sets the units and baseline efficiency of the index\), and the exponentsaQ,aD,aF≥0a\_\{Q\},a\_\{D\},a\_\{F\}\\geq 0set the importance of volume, diversity, and frontier share\. The small floorsϵD,ϵF\>0\\epsilon\_\{D\},\\epsilon\_\{F\}\>0keep the index from collapsing to zero when a population sits on a single topic \(Dt=0D\_\{t\}=0\) or entirely off the frontier \(Ft=0F\_\{t\}=0\), while preserving the complementarity between the three terms\.DtD\_\{t\}andFtF\_\{t\}are separated and can move in opposite directions\. A population can spread inquiry across many topics without reaching the frontier, or concentrate on a few topics that are mostly frontier\. More inquiry only raises performance when the added volume is not outweighed by losses in diversity or frontier\-directed effort\. This follows from Equation[10](https://arxiv.org/html/2607.06214#S5.E10): for positiveQtQ\_\{t\},DtD\_\{t\}, andFtF\_\{t\}, small proportional changes satisfy

Δ​log⁡RtS=aQ​Δ​log⁡Qt\+aD​Δ​log⁡Dt\+aF​Δ​log⁡Ft,\\Delta\\log R\_\{t\}^\{S\}=a\_\{Q\}\\,\\Delta\\log Q\_\{t\}\+a\_\{D\}\\,\\Delta\\log D\_\{t\}\+a\_\{F\}\\,\\Delta\\log F\_\{t\},so a gain in volume \(aQ​Δ​log⁡Qta\_\{Q\}\\,\\Delta\\log Q\_\{t\}\) helps net only if it is not cancelled by drops in the diversity or frontier terms\.

## 6Extensions for multi\-agent AI design

The equations in this section and the one before it can serve as design ideas for multi\-agent AI systems and can be adapted as needed\. If a swarm needs to focus on a single objective, settingaD=0a\_\{D\}=0removes topic diversity fromRtSR\_\{t\}^\{S\}\. Annealing the pursuit rulePi​\(q,t\)P\_\{i\}\(q,t\)over time moves agents from exploration toward exploitation\. Weighting the peer summaryX¯−i\\bar\{X\}\_\{\-i\}more heavily in the drift operator \(Equation[3](https://arxiv.org/html/2607.06214#S3.E3)\) couples agents more tightly, coordinating the population but homogenizing the questions they ask\.

Redundancy can sometimes be useful\. The redundancy weightζi\\zeta\_\{i\}can be made negative in cases in which replication is valuable rather than wasteful\. Re\-deriving an answer already in the stock \(d​\(q,𝒦t\)d\(q,\\mathcal\{K\}\_\{t\}\)\) wastes effort, but several agents pursuing the same open question at once \(∑j≠iPj​\(q,t\)\\sum\_\{j\\neq i\}P\_\{j\}\(q,t\)\) can on occasion help\. Parallel attempts can cross\-check and protect against single agent failure\. The framework captures this by giving the two parts ofRredR\_\{\\mathrm\{red\}\}separate weights, so peer overlap can be rewarded \(ωp<0\\omega\_\{p\}<0\) even while stock overlap is penalized\.

Other extensions require new state variables\. Adding a confidence or reliability term alongsideStS\_\{t\}may be useful for further verification, and if applied to machine intelligence, parallel compute may call for an explicit time\-to\-resolution variable, distinct from the private costCiC\_\{i\}\.

## 7Relation to prior work

The framework is inspired by growth macroeconomics\.Arrow \([1962](https://arxiv.org/html/2607.06214#bib.bib1)\)argued that inventors cannot capture all the rewards related to or produced by their discoveries\. Here, that uncaptured social value appears asκi<1\\kappa\_\{i\}<1inside an agent’s value function\. The shared knowledge stock is related toRomer \([1990](https://arxiv.org/html/2607.06214#bib.bib28)\), where knowledge is a non\-rival input whose growth raises collective returns\.

Weitzman \([1998](https://arxiv.org/html/2607.06214#bib.bib38)\)suggested that discovery depends on finding new combinations of existing ideas\. The current model does not consider recombination\. It counts only reusable, non\-redundant additions to knowledge stock and allows stock to depreciate as newer work supersedes older work through the staleness rateδ\\deltain Equation[7](https://arxiv.org/html/2607.06214#S4.E7)\(cf\. Aghion and Howitt,[1992](https://arxiv.org/html/2607.06214#bib.bib2)\)\.

The generativity term is directly inspired byMurayama \([2022](https://arxiv.org/html/2607.06214#bib.bib25)\), whose key insight is that learning can open new gaps faster than it closes old ones, so acquiring knowledge can make an agent aware of more unanswered questions\. Murayama demonstrates this with a deliberately random model in which knowledge items form a network in which each is linked to a fixed fraction of the others, and the learner acquires them in random order\. Under those assumptions, the number of open gaps rises while the learner knows less than half the items and only later declines\. Differentially, in the model in the current manuscript, generativitygig\_\{i\}is not a uniform property of interchangeable items but varies from question to question\. Also, each agent selects a question according to its own \(drifting\) curiosity policy\.

Finally, the drift component of this framework converges with the effort recalibration framework ofWiradhanyet al\.\([2026](https://arxiv.org/html/2607.06214#bib.bib43)\)\. That work was developed concurrently to address engagement with digital media and its impact on cognition\. Both proposals concentrate on experience\-dependent shifts of decision weights as an explanation for changes in human\-technology interaction\. These are modern, mechanistic forms of a concern voiced in Plato’s*Phaedrus*, in which Socrates warns that writing would reshape human memory and understanding\. Importantly, the two papers differ in scope\.Wiradhanyet al\.\([2026](https://arxiv.org/html/2607.06214#bib.bib43)\)concerns itself with a choice between media engagement and effortful tasks, such that the weight of effort changes with repeated low\-effort experience \(e\.g\., scrolling social media\)\. In the present framework, drift acts on all of the decision weights that govern a curiosity policy, of which effort is only one\. It is embedded, moreover, in a multi\-agent system with a shared knowledge stock, so that individual\-level recalibration has population\-level consequences for inquiry volume, diversity, and frontier share\. Importantly, here, unlike a single\-weight effort account, the same operator can move weights in either direction depending on the ecology\. Therefore, the drift here is not always a downward slide toward low\-effort or a maladaptive behavioral regime, it be can be adaptive or maladaptive in the nature depending on context \(Section[3](https://arxiv.org/html/2607.06214#S3)\)\. A companion paper \(under review at*Trends in Neurosciences*\) extends the framework to neurobiological mechanisms and their broader sociological implications\.

## 8Concluding remarks and limitations

An important empirical question is whether curiosity weights drift in a manner that is captured by the framework\. One could test this in an experiment which manipulates a person’s recent history of inquiry and then measures which questions they choose over the following days\. The drift termΔ​θi\\Delta\\theta\_\{i\}should eventually be tied to reinforcement learning\. This requires identifying which signals update the weights, and over what timescales\.

The shared knowledge stock is a simplified scalar indexStS\_\{t\}, paired with a set𝒦t\\mathcal\{K\}\_\{t\}of deposited results\. But quantities such as the redundancyd​\(q,𝒦t\)d\(q,\\mathcal\{K\}\_\{t\}\)and the reusable valueU​\(q,𝒦t\)U\(q,\\mathcal\{K\}\_\{t\}\)ultimately depend on the specific content of what is known and on how individual pieces of knowledge relate to one another\. Therefore, the framework needs to be expanded \(e\.g\., to represent knowledge in a a graph form\)\.

For all multi\-agent systems, the division of labor and the form of messages passed between agents are key\. Peer coupling may coordinate a group, but it may also homogenize it\. These issues need empirical study: when does social drift help, when does it instead collapse diversity, and how much information should be shared, and in what form, under different regimes?

Cheap resolution to questions may not always be inherently harmful\. If lower cost frees time to pursue the hard questions, frontier share can stay constant or even rise as volume rises\. However, ifτi\\tau\_\{i\}falls, the same freed effort can be redirected toward shallow inquiry\. External rewards may partly oppose the effects of drift: if abundant answers depreciate while rare long\-horizon questions become more valuable, agents that internalize social value throughκi\\kappa\_\{i\}orηi\\eta\_\{i\}may maintain a higher weight over long\-horizon returns\. Also, reusable by\-products of cheap inquiry can lower the cost of harder questions\(Arrow,[1962](https://arxiv.org/html/2607.06214#bib.bib1)\)and add material toStS\_\{t\}for later recombination\(Weitzman,[1998](https://arxiv.org/html/2607.06214#bib.bib38)\)\. In that case, the experience state can record speed and low cost alongside the new gaps and reusable tools the inquiry produced\. Thenτi\\tau\_\{i\}need not fall\. Whether a policyθi\\theta\_\{i\}is adaptive depends on its fit to the ecologyMtM\_\{t\}, not only on a policy by itself\. For example, high vigilance, low tolerance for open questions, or reassurance\-seeking can be useful in some settings, but not in others\.

If curiosity weights do in fact drift, it becomes important to distinguish the*empirical*drift operator, meaning how the preferences of individuals and populations actually move as a function of behavior and ecology, from the*desirable*drift operator, meaning how they would need to move to achieve a given objective, such as expanding useful knowledge\. To capture such macro\-level effects, we need to study how individuals and populations are curious and how their preferences drift, and the structure of the surrounding ecology itself, which describes which questions and answers are available, what they cost, and how answers propagate through a population\.

## Declarations

- •Funding:NIH/NIMH R01MH128344, R01MH110594, R01MH116937, and Conte Center MH10643\.
- •Conflict of interest:The author declares no competing interests\.
- •Materials availability:Not applicable\.
- •AI usage:AI tools were used in preparation of the manuscript but the author takes full responsibility for the statements made herein\.
- •The author would like to acknowledge Drs\. Gaia Tavoni, Ethan Bromberg\-Martin, Domenico Giannone, Bruno Averbeck, Michael Frank, Catherine Hartley, and Binxu Wang for many helpful comments on this article\.

## References

- A model of growth through creative destruction\.Econometrica60,pp\. 323–351\.External Links:[Document](https://dx.doi.org/10.2307/2951599)Cited by:[§7](https://arxiv.org/html/2607.06214#S7.p2.1)\.
- K\. J\. Arrow \(1962\)Economic welfare and the allocation of resources for invention\.InThe Rate and Direction of Inventive Activity,R\. R\. Nelson \(Ed\.\),pp\. 609–626\.External Links:[Document](https://dx.doi.org/10.1515/9781400879762-024)Cited by:[§5](https://arxiv.org/html/2607.06214#S5.p1.5),[§7](https://arxiv.org/html/2607.06214#S7.p1.1),[§8](https://arxiv.org/html/2607.06214#S8.p4.7)\.
- B\. B\. Averbeck \(2015\)Theory of choice in bandit, information sampling and foraging tasks\.PLoS Computational Biology11\(3\),pp\. e1004164\.External Links:[Document](https://dx.doi.org/10.1371/journal.pcbi.1004164)Cited by:[§4](https://arxiv.org/html/2607.06214#S4.p1.1)\.
- G\. S\. Becker and K\. M\. Murphy \(1988\)A theory of rational addiction\.Journal of Political Economy96,pp\. 675–700\.External Links:[Document](https://dx.doi.org/10.1086/261558)Cited by:[§3](https://arxiv.org/html/2607.06214#S3.p2.3)\.
- E\. S\. Bromberg\-Martin and I\. E\. Monosov \(2020\)Neural circuitry of information seeking\.Current Opinion in Behavioral Sciences35,pp\. 62–70\.External Links:[Document](https://dx.doi.org/10.1016/j.cobeha.2020.07.006)Cited by:[§1](https://arxiv.org/html/2607.06214#S1.p1.1)\.
- C\. J\. Charpentier, E\. S\. Bromberg\-Martin, and T\. Sharot \(2018\)Valuation of knowledge and ignorance in mesolimbic reward circuitry\.Proceedings of the National Academy of Sciences USA115,pp\. E7255–E7264\.External Links:[Document](https://dx.doi.org/10.1073/pnas.1800547115)Cited by:[§2](https://arxiv.org/html/2607.06214#S2.p1.10)\.
- C\. W\. Cobb and P\. H\. Douglas \(1928\)A theory of production\.American Economic Review18,pp\. 139–165\.Cited by:[§5](https://arxiv.org/html/2607.06214#S5.p5.1)\.
- V\. D\. Costa and B\. B\. Averbeck \(2020\)Primate orbitofrontal cortex codes information relevant for managing explore–exploit tradeoffs\.Journal of Neuroscience40\(12\),pp\. 2553–2561\.External Links:[Document](https://dx.doi.org/10.1523/JNEUROSCI.2355-19.2020)Cited by:[§4](https://arxiv.org/html/2607.06214#S4.p1.1)\.
- A\. K\. Dixit and R\. S\. Pindyck \(1994\)Investment under uncertainty\.Princeton University Press\.External Links:[Document](https://dx.doi.org/10.2307/j.ctt7sncv)Cited by:[§2](https://arxiv.org/html/2607.06214#S2.p1.10)\.
- A\. R\. Doshi and O\. P\. Hauser \(2024\)Generative AI enhances individual creativity but reduces the collective diversity of novel content\.Science Advances10,pp\. eadn5290\.External Links:[Document](https://dx.doi.org/10.1126/sciadv.adn5290)Cited by:[§4](https://arxiv.org/html/2607.06214#S4.p7.6)\.
- J\. G\. Foster, A\. Rzhetsky, and J\. A\. Evans \(2015\)Tradition and innovation in scientists’ research strategies\.American Sociological Review80,pp\. 875–908\.External Links:[Document](https://dx.doi.org/10.1177/0003122415601618)Cited by:[§5](https://arxiv.org/html/2607.06214#S5.p4.6)\.
- G\. Gigerenzer and R\. Garcia\-Retamero \(2017\)Cassandra’s regret: the psychology of not wanting to know\.Psychological Review124,pp\. 179–196\.External Links:[Document](https://dx.doi.org/10.1037/rev0000055)Cited by:[§2](https://arxiv.org/html/2607.06214#S2.p1.10)\.
- R\. Golman, D\. Hagmann, and G\. Loewenstein \(2017\)Information avoidance\.Journal of Economic Literature55,pp\. 96–135\.External Links:[Document](https://dx.doi.org/10.1257/jel.20151245)Cited by:[§2](https://arxiv.org/html/2607.06214#S2.p1.10)\.
- A\. Jezzini, E\. S\. Bromberg\-Martin, L\. R\. Trambaiolli, S\. N\. Haber, and I\. E\. Monosov \(2021\)A prefrontal network integrates preferences for advance information about uncertain rewards and punishments\.Neuron109\(14\),pp\. 2339–2352\.External Links:[Document](https://dx.doi.org/10.1016/j.neuron.2021.05.013)Cited by:[§2](https://arxiv.org/html/2607.06214#S2.p1.10)\.
- G\. Loewenstein \(1994\)The psychology of curiosity: a review and reinterpretation\.Psychological Bulletin116,pp\. 75–98\.External Links:[Document](https://dx.doi.org/10.1037/0033-2909.116.1.75)Cited by:[§2](https://arxiv.org/html/2607.06214#S2.p4.6)\.
- J\. G\. March \(1991\)Exploration and exploitation in organizational learning\.Organization Science2,pp\. 71–87\.External Links:[Document](https://dx.doi.org/10.1287/orsc.2.1.71)Cited by:[§4](https://arxiv.org/html/2607.06214#S4.p1.1)\.
- R\. McDonald and D\. Siegel \(1986\)The value of waiting to invest\.Quarterly Journal of Economics101,pp\. 707–727\.External Links:[Document](https://dx.doi.org/10.2307/1884175)Cited by:[§2](https://arxiv.org/html/2607.06214#S2.p1.10)\.
- I\. E\. Monosov \(2024\)Curiosity: primate neural circuits for novelty and information seeking\.Nature Reviews Neuroscience25,pp\. 195–208\.External Links:[Document](https://dx.doi.org/10.1038/s41583-023-00784-9)Cited by:[§1](https://arxiv.org/html/2607.06214#S1.p1.1)\.
- K\. Murayama \(2022\)A reward\-learning framework of knowledge acquisition: an integrated account of curiosity, interest, and intrinsic\-extrinsic rewards\.Psychological Review129\(1\),pp\. 175–198\.External Links:[Document](https://dx.doi.org/10.1037/rev0000349)Cited by:[§2](https://arxiv.org/html/2607.06214#S2.p4.6),[§7](https://arxiv.org/html/2607.06214#S7.p3.1)\.
- R\. R\. Nelson and S\. G\. Winter \(1982\)An evolutionary theory of economic change\.Harvard University Press\.Cited by:[§4](https://arxiv.org/html/2607.06214#S4.p1.1)\.
- V\. Padmakumar and H\. He \(2024\)Does writing with language models reduce content diversity?\.InInternational Conference on Learning Representations \(ICLR\),External Links:2309\.05196,[Link](https://arxiv.org/abs/2309.05196)Cited by:[§4](https://arxiv.org/html/2607.06214#S4.p7.6)\.
- P\. M\. Romer \(1990\)Endogenous technological change\.Journal of Political Economy98,pp\. S71–S102\.External Links:[Document](https://dx.doi.org/10.1086/261725)Cited by:[§7](https://arxiv.org/html/2607.06214#S7.p1.1)\.
- H\. E\. Ryder and G\. M\. Heal \(1973\)Optimal growth with intertemporally dependent preferences\.Review of Economic Studies40,pp\. 1–31\.External Links:[Document](https://dx.doi.org/10.2307/2296736)Cited by:[§3](https://arxiv.org/html/2607.06214#S3.p2.3)\.
- W\. Schultz, P\. Dayan, and P\. R\. Montague \(1997\)A neural substrate of prediction and reward\.Science275,pp\. 1593–1599\.External Links:[Document](https://dx.doi.org/10.1126/science.275.5306.1593)Cited by:[§3](https://arxiv.org/html/2607.06214#S3.p2.11)\.
- C\. E\. Shannon \(1948\)A mathematical theory of communication\.Bell System Technical Journal27,pp\. 379–423, 623–656\.External Links:[Document](https://dx.doi.org/10.1002/j.1538-7305.1948.tb00917.x)Cited by:[§5](https://arxiv.org/html/2607.06214#S5.p3.11)\.
- I\. Shumailov, Z\. Shumaylov, Y\. Zhao, N\. Papernot, R\. Anderson, and Y\. Gal \(2024\)AI models collapse when trained on recursively generated data\.Nature631,pp\. 755–759\.External Links:[Document](https://dx.doi.org/10.1038/s41586-024-07566-y)Cited by:[§4](https://arxiv.org/html/2607.06214#S4.p7.6)\.
- G\. J\. Stigler and G\. S\. Becker \(1977\)De gustibus non est disputandum\.American Economic Review67,pp\. 76–90\.Cited by:[§3](https://arxiv.org/html/2607.06214#S3.p8.1)\.
- R\. S\. Sutton and A\. G\. Barto \(2018\)Reinforcement learning: an introduction\.2nd edition,MIT Press\.External Links:[Link](http://incompleteideas.net/book/the-book-2nd.html)Cited by:[§3](https://arxiv.org/html/2607.06214#S3.p2.11)\.
- A\. Tversky and I\. Simonson \(1993\)Context\-dependent preferences\.Management Science39\(10\),pp\. 1179–1189\.External Links:[Document](https://dx.doi.org/10.1287/mnsc.39.10.1179)Cited by:[§1](https://arxiv.org/html/2607.06214#S1.p2.1)\.
- B\. Uzzi, S\. Mukherjee, M\. Stringer, and B\. Jones \(2013\)Atypical combinations and scientific impact\.Science342\(6157\),pp\. 468–472\.External Links:[Document](https://dx.doi.org/10.1126/science.1240474)Cited by:[§5](https://arxiv.org/html/2607.06214#S5.p4.6)\.
- M\. L\. Weitzman \(1998\)Recombinant growth\.Quarterly Journal of Economics113\(2\),pp\. 331–360\.External Links:[Document](https://dx.doi.org/10.1162/003355398555595)Cited by:[§7](https://arxiv.org/html/2607.06214#S7.p2.1),[§8](https://arxiv.org/html/2607.06214#S8.p4.7)\.
- W\. Wiradhany, D\. Parry, and J\. Aru \(2026\)An effort recalibration framework for digital media use and cognition\.Nature Human Behaviour\.External Links:[Document](https://dx.doi.org/10.1038/s41562-026-02500-w)Cited by:[§7](https://arxiv.org/html/2607.06214#S7.p4.1)\.

Similar Articles

Economy of Minds: Emerging Multi-Agent Intelligence with Economic Interactions

Hugging Face Daily Papers

This paper proposes an 'agent economy' framework inspired by Hayek's economic theory, where agents self-organize through auction-based competition and economic selection to produce emergent multi-step reasoning and collective intelligence without centralized control. The system outperforms stronger monolithic baselines across five agentic tasks including mathematical reasoning, financial research, and scientific research.

Large-scale study of curiosity-driven learning

OpenAI Blog

OpenAI presents a large-scale empirical study of curiosity-driven reinforcement learning without extrinsic rewards across 54 benchmark environments, showing strong performance and investigating the role of feature spaces in prediction-based reward signals.