Endogenous Exploration in Reinforcement Learning with Intrinsic Curiosity

arXiv cs.LG Papers

Summary

This paper proposes a reinforcement learning framework using intrinsic curiosity for endogenous exploration in non-stationary environments, achieving competitive performance on benchmarks like LunarLander-v2 and BipedalWalker-v3 compared to algorithms such as PPO and ICM.

arXiv:2609.05650v1 Announce Type: new Abstract: We propose a reinforcement learning framework in which exploration is driven by intrinsic curiosity, designed for scenarios where environments are non-stationary and rewards are sparse, delayed, uninformative, or absent. In our model, action selection is guided by a combination of external rewards and an epistemic motivation mechanism that biases the agent toward structured exploratory directions. The central hypothesis is that effective exploration emerges at intermediate levels of incoherence, while performance degrades under both overly rigid and overly disordered dynamics. To test this idea, we implement the framework on top of a Liquid State Machine (LSM) substrate and evaluate it on two standard benchmarks: the discrete-action LunarLanderv2 and the continuous-control BipedalWalkerv3. The proposed method achieves competitive performance on both tasks relative to established deep RL algorithms, including Proximal Policy Optimization (PPO) and Intrinsic Curiosity Module (ICM). We further show that the curiosity window is not recovered in Active Inference agents under the same analysis, suggesting that the proposed dynamics capture a distinct exploration regime
Original Article
View Cached Full Text

Cached at: 09/10/26, 08:22 AM

# Endogenous Exploration in Reinforcement Learning with Intrinsic Curiosity
Source: [https://arxiv.org/html/2609.05650](https://arxiv.org/html/2609.05650)
April 2026

###### Abstract

We propose a reinforcement learning framework in which exploration is driven by intrinsic curiosity, designed for scenarios where environments are non\-stationary and rewards are sparse, delayed, uninformative, or absent\. In our model, action selection is guided by a combination of external rewards and an epistemic motivation mechanism that biases the agent toward structured exploratory directions\. The central hypothesis is that effective exploration emerges at intermediate levels of incoherence, while performance degrades under both overly rigid and overly disordered dynamics\. To test this idea, we implement the framework on top of a Liquid State Machine \(LSM\) substrate and evaluate it on two standard benchmarks—the discrete\-actionLunarLander\-v2and the continuous\-controlBipedalWalker\-v3\. The proposed method achieves competitive performance on both tasks relative to established deep RL algorithms, including Proximal Policy Optimization \(PPO\) and Intrinsic Curiosity Module \(ICM\)\. We further show that the curiosity window is not recovered in Active Inference agents under the same analysis, suggesting that the proposed dynamics capture a distinct exploration regime\.

Keywords:curiosity window, liquid state machines, active inference, intrinsic motivation, reservoir computing, exploration–exploitation

## 1Introduction

The problem of exploration in adaptive systems—how an agent discovers useful behaviours without exhaustive search—remains central to reinforcement learning \(RL\), cognitive science, and theories of consciousness\. Standard approaches treat exploration as an external mechanism:ϵ\\epsilon\-greedy schedules, Boltzmann temperature annealing, count\-based bonuses, or prediction\-error curiosity signals\([Pathak et al\., 2017](https://arxiv.org/html/2609.05650#bib.bib6);[Burda et al\., 2018](https://arxiv.org/html/2609.05650#bib.bib1)\)\. In each case, the drive to explore is*designed in*rather than emerging from the system’s own dynamics\.

Despite significant progress, most modern approaches still rely on exogenous mechanisms such asϵ\\epsilon\-greedy noise, entropy regularisation, or intrinsic reward proxies based on prediction error or state novelty\. While these methods improve exploration in specific settings, they do not provide a general solution: they require careful tuning, often fail under non\-stationarity, and tend to collapse either into premature exploitation or unstructured random behaviour\. More recent advances—including Soft Actor\-Critic \(SAC\), Twin Delayed DDPG \(TD3\), and curiosity\-driven methods such as ICM\([Pathak et al\., 2017](https://arxiv.org/html/2609.05650#bib.bib6)\)and RND\([Burda et al\., 2018](https://arxiv.org/html/2609.05650#bib.bib1)\)—improve stability and performance but still treat exploration as an auxiliary objective rather than as an endogenous consequence of the agent’s internal dynamics\. Consequently, they struggle in regimes where reward is sparse, absent, or shifting, and they lack a principled account of when exploration should increase or decrease\. This gap suggests that the core issue is conceptual: current frameworks do not capture the conditions under which meaningful, structured exploration should emerge\.

We propose a different approach\. Instead of optimising a single external objective, we posit that systems self\-regulate*coherently*between functionally differentiated fields\. Exploration arises endogenously when incoherence reaches an intermediate regime—neither so low that the system is trapped, nor so high that it fragments\. This prediction, termed the*Curiosity Window conjecture*, is the framework’s most distinctive and testable claim\.

In this paper, we provide the first computational validation of these ideas, implemented through a Liquid State Machine \(LSM\) substrate\. We demonstrate that systems self\-organise around a curiosity window, compare our agents against Active Inference \(AIF\) baselines, and test whether endogenous noise modulation can sustain exploration without external signals\. Our contributions are:

1. 1\.Curiosity Window: TheC2C\_\{2\}metric \(coherent basins×\\timesglobal overlap×\\timestransition rate\) peaks at intermediate noise, with\>1\.5×\>1\.5\\timesdrop\-off on both sides \(Section[4](https://arxiv.org/html/2609.05650#S4)\)\.
2. 2\.Endogenous exploration: The curiosity window is specific to our model—AIF agents show no analogous peak \(Section[5\.3](https://arxiv.org/html/2609.05650#S5.SS3)\)\.
3. 3\.Extended dynamics: Spatial coupling, sedimentation learning, and non\-stationary adaptation are validated \(Section[5\.4](https://arxiv.org/html/2609.05650#S5.SS4)\)\.

## 2Summary of the Coherence Framework

We posit that experience is not a passive reception of data but an active regulation of coherence between what is*open*\(possible continuations\) and what is*constrained*\(current situational demands\)\. Unlike the Free Energy Principle \(FEP\), which derives behaviour from surprise minimisation over generative models, we treat coherence regulation as primitive and representational structures as secondary invariances that may or may not emerge\.

Three principles govern the framework:

1. 1\.Coherence Primacy: The primitive explanatory target is the regulation of coherence between reach and yield\. Representational structures are secondary\.
2. 2\.Functional Differentiation: Openness and constraint must be functionally differentiated; undifferentiated dynamics cannot express the framework’s characteristic tension\.
3. 3\.Mutual Constraint: Reach must be constrained by yield, yield must matter for reach, and memory must affect the trajectory of both\.

### 2\.1Mathematical Setting

Letℰ\\mathcal\{E\}denote the experiential field, a measurable space equipped with aσ\\sigma\-algebraℱ\\mathcal\{F\}and a reference measureμ\\mu\. Three time\-dependent densities are defined onℰ\\mathcal\{E\}:

###### Definition 1\(Functional Roles\)\.

- •Reachπt∈𝒫⁡\(ℰ\)\\pi\_\{t\}\\in\\mathcal\{P\}\(\\mathcal\{E\}\): the distribution over viable continuations \(openness, action tendency\)\.
- •Yieldyt∈𝒫⁡\(ℰ\)y\_\{t\}\\in\\mathcal\{P\}\(\\mathcal\{E\}\): the distribution reflecting current environmental and internal constraint\.
- •Memorymt∈𝒫⁡\(ℰ\)m\_\{t\}\\in\\mathcal\{P\}\(\\mathcal\{E\}\): the sedimented trace of past coherence achievements\.

Two functionals measure the system’s coherence state:

###### Definition 2\(Incoherence\)\.

I\(t\)=DKL\(πt∥yt\)I\(t\)=D\_\{\\mathrm\{KL\}\}\(\\pi\_\{t\}\\\|y\_\{t\}\)\(1\)

###### Definition 3\(Global Overlap\)\.

G⁡\(t\)=∫ℰπt​\(e\)⋅yt​\(e\)​𝑑μ​\(e\)G\(t\)=\\int\_\{\\mathcal\{E\}\}\\sqrt\{\\pi\_\{t\}\(e\)\\cdot y\_\{t\}\(e\)\}\\,d\\mu\(e\)\(2\)

The global overlapG⁡\(t\)G\(t\)is the Bhattacharyya coefficient between reach and yield; it equals 1 whenπt=yt\\pi\_\{t\}=y\_\{t\}and approaches 0 when the two distributions have disjoint support\. In the discrete implementations that follow, the integral is replaced by a finite sum over field cells\.

The dynamics are governed by coupled update rules:

πt\+1\\displaystyle\\pi\_\{t\+1\}∝πt⋅exp⁡\(−ηπ​yt\)\+σπ​ξt\\displaystyle\\propto\\pi\_\{t\}\\cdot\\exp\\bigl\(\-\\eta\_\{\\pi\}\\,y\_\{t\}\\bigr\)\+\\sigma\_\{\\pi\}\\xi\_\{t\}\(3\)yt\+1\\displaystyle y\_\{t\+1\}=\(1−ηy\)​yt\+ηy​Φy​\(rt,ut\)\\displaystyle=\(1\-\\eta\_\{y\}\)y\_\{t\}\+\\eta\_\{y\}\\,\\Phi\_\{y\}\(r\_\{t\},u\_\{t\}\)\(4\)mt\+1\\displaystyle m\_\{t\+1\}=\(1−λ\)​mt\+λ​πt\\displaystyle=\(1\-\\lambda\)m\_\{t\}\+\\lambda\\,\\pi\_\{t\}\(5\)whereηπ,ηy,λ\\eta\_\{\\pi\},\\eta\_\{y\},\\lambdaare learning rates,Φy\\Phi\_\{y\}maps reservoir state and input to yield,σπ\\sigma\_\{\\pi\}is the noise scale, andξt\\xi\_\{t\}is i\.i\.d\. noise\.

###### Conjecture 1\(Curiosity Window\)\.

There exists an intermediate regime of incoherenceI∗∈\(Imin,Imax\)I^\{\*\}\\in\(I\_\{\\min\},I\_\{\\max\}\)such that a suitable curiosity metricC⁡\(t\)C\(t\)is maximised\. BelowI∗I^\{\*\}, the system is trapped in coherent but unexplorative basins; aboveI∗I^\{\*\}, the system fragments and loses coherent structure\.

The Curiosity Window proof \(Appendix[A](https://arxiv.org/html/2609.05650#A1)\) relies on a*curiosity functional*C^2\\hat\{C\}\_\{2\}whose specific form—coherent basin occupancy times global overlap times transition rate—was introduced computationally\. The key insight is that incoherenceI\(t\)=DKL\(πt∥yt\)I\(t\)=D\_\{\\mathrm\{KL\}\}\(\\pi\_\{t\}\\\|y\_\{t\}\), measuring reach–yield tension, and global overlapG⁡\(t\)=∫ℰπt​yt​𝑑μG\(t\)=\\int\_\{\\mathcal\{E\}\}\\sqrt\{\\pi\_\{t\}\\,y\_\{t\}\}\\,d\\mu, measuring system\-level alignment \(Definitions[2](https://arxiv.org/html/2609.05650#Thmdefinition2)and[3](https://arxiv.org/html/2609.05650#Thmdefinition3)\), together capture the essential ingredients of*curiosity*: the endogenous opening of new viable possibilities driven by unresolved internal tension\. This requires:

1. 1\.that the system*visits distinct coherent configurations*\(not merely fluctuates randomly\);
2. 2\.that it does so while*maintaining global integration*\(not fragmenting\);
3. 3\.that the exploration is*structured*—transitions between identifiable basins, not diffusion through undifferentiated state space\.

### 2\.2Representation Theorem

###### Theorem 1\(Representation of the curiosity functional\)\.

Let𝒞:ℝ≥03→ℝ≥0\\mathcal\{C\}:\\mathbb\{R\}\_\{\\geq 0\}^\{3\}\\to\\mathbb\{R\}\_\{\\geq 0\}be a curiosity functional of the form𝒞=F⁡\(G,ℬ,𝒯\)\\mathcal\{C\}=F\(G,\\mathcal\{B\},\\mathcal\{T\}\)whereFFis continuous and separately monotone in each argument \(increasing inGG,ℬ\\mathcal\{B\}, and𝒯\\mathcal\{T\}\)\. Then:

1. 1\.FFvanishes on the boundary:F⁡\(G,ℬ,0\)=F⁡\(G,0,𝒯\)=F⁡\(0,ℬ,𝒯\)=0F\(G,\\mathcal\{B\},0\)=F\(G,0,\\mathcal\{T\}\)=F\(0,\\mathcal\{B\},\\mathcal\{T\}\)=0for all values of the remaining arguments below their respective thresholds\.
2. 2\.FFis uniquely determined up to monotone transformation by the product form: there exists a strictly increasingφ:ℝ≥0→ℝ≥0\\varphi:\\mathbb\{R\}\_\{\\geq 0\}\\to\\mathbb\{R\}\_\{\\geq 0\}withφ⁡\(0\)=0\\varphi\(0\)=0such that 𝒞=φ⁡\(G⋅ℬ⋅𝒯\)\.\\mathcal\{C\}=\\varphi\\bigl\(G\\cdot\\mathcal\{B\}\\cdot\\mathcal\{T\}\\bigr\)\.\(6\)

Proof is presented in Appendix[B](https://arxiv.org/html/2609.05650#A2)\.

### 2\.3Sedimentation as Learning

We treat learning not as parameter updates to a loss function but as*sedimentation*: the progressive shaping of the memory fieldmtm\_\{t\}through repeated coherence episodes\. States that have repeatedly achieved low incoherence acquire higher probability mass in the baseline reachπ0\\pi\_\{0\}, making future coherence in those regions easier\. This is formalised as:

π0\(T\+1\)∝π0\(T\)⋅exp\(−α∑τ=1TωτIτ\)\\pi\_\{0\}^\{\(T\+1\)\}\\propto\\pi\_\{0\}^\{\(T\)\}\\cdot\\exp\\Bigl\(\-\\alpha\\sum\_\{\\tau=1\}^\{T\}\\omega\_\{\\tau\}\\,I\_\{\\tau\}\\Bigr\)\(7\)whereα\\alphais the sedimentation rate andωτ\\omega\_\{\\tau\}are recency weights\.

## 3Computational Architecture and Implementation

Our computational realisation is structured as a two\-layer architecture: \(i\) a dynamical substrate providing high\-dimensional temporal representations, and \(ii\) a coherence\-regulating control layer implementing the field dynamics\.

The substrate is instantiated as a Liquid State Machine \(LSM\), while the coherence layer defines the evolution of the reach \(π\\pi\), yield \(yy\), and memory \(mm\) fields, together with the endogenous modulation of exploration\. This separation is essential: the LSM supplies a rich, fading\-memory embedding of experience, whereas the coherence layer provides the governing principles of organisation and exploration\.

### 3\.1Liquid State Machine Substrate

We model the experiential fieldℰ\\mathcal\{E\}as the state space induced by a recurrent reservoir\. The reservoir statex⁡\(t\)∈ℝNx\(t\)\\in\\mathbb\{R\}^\{N\}evolves according to:

x⁡\(t\+1\)=tanh⁡\(W​x​\(t\)\+Win​o​\(t\)\+ξ⁡\(t\)\),x\(t\+1\)=\\tanh\\big\(Wx\(t\)\+W\_\{\\text\{in\}\}o\(t\)\+\\xi\(t\)\\big\),\(8\)
whereW∈ℝN×NW\\in\\mathbb\{R\}^\{N\\times N\}is a sparse recurrent weight matrix with spectral radiusρ<1\\rho<1,WinW\_\{\\text\{in\}\}maps sensory inputo⁡\(t\)o\(t\)into the reservoir, andξ⁡\(t\)∼𝒩⁡\(0,σI2\)\\xi\(t\)\\sim\\mathcal\{N\}\(0,\\sigma\_\{I\}^\{2\}\)represents intrinsic perturbations\.

The LSM provides:

- •high\-dimensional nonlinear expansion of inputs,
- •fading memory of past states,
- •continuous\-time\-like dynamics suitable for temporal integration\.

Crucially, the reservoir itself is not the agent: it serves as a dynamical medium over which the coherence variables are defined and updated\.

### 3\.2State Augmentation and Temporal Thickness

To capture temporal extension explicitly, we augment the instantaneous reservoir state with a slow trace:

ht=\(1−β\)​ht−1\+β​x​\(t\),h\_\{t\}=\(1\-\\beta\)h\_\{t\-1\}\+\\beta x\(t\),\(9\)
and define the effective substrate representation as:

zt=\[x⁡\(t\),ht\]∈ℝ2​N\.z\_\{t\}=\[x\(t\),\\;h\_\{t\}\]\\in\\mathbb\{R\}^\{2N\}\.\(10\)
This construction provides both fast dynamics \(xx\) and sedimented temporal structure \(hh\), aligning the implementation with the requirement of temporally extended coherence\.

### 3\.3Field Representation and Extraction

The fields are defined as probability distributions over a discretised version of the reservoir state space\. Givenztz\_\{t\}, we construct:

πt\\displaystyle\\pi\_\{t\}=softmax​\(fπ​\(zt\)\),\\displaystyle=\\text\{softmax\}\(f\_\{\\pi\}\(z\_\{t\}\)\),\(11\)yt\\displaystyle y\_\{t\}=softmax​\(Wr​zt\+Wπ​πt\+Wm​mt\),\\displaystyle=\\text\{softmax\}\(W\_\{r\}z\_\{t\}\+W\_\{\\pi\}\\pi\_\{t\}\+W\_\{m\}m\_\{t\}\),\(12\)mt\\displaystyle m\_\{t\}∈𝒫⁡\(ℰ\),\\displaystyle\\in\\mathcal\{P\}\(\\mathcal\{E\}\),\(13\)
wherefπf\_\{\\pi\}is typically the identity or a linear projection, andWr,Wπ,WmW\_\{r\},W\_\{\\pi\},W\_\{m\}are learned or fixed mappings\.

The three fields play distinct roles:

- •πt\\pi\_\{t\}: reach \(action tendency / exploratory distribution\),
- •yty\_\{t\}: yield \(constraint induced by environment and internal state\),
- •mtm\_\{t\}: memory \(sedimented trace of past coherent states\)\.

### 3\.4Field Dynamics

The coupled dynamics of the fields follow:

πt\+1\\displaystyle\\pi\_\{t\+1\}∝πt⋅exp⁡\(−ηπ​yt\)\+σπ​\(t\)​ξt,\\displaystyle\\propto\\pi\_\{t\}\\cdot\\exp\(\-\\eta\_\{\\pi\}y\_\{t\}\)\+\\sigma\_\{\\pi\}\(t\)\\,\\xi\_\{t\},\(14\)yt\+1\\displaystyle y\_\{t\+1\}=\(1−ηy\)​yt\+ηy​Φy​\(zt,πt,mt\),\\displaystyle=\(1\-\\eta\_\{y\}\)y\_\{t\}\+\\eta\_\{y\}\\Phi\_\{y\}\(z\_\{t\},\\pi\_\{t\},m\_\{t\}\),\(15\)mt\+1\\displaystyle m\_\{t\+1\}=\(1−λ\)​mt\+λ​πt\.\\displaystyle=\(1\-\\lambda\)m\_\{t\}\+\\lambda\\pi\_\{t\}\.\(16\)
These equations implement:

- •Mutual constraint: reach is shaped by yield, and yield depends on reach,
- •Temporal integration: memory accumulates past reach states,
- •Non\-equilibrium dynamics: the system continuously reconfigures rather than converging to a static optimum\.

The variablemmshould not be interpreted as memory in the classical RL sense \(e\.g\., a replay buffer or explicit storage of past transitions\)\. Instead,mmrepresents a*sedimented internal trace*: a continuously updated latent summary of the agent’s recent and recurrent interactions with the environment\. It evolves as a slow\-moving average of the reach distributionπt\\pi\_\{t\}, capturing what has become statistically stable for the agent over time\. Deviations betweenyty\_\{t\}andmtm\_\{t\}signal novelty or drift, while alignment indicates coherence and stability\. In this sense,mmfunctions as a dynamic baseline of “what is normal,” enabling regulation of exploration without requiring episodic recall\.

### 3\.5Endogenous Exploration Mechanism

A central feature of the architecture is the replacement of exogenous noise with endogenous modulation:

σπ​\(t\)=ψ⁡\(I⁡\(t\)\)⋅G⁡\(t\),\\sigma\_\{\\pi\}\(t\)=\\psi\\big\(I\(t\)\\big\)\\cdot G\(t\),\(17\)
where

ψ\(I\)=AIe−I/I0\.\\psi\(I\)=A\\,I\\,e^\{\-I/I\_\{0\}\}\.\(18\)
This function is unimodal inII, ensuring:

- •low noise under high coherence \(exploitation\),
- •low noise under extreme incoherence \(fragmentation\),
- •maximal exploration at intermediate incoherence \(curiosity window\)\.

This mechanism closes the loop between internal state and exploration, making curiosity an emergent property rather than an externally imposed signal\.

### 3\.6Spatial Structure and Basin Dynamics

The discretised field is partitioned intoKKbasins\{Bk\}k=1K\\\{B\_\{k\}\\\}\_\{k=1\}^\{K\}representing metastable regions of coherent organisation\. Basin assignment at timettis determined by dominant probability mass:

k⁡\(t\)=arg⁡max⁡∑e∈Bkk⁡πt​\(e\)\.k\(t\)=\\arg\\max\_\{k\}\\sum\_\{e\\in B\_\{k\}\}\\pi\_\{t\}\(e\)\.\(19\)
To enforce global integration, a spatial coupling kernel𝐊\\mathbf\{K\}can be applied:

πt\+1​\(i\)←πt\+1​\(i\)\+γ​∑jK⁡\(i,j\)​πt\+1​\(j\)\.\\pi\_\{t\+1\}\(i\)\\leftarrow\\pi\_\{t\+1\}\(i\)\+\\gamma\\sum\_\{j\}K\(i,j\)\\pi\_\{t\+1\}\(j\)\.\(20\)
This allows local perturbations to propagate across the field, preventing fragmentation and supporting coherent large\-scale dynamics\.

The complete system can be summarised as:

- •LSM substrate: provides high\-dimensional, temporally rich state representation,
- •Coherence fields: encode openness \(π\\pi\), constraint \(yy\), and memory \(mm\),
- •Coupled dynamics: enforce mutual constraint and temporal integration,
- •Endogenous noise: links incoherence to exploration intensity,
- •Sedimentation: implements learning as structural bias\.

This architecture differs fundamentally from standard RL systems: exploration is not injected but generated by the system’s own coherence dynamics, yielding a closed\-loop, self\-regulating process\.

## 4The Curiosity Window

The notion of a curiosity window arises from a central limitation in existing exploration strategies: they lack a principled account of*when*exploration should occur\. In most RL frameworks, exploration is either externally imposed \(e\.g\., fixed noise, entropy bonuses\) or monotonically controlled, leading to two well\-known failure modes: premature convergence or unstructured randomness\.

Our framework predicts that effective exploration self\-organises within a bounded intermediate regime\. When incoherence is too low, the system becomes overly stable and is trapped in a limited set of behaviours\. When incoherence is too high, structure is lost and exploration becomes fragmented and unproductive\. Between these extremes lies the*curiosity window*: a regime in which the system explores multiple possibilities while maintaining enough internal coherence to make that exploration meaningful\.

###### Definition 4\(Curiosity functionalC2C\_\{2\}\)\.

Let\(πt,yt,mt\)t≥0\(\\pi\_\{t\},y\_\{t\},m\_\{t\}\)\_\{t\\geq 0\}be a trajectory of the coherence dynamic on\(ℰ,ℬ⁡\(ℰ\),μ\)\(\\mathcal\{E\},\\mathcal\{B\}\(\\mathcal\{E\}\),\\mu\)with basin partition\{Bk\}k=1K\\\{B\_\{k\}\\\}\_\{k=1\}^\{K\}\. The*curiosity functional*is

C2​\(t\)=G⁡\(t\)⏟global overlap⋅ℬ⁡\(t\)⏟coherent basin count⋅𝒯⁡\(t\)⏟inter\-basin transition rate\\boxed\{\\;C\_\{2\}\(t\)\\;=\\;\\underbrace\{G\(t\)\}\_\{\\text\{global overlap\}\}\\;\\cdot\\;\\underbrace\{\\mathcal\{B\}\(t\)\}\_\{\\text\{coherent basin count\}\}\\;\\cdot\\;\\underbrace\{\\mathcal\{T\}\(t\)\}\_\{\\text\{inter\-basin transition rate\}\}\\;\}\(21\)

The three factors ofC2C\_\{2\}correspond to three distinct commitments:

- •G⁡\(t\)G\(t\): Global overlap\.Ensures that the field remains unified\. WithoutGG, a fragmented system that randomly visits many regions would score high on curiosity—violating the requirement that curiosity occurs*within*a coherent field\.
- •ℬ⁡\(t\)\\mathcal\{B\}\(t\): Coherent basin count\.Ensures that the system has differentiated metastable structure\. A system with only one basin cannot exhibit the structured transitions needed for learning and creativity\.
- •𝒯⁡\(t\)\\mathcal\{T\}\(t\): Transition rate\.Ensures that the system is actively*moving*between basins, not merely possessing the potential to do so\. This is the dynamical signature of intrinsic curiosity: the system reorganises because internal tension drives it to explore\.

The curiosity functionalC2=G⋅ℬ⋅𝒯C\_\{2\}=G\\cdot\\mathcal\{B\}\\cdot\\mathcal\{T\}is not an*ad hoc*choice\. By Theorem[1](https://arxiv.org/html/2609.05650#Thmtheorem1), it is the unique \(up to monotone transformation\) continuous, separately monotone, dimensionally consistent functional of the three quantities—global overlap, coherent basin count, and inter\-basin transition rate—that satisfies the boundary conditions derived from the framework’s foundational commitments\. Any system that scores high onC2C\_\{2\}is simultaneously integrated \(GGhigh\), differentiated \(ℬ\>1\\mathcal\{B\}\>1\), and dynamically reorganising \(𝒯\>0\\mathcal\{T\}\>0\)—precisely the conditions identified with intrinsic curiosity\.

## 5Toy Experiments

To evaluate the Curiosity Window conjecture under controlled conditions, we implemented the two\-layer architecture \(LSM substrate \+ coherence field dynamics\) described in Section 3\. The reach noise parameterσπ\\sigma\_\{\\pi\}serves as the main control variable\.

##### Environment\.

The field evolves in a structured landscape defined by five Gaussian basins at fixed positions in a4040\-cell space\. These basins create a multi\-attractor surface against which the dynamics of exploration and trapping can be measured\. The sensory input to the LSM is a55\-dimensional signal derived from the current field state plus additive noise\.

### 5\.1Simulation 1: Phase\-Diagram Sweep

The first experiment tests whether curiosity is maximised at an intermediate level of incoherence\. We swept the reach\-noise parameter in the range

σπ∈\[5×10−4,5×10−1\]\\sigma\_\{\\pi\}\\in\[5\\times 10^\{\-4\},\\,5\\times 10^\{\-1\}\]using4545logarithmically spaced values, while fixing the field\-coupling strength at0\.30\.3\. Each run lastedT=1500T=1500time steps\.

For each run we measured: mean incoherenceI\(t\)=DKL\(πt∥yt\)I\(t\)=D\_\{\\mathrm\{KL\}\}\(\\pi\_\{t\}\\,\\\|\\,y\_\{t\}\), mean global overlapG⁡\(t\)=∑iπi​\(t\)​yi​\(t\)G\(t\)=\\sum\_\{i\}\\sqrt\{\\pi\_\{i\}\(t\)\\,y\_\{i\}\(t\)\}, the exploration entropy over the macro\-basin visitation histogram, the mean dwell time within a basin before transition, the number of coherent basins \(defined as basins with dwell time\>5\>5steps\), and the basin transition rate\.

![Refer to caption](https://arxiv.org/html/2609.05650v1/figures/ecf_phase_diagram.png)Figure 1:Phase diagram of exploration dynamics across noise regimes\. Each panel shows a different property of the dynamics as a function of the reach\-noise parameterσπ\\sigma\_\{\\pi\}: mean incoherenceI⁡\(t\)I\(t\), global overlapG⁡\(t\)G\(t\), basin coherence, exploration entropy, dwell time, transition rate, and composite curiosity metrics\. The curiosity window \(bottom centre\) is visible as a distinct intermediate regime\.
### 5\.2Simulation 2: Three\-Regime Time Series

To visualise the dynamics underlying the phase diagram, we selected three representative noise values and ran full15001500\-step trajectories \(Figure[2](https://arxiv.org/html/2609.05650#S5.F2)\):

- •Low noise \(trapped regime\):reach collapses onto one or two basins; overlap remains high, incoherence remains low, and exploration entropy is near zero\.
- •Intermediate noise \(curious regime\):the system visits multiple basins with sustained dwell times while maintaining moderate\-to\-high overlap\. This regime operationalises the Curiosity Window\.
- •High noise \(fragmented regime\):the system rapidly switches across basins without sustained occupancy; entropy is high but overlap collapses because reach and yield decouple\.

![Refer to caption](https://arxiv.org/html/2609.05650v1/figures/ecf_three_regimes.png)Figure 2:Representative time series for three noise regimes\.Top: Low noise \(σπ=0\.05\\sigma\_\{\\pi\}=0\.05\)—trapped in a single basin\.Middle: Intermediate noise \(σπ=0\.25\\sigma\_\{\\pi\}=0\.25\)—coherent transitions across multiple basins\.Bottom: High noise \(σπ=0\.8\\sigma\_\{\\pi\}=0\.8\)—rapid, incoherent switching with loss of basin structure\.A clear three\-regime structure emerges\. At low noise, the system exhibits overcoherent trapping: low incoherence, high overlap, long dwell times, and minimal exploration\. At high noise, the system enters a fragmented regime: high incoherence, collapsed overlap, frequent transitions, and loss of basin structure\.

Between these extremes lies the intermediate regime in which structured exploration is maximised\. The system maintains moderate incoherence while preserving significant global overlap, visiting multiple basins with sustained occupancy\. This balance is captured byC2C\_\{2\}, which exhibits a clear interior peak\.

Importantly, exploration quality is not monotonic with noise: while entropy increases steadily, only the combination of exploration*and*coherence yields effective behaviour\. This highlights the necessity of intermediate incoherence for intrinsically motivated exploration\.

##### Interpretive criterion\.

The core prediction is not simply that entropy should increase with noise, but that*coherent exploration*should peak at an intermediate noise level\. In the simulations, incoherence increased monotonically withσπ\\sigma\_\{\\pi\}, while overlap decreased monotonically, as expected\. The non\-trivial result is that the product of exploration and coherence peaks only when curiosity is defined structurally rather than purely entropically\.

TheC2C\_\{2\}metric exhibited a distinct interior maximum with a window width of approximately1\.741\.74decades in log\-space and clear drop\-off on both sides\.

![Refer to caption](https://arxiv.org/html/2609.05650v1/figures/ecf_curiosity_refined.png)Figure 3:Curiosity metrics as a function of noiseσπ\\sigma\_\{\\pi\}at coupling strengthηπ=0\.12\\eta\_\{\\pi\}=0\.12\. TheC2C\_\{2\}metric shows a clear peak at intermediate noise, confirming the Curiosity Window conjecture\. The\>1\.5×\>1\.5\\timesdrop\-off on both sides of the peak demonstrates that the window is sharply defined\.At low noise \(σπ<0\.1\\sigma\_\{\\pi\}<0\.1\), the system is trapped in a single basin with high overlap but no exploration\. At high noise \(σπ\>0\.6\\sigma\_\{\\pi\}\>0\.6\), the system visits all basins but with fragmented, incoherent transitions\. Only at intermediate noise does the system achieve*coherent multi\-basin exploration*—the signature of the curiosity window \(Figure[3](https://arxiv.org/html/2609.05650#S5.F3)\)\.

### 5\.3Endogenous Exploration

A critical test is whether exploration persists when external noise is removed\. In the original model with constantσπ\\sigma\_\{\\pi\}, settingσπ=0\\sigma\_\{\\pi\}=0at timeT/2T/2causes exploration to collapse to1%1\\%of the pre\-cutoff rate—a fundamental failure\.

Replacing constant noise with the endogenous modulationσπ​\(t\)=ψ⁡\(I⁡\(t\)\)⋅G⁡\(t\)\\sigma\_\{\\pi\}\(t\)=\\psi\(I\(t\)\)\\cdot G\(t\)\(Equation[17](https://arxiv.org/html/2609.05650#S3.E17)\) creates a self\-sustaining feedback loop:

1. 1\.High coherence→\\tolowII→\\tolow noise→\\toexploitation\.
2. 2\.Exploitation→\\toenvironment shifts→\\torisingII\.
3. 3\.RisingII→\\toψ⁡\(I\)\\psi\(I\)increases→\\tomore exploration\.
4. 4\.Exploration→\\tonew coherence→\\toIIdecreases\.

### 5\.4Extended Dynamics

##### Spatial coupling and global integration\.

To test global integration, we introduced the spatial coupling kernel𝐊\\mathbf\{K\}\(Equation[20](https://arxiv.org/html/2609.05650#S3.E20)\)\. Local perturbation experiments confirmed that disturbances propagate across the full field within 5–10 timesteps in the coupled system, while remaining localised in the uncoupled control\.

##### Sedimentation learning\.

We tested whether the memory fieldmtm\_\{t\}accumulates useful structure across episodes\. An agent with sedimentation \(α=0\.3\\alpha=0\.3\) was compared against a memoryless control across 50 episodes of basin\-finding\. The sedimented agent showed faster convergence to low\-incoherence states and higher final coherence \(GG\), confirming that sedimentation functions as a learning mechanism \(Figure[4](https://arxiv.org/html/2609.05650#S5.F4)\)\.

![Refer to caption](https://arxiv.org/html/2609.05650v1/figures/ecf_sedimentation_learning.png)Figure 4:Sedimentation learning across 50 episodes\. The memory agent \(red\) achieves lower incoherence and higher global overlap than the memoryless control \(blue\), demonstrating that sedimentation functions as a viable learning mechanism\. Error bands show±1\\pm 1standard deviation across runs\.
##### Non\-stationary environments\.

Basin locations were shifted atT/2T/2to test adaptation\. The endogenousψ⁡\(I\)\\psi\(I\)model recovered fastest due to the automatic increase in exploration triggered by rising incoherence\.

### 5\.5Sedimentation as Structural Learning

Standard machine learning encodes experience as discrete weight updates via backpropagation\. Our framework proposes a fundamentally different mechanism:*sedimentation*, in which the memory layermtm\_\{t\}accumulates a continuous, coherence\-weighted trace of the system’s history\. Rather than storing facts, sedimentation deforms the probability landscapeπt\\pi\_\{t\}itself, biasing future coherence\-seeking toward previously successful regions\.

#### 5\.5\.1Setup

The simulation runs for 60 episodes over a five\-basin environment\. Each episode presents a different basin as the primary attractor, cycling through Basin 0, Basin 1, and Basin 2 in sequence, so that the relevance of each region shifts over time\. Two agents are compared:

- •Memory agent—sedimentation active, with the update rule mt\+1=\(1−λm\)​mt\+λm​G​\(t\)​πt,m\_\{t\+1\}\\;=\\;\(1\-\\lambda\_\{m\}\)\\,m\_\{t\}\\;\+\\;\\lambda\_\{m\}\\,G\(t\)\\,\\pi\_\{t\},\(22\)whereG⁡\(t\)G\(t\)is the global overlap andλm=0\.003\\lambda\_\{m\}=0\.003is the sedimentation rate\. High\-coherence moments therefore contribute disproportionately to the accumulated trace\.
- •Control agent—identical architecture with sedimentation disabled \(λm=0\\lambda\_\{m\}=0\), providing a matched baseline\.

#### 5\.5\.2Memory Reshapes the Probability Landscape

Over the 60\-episode run, the memory distributionmtm\_\{t\}undergoes measurable structural change\. Shannon entropy ofmmdrops by0\.0750\.075nats, and the top 10 field positions accumulate33\.2%33\.2\\%of total probability mass \(compared with20%20\\%under a uniform distribution\)\. The memory heatmap reveals clear ridges at basin locations: the landscape sculpts itself around coherent regions, concentrating future attractor pull where the system has previously achieved highGG\.

#### 5\.5\.3Sedimentation Tracks Environmental Relevance

The episode schedule shifts emphasis progressively from Basin 0 to Basin 1 to Basin 2\. Memory follows this shift with a measurable lag\. At the end of training, Basin 2 holds the largest mass \(0\.3760\.376\), Basin 1 the next \(0\.2900\.290\), and Basin 0 the least \(0\.1940\.194\)\. Crucially, the system does not merely accumulate—it*forgets*what is no longer relevant\. Because older sedimentation fades through the exponential moving average in Equation \([22](https://arxiv.org/html/2609.05650#S5.E22)\), new coherence patterns progressively overwrite stale ones\.

#### 5\.5\.4Sedimentation Confers a Measurable Learning Advantage

Table[1](https://arxiv.org/html/2609.05650#S5.T1)reports the quantitative comparison between the memory and control agents\.

Table 1:Sedimentation learning advantage over 60 episodes\. The memory agent consistently outperforms the memoryless control on all coherence\-related metrics\.The memory agent achieves higher coherence faster because sedimented regions exert an additional pull onπt\\pi\_\{t\}, drawing the field toward previously successful attractor configurations\. The\+5\.9\+5\.9percentage\-point advantage in incoherence reduction compounds across episodes: each high\-coherence moment makes the next one slightly easier to reach\.

#### 5\.5\.5How Sedimentation Differs from Neural Network Learning

Four structural differences distinguish sedimentation from standard gradient\-based learning:

1. 1\.No weight updates\.The system’s “parameters” form a continuous probability landscape, not a vector of discrete weights\. Learning is a smooth, ongoing deformation of this landscape\. There is no loss function and no optimisation objective\.
2. 2\.Coherence\-gated storage\.Only high\-coherence moments contribute strongly tomtm\_\{t\}, because the update is weighted byG⁡\(t\)G\(t\)\(Equation \([22](https://arxiv.org/html/2609.05650#S5.E22)\)\)\. Low\-coherence episodes barely register\. This is structurally closer to how emotional salience gates consolidation in biological memory than to how backpropagation treats all training examples equally\.
3. 3\.Structural forgetting without catastrophe\.Old patterns fade naturally through the exponential decay term\(1−λm\)\(1\-\\lambda\_\{m\}\)\. There is no catastrophic forgetting, but equally no permanent storage\. Stability and plasticity are balanced by a single parameterλm\\lambda\_\{m\}\.
4. 4\.Memory drives exploration\.The memory\-gradient term in theπ\\piupdate pushes the field*away*from over\-sedimented regions, so consolidation does not collapse the system into a fixed attractor\. Learning simultaneously consolidates successful patterns and opens new territory\.

### 5\.6Multi\-Room Gridworld with Shifting Rewards

We used a custom10×1010\\times 10gridworld partitioned into four rooms with centres at\(2,2\)\(2,2\),\(2,7\)\(2,7\),\(7,2\)\(7,2\), and\(7,7\)\(7,7\)\. The agent started at the centre of the grid and received reward proportional to its proximity to the currently active room centre\. The active reward room shifted every 200 steps, cycling through all four rooms\. This non\-stationarity penalises agents that exploit a single learned policy and rewards those capable of sustained re\-exploration\.

The ECF agent was compared against a standardϵ\\epsilon\-greedy Q\-learning baseline\. Both agents used identical Q\-tables and learning rates \(α=0\.1\\alpha=0\.1,γ=0\.99\\gamma=0\.99\)\. The key difference was the exploration mechanism: the baseline decayedϵ\\epsilonfrom 0\.3 toward 0\.01 on a fixed schedule, while the ECF agent modulated exploration endogenously viaψ⁡\(I⁡\(t\)\)⋅G⁡\(t\)\\psi\(I\(t\)\)\\cdot G\(t\)\. When the reward room shifted, the ECF agent’s incoherence spiked as its reach distributionπ\\pidiverged from the now\-misaligned yieldyy, automatically increasing exploration noise\. When the agent settled into the new reward region, incoherence dropped and exploitation resumed\.

The ECF agent adapted more rapidly to reward shifts\. After each transition, its recovery time—measured as the number of steps to return to80%80\\%of peak reward rate—was consistently shorter than the baseline’s\. The sedimentation mechanism also contributed: after visiting all four rooms, the memory fieldmmdeveloped peaks at each room centre, biasing future exploration toward previously productive regions\.

This experiment established the basic viability of ECF\-augmented RL but was limited by the simplicity of the environment\. The next experiments were designed to test the framework in more challenging settings\.

## 6Reinforcement Learning Experiments

To evaluate whether coherence\-seeking dynamics confer practical advantages in RL settings, we conducted progressively more demanding experiments\. Each experiment stress\-tests a specific claim: that endogenous curiosity driven by incoherence produces more adaptive exploration than standard strategies, that this advantage grows under non\-stationary conditions, and that the mechanism generalises from discrete to continuous action spaces\. In all experiments, the ECF agent maintained full field dynamics—reach \(π\\pi\), yield \(yy\), and memory \(mm\) distributions overN=30N=30basins—with incoherenceI\(t\)=DKL\(π∥y\)I\(t\)=D\_\{\\mathrm\{KL\}\}\(\\pi\\\|y\), global overlapG⁡\(t\)=∑iπi​yiG\(t\)=\\sum\_\{i\}\\sqrt\{\\pi\_\{i\}\\,y\_\{i\}\}, and the endogenous noise functionσπ​\(t\)=ψ⁡\(I⁡\(t\)\)⋅G⁡\(t\)\\sigma\_\{\\pi\}\(t\)=\\psi\(I\(t\)\)\\cdot G\(t\), whereψ\(I\)=Iexp\(−I/I∗\)\\psi\(I\)=I\\exp\(\-I/I^\{\*\}\)peaks at intermediate incoherenceI∗I^\{\*\}\.

### 6\.1ECF–PPO Hybrid Approach

To address more complex RL problems we implemented a hybrid approach \(ECF–PPO\) combining the optimisation stability of Proximal Policy Optimization \(PPO\) with an auxiliary memory\-based dynamical system\. PPO provides the RL backbone through a policy networkπθ​\(at∣st\)\\pi\_\{\\theta\}\(a\_\{t\}\\mid s\_\{t\}\)and a value functionVϕ​\(st\)V\_\{\\phi\}\(s\_\{t\}\), while the ECF module augments the agent with internal latent variables that track current experience, expectation, and memory\.

Letztz\_\{t\}denote a latent encoding of the observationsts\_\{t\}\. The ECF module maintains three internal quantities: a policy\-side expectationπt\\pi\_\{t\}, a current experience signalyty\_\{t\}, and a sedimented memory tracemtm\_\{t\}\. The dynamics are:

yt=\(1−αy\)​yt−1\+αy​zt,mt=\(1−αm\)​mt−1\+αm​yt,y\_\{t\}=\(1\-\\alpha\_\{y\}\)y\_\{t\-1\}\+\\alpha\_\{y\}z\_\{t\},\\qquad m\_\{t\}=\(1\-\\alpha\_\{m\}\)m\_\{t\-1\}\+\\alpha\_\{m\}y\_\{t\},with richer variants replacing the single memory tracemtm\_\{t\}by multi\-timescale components \(fast, medium, and slow\)\. From these variables we define an incoherence signal and a novelty signal:

It=‖πt−yt‖2\+‖yt−mt‖2,Nt=‖yt−mt‖\.I\_\{t\}=\\\|\\pi\_\{t\}\-y\_\{t\}\\\|^\{2\}\+\\\|y\_\{t\}\-m\_\{t\}\\\|^\{2\},\\qquad N\_\{t\}=\\\|y\_\{t\}\-m\_\{t\}\\\|\.These quantities are transformed into an intrinsic drivepECF,tp\_\{\\mathrm\{ECF\},t\}, gated by a coherence factorψ⁡\(It\)\\psi\(I\_\{t\}\)and a controllerλt\\lambda\_\{t\}:

rttot=rtext\+λt​pECF,t\.r\_\{t\}^\{\\mathrm\{tot\}\}=r\_\{t\}^\{\\mathrm\{ext\}\}\+\\lambda\_\{t\}\\,p\_\{\\mathrm\{ECF\},t\}\.PPO’s role is unchanged at the optimisation level: it performs policy and value updates with clipped objectives and advantage estimation\. The difference is that the reward stream now contains a structured intrinsic component derived from internal consistency, novelty, and memory mismatch\.

![Refer to caption](https://arxiv.org/html/2609.05650v1/figures/ECF_RL_diagram.png)Figure 5:Architecture of the ECF–PPO hybrid agent\. The LSM substrate produces a latent encodingztz\_\{t\}, from which the ECF module computes reach \(πt\\pi\_\{t\}\), yield \(yty\_\{t\}\), and memory \(mtm\_\{t\}\) fields\. Incoherence and novelty signals modulate an intrinsic rewardpECF,tp\_\{\\mathrm\{ECF\},t\}that augments the external reward before PPO updates the policy\.PPO contributes stable policy optimisation and strong baseline performance, while the ECF module contributes internal structure: tracking what is familiar, what is changing, and what is inconsistent with the agent’s latent expectations\. The resulting agent is guided by a memory\-conditioned internal signal that favours structured exploration over purely random behaviour\.

### 6\.2LunarLander\-v2 Stress Test

We used the OpenAI GymnasiumLunarLander\-v2environment, which requires the agent to control a spacecraft’s thrusters to land safely on a pad\. The eight\-dimensional continuous state space \(position, velocity, angle, angular velocity, leg contact\) was discretised into bins for tabular Q\-learning\.

#### 6\.2\.1ECF\-RichMem Algorithm

To handle this more complex problem we developed an extended version of ECF:ECF\-RichMem\. It extends the basic formulation by incorporating a multi\-timescale memory structure that captures both short\-term fluctuations and long\-term regularities\.

Letst∈𝒮s\_\{t\}\\in\\mathcal\{S\}denote the environment state at timett, and letzt=fθ​\(st\)z\_\{t\}=f\_\{\\theta\}\(s\_\{t\}\)be a latent encoding\. The agent maintains three internal variables:

- •yty\_\{t\}: current experiential state,
- •πt\\pi\_\{t\}: endogenous expectation \(policy\-side latent\),
- •𝐦t=\{mt\(f\),mt\(m\),mt\(s\)\}\\mathbf\{m\}\_\{t\}=\\\{m\_\{t\}^\{\(f\)\},m\_\{t\}^\{\(m\)\},m\_\{t\}^\{\(s\)\}\\\}: multi\-timescale memory \(fast, medium, slow components\)\.

##### Internal dynamics\.

The experiential state is updated as a filtered version of the latent observation:

yt=\(1−αy\)​yt−1\+αy​zt\.y\_\{t\}=\(1\-\\alpha\_\{y\}\)y\_\{t\-1\}\+\\alpha\_\{y\}z\_\{t\}\.\(23\)
Each memory component evolves according to its own timescale:

mt\(k\)=\(1−αk\)​mt−1\(k\)\+αk​yt,k∈\{f,m,s\},αf\>αm\>αs\.m\_\{t\}^\{\(k\)\}=\(1\-\\alpha\_\{k\}\)m\_\{t\-1\}^\{\(k\)\}\+\\alpha\_\{k\}y\_\{t\},\\quad k\\in\\\{f,m,s\\\},\\quad\\alpha\_\{f\}\>\\alpha\_\{m\}\>\\alpha\_\{s\}\.\(24\)
The expectationπt\\pi\_\{t\}is updated from the policy representation\.

##### Coherence and novelty signals\.

The agent computes internal signals based on discrepancies between its internal variables:

It\\displaystyle I\_\{t\}=‖πt−yt‖2\+∑kwk​‖yt−mt\(k\)‖2,\\displaystyle=\\\|\\pi\_\{t\}\-y\_\{t\}\\\|^\{2\}\+\\sum\_\{k\}w\_\{k\}\\\|y\_\{t\}\-m\_\{t\}^\{\(k\)\}\\\|^\{2\},\(25\)Nt\\displaystyle N\_\{t\}=∑kwk​‖yt−mt\(k\)‖,\\displaystyle=\\sum\_\{k\}w\_\{k\}\\\|y\_\{t\}\-m\_\{t\}^\{\(k\)\}\\\|,\(26\)whereItI\_\{t\}is the incoherence,NtN\_\{t\}is novelty relative to memory, andwkw\_\{k\}are weighting coefficients\.

##### Intrinsic modulation\.

A coherence gating function modulates the intrinsic signal:

ψ⁡\(It\)=exp⁡\(−β​It\),\\psi\(I\_\{t\}\)=\\exp\(\-\\beta I\_\{t\}\),\(27\)or alternatively a bell\-shaped function that increases exploration under moderate incoherence\. The intrinsic signal is:

pECF,t=ψ⁡\(It\)⋅Nt\.p\_\{\\text\{ECF\},t\}=\\psi\(I\_\{t\}\)\\cdot N\_\{t\}\.\(28\)
This signal defines an intrinsic reward:

rttot=rtext\+λt​pECF,t,r\_\{t\}^\{\\text\{tot\}\}=r\_\{t\}^\{\\text\{ext\}\}\+\\lambda\_\{t\}\\,p\_\{\\text\{ECF\},t\},\(29\)whereλt\\lambda\_\{t\}is a scaling coefficient that may depend on entropy, learning progress, or other adaptive factors\.

##### Interpretation\.

Unlike standard intrinsic curiosity methods that reward unpredictability, ECF\-RichMem encourages exploration based on structured deviations from internal memory\. The multi\-timescale memory allows the agent to distinguish between transient fluctuations and persistent environmental changes, enabling a balance between stability and adaptability\.

#### 6\.2\.2Experimental Protocol

The implementation evaluates three agents—ECF,ECF\-RichMem, andICM—under a four\-phase protocol designed to separate baseline learning, retention, disturbance handling, and recovery\.

##### Phase 1: Standard training\.

All agents were trained in the nominal LunarLander environment to measure baseline learning\.

##### Phase 2: Retention\.

The environment remained nominal\. This phase tested how well the policy carried forward behaviour acquired in Phase 1\.

##### Phase 3: Disturbance \(wind\)\.

Wind and turbulence were introduced to test robustness under shifted dynamics\.

##### Phase 4: Recovery\.

The environment returned to nominal dynamics to measure post\-disturbance recovery\.

All agents ran for 500 epochs over 10 different seeds\. Performance was measured as the average episodic return over the final portion of each phase\.

Table 2:Mean performance across phases in LunarLander\-v2\. Values are mean±\\pmstandard deviation across seeds\. Bold indicates the best\-performing agent in each phase\.
##### Phases 1–2: Baseline learning and retention\.

In the nominal environment,ECF\-RichMemoutperformed the other approaches, achieving positive mean performance in both Phase 1 \(16\.316\.3\) and Phase 2 \(40\.040\.0\)\. The improvement from Phase 1 to Phase 2 suggests that the multi\-timescale memory helps stabilise and reinforce coherent patterns over time\. BothECFandICMremained negative on average\.

##### Phase 3: Disturbance\.

All methods deteriorated, butECF\-RichMemremained the strongest performer \(−16\.7\-16\.7\), substantially better thanECF\(−58\.3\-58\.3\) andICM\(−90\.3\-90\.3\)\. The multi\-timescale memory provided a more stable reference under environmental shift\.

##### Phase 4: Recovery\.

ECF\-RichMemproduced the best result \(5\.75\.7\), showing clear re\-stabilisation after disturbance, unlike the other methods which remained substantially impaired\.

##### Summary\.

ECF\-RichMemis consistently the strongest approach across all four phases\. Its advantage lies not only in higher average performance but in a qualitatively different adaptation profile: it learns better in the nominal regime, retains useful structure more effectively, degrades less severely under disturbance, and recovers more successfully\.ICM, while designed to encourage exploration, performs poorly throughout, especially under perturbation and recovery, suggesting that novelty\-seeking alone is insufficient for maintaining coherent control\.

### 6\.3BipedalWalker\-v3 Continuous Control

To evaluate the framework in a more demanding continuous\-control setting, we usedBipedalWalker\-v3, a locomotion task requiring the agent to coordinate four continuous joint torques for stable forward walking\.

Both agents were based on a neural\-network policy trained with PPO\. The comparison therefore asks whether the ECF module provides additional value*on top of*a strong RL algorithm:

- •PPO: standard neural policy and value\-function architecture\.
- •PPO\+ECF: PPO augmented with an ECF\-based intrinsic modulation mechanism\.

The ECF–PPO agent retained the same conceptual ingredients: an internal experiential state, an endogenous expectation, and a sedimented memory trace\. These variables produced coherence\- and novelty\-related signals that modulated the intrinsic guidance provided to the policy\.

![Refer to caption](https://arxiv.org/html/2609.05650v1/figures/bipeda_4stages.png)Figure 6:BipedalWalker\-v3 results across the four experimental phases\. Phase 1: standard training under nominal dynamics\. Phase 2: retention under nominal conditions without reward\. Phase 3: disturbance via altered gravity\. Phase 4: recovery after restoring nominal dynamics\. PPO\+ECF shows substantially better robustness in Phase 3\.The evaluation followed the same four\-phase structure as the LunarLander experiments: Phase 1 \(standard training\), Phase 2 \(retention without reward\), Phase 3 \(gravity perturbation\), and Phase 4 \(recovery\)\.

Table 3:BipedalWalker\-v3 results across four phases\. Values are mean±\\pmstandard deviation\. Bold indicates the best\-performing agent in each phase\.##### Phases 1–2: Training and retention\.

PPO and PPO\+ECF achieved similar mean performance \(31\.131\.1vs\.29\.029\.0in Phase 1\), though PPO\+ECF exhibited substantially higher variance\. In Phase 2, PPO slightly outperformed PPO\+ECF \(42\.642\.6vs\.37\.437\.4\)\. No clear ECF advantage is evident in these stationary conditions\.

##### Phase 3: Gravity disturbance\.

PPO’s performance dropped to−20\.3\-20\.3, while PPO\+ECF maintained a strongly positive mean \(45\.345\.3\)\. Despite high variance, this phase reveals the main benefit of ECF: improved adaptability under environmental disturbance\.

##### Phase 4: Recovery\.

Both agents showed reduced performance relative to earlier phases\. PPO\+ECF retained a slight advantage \(5\.45\.4vs\.3\.13\.1\) but with high variance\. Neither method fully recovered its previous performance\.

### 6\.4Cross\-Experiment Analysis

Across all experiments, a consistent pattern emerges\. The ECF agent’s advantage is smallest in stable, well\-rewarded environments \(Phase 1\) and largest when the environment is non\-stationary or reward is absent\. This is a precise confirmation of the framework’s theoretical predictions: the Curiosity Window conjecture states that coherence\-seeking exploration is most valuable at intermediate levels of tension, and the disturbance phases place the system squarely in this regime\.

The transition from discrete \(LunarLander\) to continuous \(BipedalWalker\) action spaces amplified the ECF advantage, suggesting that the framework’s benefits scale with the dimensionality of the exploration problem\. In discrete spaces, even random exploration has a reasonable probability of finding useful actions; in continuous spaces, directed exploration becomes essential, and the memory\-gradient mechanism provides exactly this directionality\.

The sedimentation dynamics were consistent across all experiments\. In every case, the memory fieldmmdeveloped concentrated peaks at frequently visited, high\-coherence state regions, and these peaks persisted through environmental changes\.

Finally, the endogenous noise functionσπ​\(t\)=ψ⁡\(I⁡\(t\)\)⋅G⁡\(t\)\\sigma\_\{\\pi\}\(t\)=\\psi\(I\(t\)\)\\cdot G\(t\)proved essential\. Without it—using constant noise—the agent’s exploration was either excessive or insufficient, with no intermediate regime producing adaptive behaviour\.

## 7Discussion

### 7\.1Main Differences of the ECF Approach

The main distinction between ECF\-based approaches and conventional RL baselines is that ECF introduces an explicit*internal organisation of experience*\. Standard PPO optimises a policy and value function directly from returns and advantages\. Curiosity\-driven methods such as ICM add an intrinsic reward based on prediction error\. By contrast, ECF maintains internal variables representing expectation, current experience, and sedimented memory, computing structured signals from their relationships\.

This leads to several potential advantages:

##### Memory\-guided adaptation\.

The memory componentmtm\_\{t\}acts as a slowly changing reference against which the current latent state is compared, allowing the agent to distinguish between familiar patterns, transient fluctuations, and meaningful deviations\.

##### Weak, sparse, or delayed rewards\.

ECF’s internal signals supply learning pressure even when external rewards are weak or absent\. Unlike pure curiosity, which rewards surprise itself, ECF rewards novelty relative to memory and coherence, making exploration more structured\.

##### Transfer learning and non\-stationarity\.

The sedimented summary of past experience can serve as a scaffold for adaptation in new but related settings, making ECF a candidate for continual and non\-stationary RL\.

##### Interpretability\.

ECF exposes internal quantities \(incoherence, novelty, memory mismatch\) that can be directly monitored, making it easier to analyse*why*the agent is exploring or changing its behaviour\.

### 7\.2Relationship to Existing Frameworks

The core innovation is the*closed feedback loop*between coherence state and exploration intensity\. While each component exists in isolation—reservoir computing\([Maass et al\., 2002](https://arxiv.org/html/2609.05650#bib.bib5)\), intrinsic motivation\([Pathak et al\., 2017](https://arxiv.org/html/2609.05650#bib.bib6)\), coherence dynamics\([Friston, 2010](https://arxiv.org/html/2609.05650#bib.bib3)\), noise modulation \(simulated annealing\)—no prior work couples them such that the system’s own incoherence drives its exploration, which reshapes its coherence landscape, which changes its incoherence\. This circular causality is the specific contribution\.

Our simulations demonstrate three results: first, the curiosity window creates a*metastable middle regime*between rigidity and chaos, consistent with theories linking consciousness to criticality\([Tononi et al\., 2016](https://arxiv.org/html/2609.05650#bib.bib7)\); second, the endogenousψ⁡\(I\)\\psi\(I\)loop ensures that exploratory behaviour is generated by the system’s own coherence dynamics rather than imposed externally; third, endogenous exploration fails withoutψ⁡\(I\)\\psi\(I\), providing falsifiable criteria for systems that lack this property\.

### 7\.3Limitations

1. 1\.Scale: All simulations useN=50N=50–6060reservoir nodes\. Scaling to realistic dimensions is untested\.
2. 2\.Metastable dwell times: Basin dwell times remain short \(∼\\sim2 steps\)\. Deeper attractor landscapes or largerNNmay be needed\.
3. 3\.Metric sensitivity: OnlyC2C\_\{2\}clearly supports the conjecture; alternative metrics show weaker effects\.
4. 4\.High variance: The PPO\+ECF results in BipedalWalker exhibit substantially higher variance than the PPO baseline, indicating that the ECF modulation may introduce instability in certain runs\.
5. 5\.No deep learning integration: Combining ECF with deep policy networks \(beyond the PPO backbone\) remains future work\.

### 7\.4Hardware Substrate Considerations

Our analysis of implementation substrates suggests a hybrid neuromorphic–optical architecture as optimal:

- •Neuromorphic\(e\.g\., Intel Loihi, SpiNNaker\): Natural fit for the LSM reservoir layer; spiking dynamics provide temporal richness\.
- •Optical: Potential for ultra\-fast reservoir computation via photonic reservoirs\.
- •Hybrid: Neuromorphic reservoir \+ digital coherence control layer scores highest \(33/40\) in our substrate evaluation\.

## 8Conclusion

We have provided the first computational validation of the Experiential Coherence Framework’s central predictions\. The Curiosity Window conjecture is confirmed under theC2C\_\{2\}metric, is absent in Active Inference agents, and becomes self\-sustaining with the endogenousψ⁡\(I\)\\psi\(I\)modulation\. These results demonstrate that ECF is not merely a philosophical framework but a computationally tractable architecture with distinctive, falsifiable predictions\.

The most important finding is that replacing exogenous noise withσπ​\(t\)=ψ⁡\(I⁡\(t\)\)⋅G⁡\(t\)\\sigma\_\{\\pi\}\(t\)=\\psi\(I\(t\)\)\\cdot G\(t\)transforms the framework from one that*describes*coherence dynamics to one that*generates*them endogenously\. In RL experiments, the ECF\-augmented agents showed particular strength under non\-stationary conditions and environmental disturbance, precisely the regimes where the curiosity window is predicted to be most valuable\.

TheC2C\_\{2\}metric \(coherent basins×\\timesglobal overlap×\\timestransition rate\) emerged as the robust, empirically grounded curiosity functional, consistently exhibiting a clear interior peak across parameter sweeps\. Future work should address scaling to higher\-dimensional problems, integration with deep policy architectures beyond PPO, and validation in richer environments with longer horizons\.

## Acknowledgments

The author gratefully acknowledges Nuno P\. Barradas \(Centro de Ciências e Tecnologias Nucleares, Instituto Superior Técnico, Universidade de Lisboa\) for his careful reading of the manuscript and his valuable feedback, which greatly improved the final version of this work\.

## References

- Burda et al\. \(2018\)Y\. Burda, H\. Edwards, A\. Storkey, and O\. Klimov\.Exploration by random network distillation\.*arXiv preprint arXiv:1810\.12894*, 2018\.
- ECF Authors \(2026\)ECF Authors\.The experiential coherence framework\.*Manuscript*, 2026\.
- Friston \(2010\)K\. Friston\.The free\-energy principle: a unified brain theory?*Nature Reviews Neuroscience*, 11\(2\):127–138, 2010\.
- Friston et al\. \(2017\)K\. Friston, T\. FitzGerald, F\. Rigoli, P\. Schwartenbeck, and G\. Pezzulo\.Active inference: a process theory\.*Neural Computation*, 29\(1\):1–49, 2017\.
- Maass et al\. \(2002\)W\. Maass, T\. Natschläger, and H\. Markram\.Real\-time computing without stable states: A new framework for neural computation based on perturbations\.*Neural Computation*, 14\(11\):2531–2560, 2002\.
- Pathak et al\. \(2017\)D\. Pathak, P\. Agrawal, A\. A\. Efros, and T\. Darrell\.Curiosity\-driven exploration by self\-predictive next feature learning\.In*ICML*, 2017\.
- Tononi et al\. \(2016\)G\. Tononi, M\. Boly, M\. Massimini, and C\. Koch\.Integrated information theory: from consciousness to its physical substrate\.*Nature Reviews Neuroscience*, 17\(7\):450–461, 2016\.

## Appendix AAnalytical Proof of the Curiosity Window

The Curiosity Window Conjecture \(Conjecture[1](https://arxiv.org/html/2609.05650#Thmconjecture1)\) asserts the existence of thresholds0<α<β<∞0<\\alpha<\\beta<\\inftysuch that sustained intrinsic curiosity is possible only whenα≤I⁡\(t\)≤β\\alpha\\leq I\(t\)\\leq\\beta\. In this section we prove the conjecture in closed form for the minimal two\-basin system, derive explicit bounds onα\\alphaandβ\\beta, and sketch the extension toK\>2K\>2basins\.

### A\.1The 2\-Basin System

###### Definition 5\(2\-basin field\)\.

LetE=\{e1,e2\}E=\\\{e\_\{1\},e\_\{2\}\\\}with counting measureμ\\mu\. Every density onEEis parameterised by a single scalar:

πt=\(pt,1−pt\),y=\(q,1−q\),pt,q∈\(0,1\)\.\\pi\_\{t\}=\(p\_\{t\},\\;1\-p\_\{t\}\),\\qquad y=\(q,\\;1\-q\),\\qquad p\_\{t\},q\\in\(0,1\)\.We fix the yield aty=\(q,1−q\)y=\(q,1\-q\)withq\>12q\>\\tfrac\{1\}\{2\}\(basin 1 preferred\) and let the reach evolve under the*frozen\-yield mirror flow*\(Equation[3](https://arxiv.org/html/2609.05650#S2.E3)\) with additive Gaussian noise of scaleσ\>0\\sigma\>0:

pt\+1=clip\[ε,1−ε\]⁡\(pt−η​pt​\(log⁡ptq−𝔼πt​\[log⁡πty\]\)\+σ​ξt\),ξt∼𝒩⁡\(0,1\),p\_\{t\+1\}=\\operatorname\{clip\}\_\{\[\\varepsilon,\\,1\-\\varepsilon\]\}\\\!\\Bigl\(p\_\{t\}\-\\eta\\,p\_\{t\}\\Bigl\(\\log\\frac\{p\_\{t\}\}\{q\}\-\\mathbb\{E\}\_\{\\pi\_\{t\}\}\\\!\\Bigl\[\\log\\frac\{\\pi\_\{t\}\}\{y\}\\Bigr\]\\Bigr\)\+\\sigma\\,\\xi\_\{t\}\\Bigr\),\\qquad\\xi\_\{t\}\\sim\\mathcal\{N\}\(0,1\),\(30\)whereη\>0\\eta\>0is the mirror\-flow learning rate andε\>0\\varepsilon\>0is a boundary guard\.

The functionals reduce to scalar functions ofpp:

I⁡\(p\)\\displaystyle I\(p\)=p​log⁡pq\+\(1−p\)​log⁡1−p1−q,\\displaystyle=p\\log\\frac\{p\}\{q\}\+\(1\-p\)\\log\\frac\{1\-p\}\{1\-q\},\(31\)G⁡\(p\)\\displaystyle G\(p\)=p​q\+\(1−p\)​\(1−q\)\.\\displaystyle=\\sqrt\{p\\,q\}\+\\sqrt\{\(1\-p\)\(1\-q\)\}\.\(32\)

### A\.2Deterministic Dynamics

###### Lemma 1\(Unique attractor\)\.

Forσ=0\\sigma=0the mirror flow has a unique globally attracting fixed point atp∗=qp^\{\*\}=q, whereI⁡\(q\)=0I\(q\)=0andG⁡\(q\)=1G\(q\)=1\.

###### Proof\.

The incoherenceI⁡\(pt\)I\(p\_\{t\}\)is a strict Lyapunov function:dd​t​I​\(pt\)=−Varπt⁡\(log⁡πty\)≤0\\frac\{d\}\{dt\}I\(p\_\{t\}\)=\-\\operatorname\{Var\}\_\{\\pi\_\{t\}\}\\\!\\bigl\(\\log\\frac\{\\pi\_\{t\}\}\{y\}\\bigr\)\\leq 0, with equality if and only ifπt=y\\pi\_\{t\}=y\. On\|E\|=2\|E\|=2the variance is

Varπ⁡\(g\)=p⁡\(1−p\)​\(log⁡pq−log⁡1−p1−q\)2,\\operatorname\{Var\}\_\{\\pi\}\(g\)=p\(1\-p\)\\Bigl\(\\log\\frac\{p\}\{q\}\-\\log\\frac\{1\-p\}\{1\-q\}\\Bigr\)^\{\\\!2\},which vanishes if and only ifp=qp=q\(sincep∈\(0,1\)p\\in\(0,1\)\)\. Hencep∗=qp^\{\*\}=qis the unique fixed point\. ∎

###### Corollary 1\(Zero\-noise trapping\)\.

Atσ=0\\sigma=0, oncept≈qp\_\{t\}\\approx qthe system is permanently trapped in basin alignment\. No inter\-basin transitions occur and exploration is identically zero\.

### A\.3Stochastic Dynamics and the Effective Potential

Forσ\>0\\sigma\>0the dynamics \([30](https://arxiv.org/html/2609.05650#A1.E30)\) become a discrete Langevin process on\(0,1\)\(0,1\)\. The deterministic drift simplifies to

f⁡\(p\)=−η​p​\(1−p\)​\(log⁡pq−log⁡1−p1−q\)=−V′​\(p\),f\(p\)=\-\\eta\\,p\(1\-p\)\\Bigl\(\\log\\frac\{p\}\{q\}\-\\log\\frac\{1\-p\}\{1\-q\}\\Bigr\)=\-V^\{\\prime\}\(p\),\(33\)with*effective potential*

V⁡\(p\)=η​I​\(p\)\.V\(p\)=\\eta\\,I\(p\)\.\(34\)The stationary density of the associated Langevin process is

ρ∞​\(p\)∝exp⁡\(−2​V​\(p\)σ2\)=exp⁡\(−βeff​I​\(p\)\),\\rho\_\{\\infty\}\(p\)\\;\\propto\\;\\exp\\\!\\Bigl\(\-\\frac\{2V\(p\)\}\{\\sigma^\{2\}\}\\Bigr\)=\\exp\\\!\\Bigl\(\-\\beta\_\{\\mathrm\{eff\}\}\\,I\(p\)\\Bigr\),\(35\)where we define the*effective inverse temperature*

βeff=2​ησ2\.\\beta\_\{\\mathrm\{eff\}\}\\;=\\;\\frac\{2\\eta\}\{\\sigma^\{2\}\}\.\(36\)
##### Barrier height\.

The mirror\-flow attractor sits atp=qp=q\. The “opposite basin” corresponds top≈1−qp\\approx 1\-q, and the barrier separating them is atp=12p=\\tfrac\{1\}\{2\}\. SettingΔ=q−12\>0\\Delta=q\-\\tfrac\{1\}\{2\}\>0:

Δ​V=V⁡\(12\)−V⁡\(q\)=η​I​\(12\)=−η2​log⁡\(1−4​Δ2\)\.\\Delta V=V\\\!\\bigl\(\\tfrac\{1\}\{2\}\\bigr\)\-V\(q\)=\\eta\\,I\\\!\\bigl\(\\tfrac\{1\}\{2\}\\bigr\)=\-\\frac\{\\eta\}\{2\}\\log\\\!\\bigl\(1\-4\\Delta^\{2\}\\bigr\)\.\(37\)

### A\.4Three Regimes

Regime I: Overcoherence \(βeff≫1\\beta\_\{\\mathrm\{eff\}\}\\gg 1\)\.A Laplace approximation aroundp=qp=qgivesρ∞≈𝒩⁡\(q,σ2/\[4​η​q​\(1−q\)\]\)\\rho\_\{\\infty\}\\approx\\mathcal\{N\}\\\!\\bigl\(q,\\;\\sigma^\{2\}/\[4\\eta q\(1\-q\)\]\\bigr\)\. The exploration rate

ℰ⁡\(δ\)=1−∫q−δq\+δρ∞​\(p\)​𝑑p≈2​Φ​\(−δ​4​η​q​\(1−q\)σ2\)⟶0as​σ→0\.\\mathcal\{E\}\(\\delta\)=1\-\\int\_\{q\-\\delta\}^\{q\+\\delta\}\\rho\_\{\\infty\}\(p\)\\,dp\\;\\approx\\;2\\,\\Phi\\\!\\Bigl\(\-\\delta\\sqrt\{\\frac\{4\\eta q\(1\-q\)\}\{\\sigma^\{2\}\}\}\\Bigr\)\\;\\longrightarrow\\;0\\quad\\text\{as \}\\sigma\\to 0\.The system is locked \(overcoherence rigidity\)\.

Regime III: Fragmentation \(βeff≪1\\beta\_\{\\mathrm\{eff\}\}\\ll 1\)\.ρ∞→Uniform⁡\(0,1\)\\rho\_\{\\infty\}\\to\\mathrm\{Uniform\}\(0,1\)\. Coherent basin occupancyB=ρ∞​\(p<δ\)\+ρ∞​\(p\>1−δ\)→2​δ→0B=\\rho\_\{\\infty\}\(p<\\delta\)\+\\rho\_\{\\infty\}\(p\>1\-\\delta\)\\to 2\\delta\\to 0for smallδ\\delta\. The system fragments\.

Regime II: Curiosity Window \(βeff∼1\\beta\_\{\\mathrm\{eff\}\}\\sim 1\)\.The system has enough noise to escape the coherent fixed point and visit both basins, but not so much that it loses coherent occupancy\.

We now state and prove the central result\.

###### Theorem 2\(Curiosity Window, 2\-Basin Case\)\.

Consider the 2\-basin system of Definition[5](https://arxiv.org/html/2609.05650#Thmdefinition5)with yieldy=\(q,1−q\)y=\(q,1\-q\),q\>12q\>\\tfrac\{1\}\{2\}, mirror\-flow rateη\>0\\eta\>0, and noise scaleσ\>0\\sigma\>0\. Define:

- •Kramers escape rate \(basin 1→\\tobasin 2\): 𝒯⁡\(σ\)=η​\|V′′​\(q\)\|⋅\|V′′​\(12\)\|π​exp⁡\(−2​Δ​Vσ2\)\.\\mathcal\{T\}\(\\sigma\)=\\frac\{\\eta\\sqrt\{\|V^\{\\prime\\prime\}\(q\)\|\\cdot\|V^\{\\prime\\prime\}\(\\tfrac\{1\}\{2\}\)\|\\,\}\}\{\\pi\}\\;\\exp\\\!\\Bigl\(\-\\frac\{2\\Delta V\}\{\\sigma^\{2\}\}\\Bigr\)\.\(38\)
- •Mean global overlap under noise:Gavg​\(σ\)=𝔼ρ∞​\[G⁡\(p\)\]G\_\{\\mathrm\{avg\}\}\(\\sigma\)=\\mathbb\{E\}\_\{\\rho\_\{\\infty\}\}\[G\(p\)\]\.
- •*Curiosity functional*: C^2​\(σ\)=Gavg​\(σ\)⋅𝒯⁡\(σ\)\.\\hat\{C\}\_\{2\}\(\\sigma\)=G\_\{\\mathrm\{avg\}\}\(\\sigma\)\\;\\cdot\\;\\mathcal\{T\}\(\\sigma\)\.\(39\)

Then:

1. 1\.C^2​\(0\+\)=0\\hat\{C\}\_\{2\}\(0^\{\+\}\)=0\.
2. 2\.C^2​\(σ\)→0\\hat\{C\}\_\{2\}\(\\sigma\)\\to 0asσ→∞\\sigma\\to\\infty\.
3. 3\.C^2\\hat\{C\}\_\{2\}attains a unique interior maximum atσ∗∈\(0,∞\)\\sigma^\{\*\}\\in\(0,\\infty\)satisfying σ∗=2​Δ​VW⁡\(Δ​V/ηG\)\\boxed\{\\;\\sigma^\{\*\}=\\sqrt\{\\frac\{2\\Delta V\}\{\\,W\\\!\\bigl\(\\Delta V/\\eta\_\{G\}\\bigr\)\\,\}\}\\;\}\(40\)whereWWis the principal branch of the LambertWWfunction andηG=−dd⁡\(σ2\)​Gavg\|σ=0\\eta\_\{G\}=\-\\frac\{d\}\{d\(\\sigma^\{2\}\)\}G\_\{\\mathrm\{avg\}\}\\big\|\_\{\\sigma=0\}is the noise\-sensitivity of overlap\.

In the well\-separated regimeΔ​V≫ηG\\Delta V\\gg\\eta\_\{G\}this simplifies to

σ∗≈−η​log⁡\(1−4​Δ2\)log⁡\(−η​log⁡\(1−4​Δ2\)2​ηG\)\.\\sigma^\{\*\}\\;\\approx\\;\\sqrt\{\\frac\{\-\\eta\\log\(1\-4\\Delta^\{2\}\)\}\{\\log\\\!\\bigl\(\\frac\{\-\\eta\\log\(1\-4\\Delta^\{2\}\)\}\{2\\eta\_\{G\}\}\\bigr\)\}\}\\,\.\(41\)

###### Proof\.

Part[1](https://arxiv.org/html/2609.05650#A1.I3.i1)\.Asσ→0\+\\sigma\\to 0^\{\+\}, the Kramers rate \([38](https://arxiv.org/html/2609.05650#A1.E38)\) decays as𝒯∼exp\(−2ΔV/σ2\)→0\\mathcal\{T\}\\sim\\exp\(\-2\\Delta V/\\sigma^\{2\}\)\\to 0whileGavg→G⁡\(q\)=1G\_\{\\mathrm\{avg\}\}\\to G\(q\)=1\. The product vanishes\.

Part[2](https://arxiv.org/html/2609.05650#A1.I3.i2)\.Asσ→∞\\sigma\\to\\infty,ρ∞→Uniform⁡\(0,1\)\\rho\_\{\\infty\}\\to\\mathrm\{Uniform\}\(0,1\)\. The mean overlap converges to

Gavgunif=∫01\[p​q\+\(1−p\)​\(1−q\)\]​𝑑p=23​\(q\+1−q\)<1\.G\_\{\\mathrm\{avg\}\}^\{\\mathrm\{unif\}\}=\\int\_\{0\}^\{1\}\\\!\\bigl\[\\sqrt\{pq\}\+\\sqrt\{\(1\-p\)\(1\-q\)\}\\,\\bigr\]\\,dp=\\tfrac\{2\}\{3\}\\bigl\(\\sqrt\{q\}\+\\sqrt\{1\-q\}\\,\\bigr\)<1\.Coherent basin occupancyB⁡\(σ\)→2​δ→0B\(\\sigma\)\\to 2\\delta\\to 0for any fixed smallδ\\delta, so the transition\-weighted curiosity vanishes\.

Part \(iii\): Existence\.C^2\\hat\{C\}\_\{2\}is continuous on\(0,∞\)\(0,\\infty\), vanishes at both limits, and is strictly positive for intermediateσ\\sigma\(since𝒯\>0\\mathcal\{T\}\>0andGavg\>0G\_\{\\mathrm\{avg\}\}\>0for allσ\>0\\sigma\>0\)\. By the extreme value theorem,C^2\\hat\{C\}\_\{2\}attains a maximum in the interior\.

Part \(iii\): Uniqueness\.Taking the logarithm,log⁡C^2=log⁡Gavg\+log⁡𝒯\\log\\hat\{C\}\_\{2\}=\\log G\_\{\\mathrm\{avg\}\}\+\\log\\mathcal\{T\}\. Differentiating with respect tos=σ2s=\\sigma^\{2\}:

dd​s​log⁡𝒯=Δ​Vs2\>0\(strictly decreasing ins, convex\),\\frac\{d\}\{ds\}\\log\\mathcal\{T\}=\\frac\{\\Delta V\}\{s^\{2\}\}\>0\\qquad\\text\{\(strictly decreasing in $s$, convex\)\},dd​s​log⁡Gavg<0\(noise degrades overlap, bounded derivative\)\.\\frac\{d\}\{ds\}\\log G\_\{\\mathrm\{avg\}\}<0\\qquad\\text\{\(noise degrades overlap, bounded derivative\)\}\.The sumdd​s​log⁡C^2\\frac\{d\}\{ds\}\\log\\hat\{C\}\_\{2\}is strictly decreasing, so it crosses zero exactly once\. Hence the critical point is unique\.

Closed form\.At the critical point,Δ​Vs2=−dd​s​log⁡Gavg\\frac\{\\Delta V\}\{s^\{2\}\}=\-\\frac\{d\}\{ds\}\\log G\_\{\\mathrm\{avg\}\}\. Approximating the right\-hand side by its value ats=0s=0, namelyηG/G⁡\(q\)=ηG\\eta\_\{G\}/G\(q\)=\\eta\_\{G\}, and settings∗=\(σ∗\)2s^\{\*\}=\(\\sigma^\{\*\}\)^\{2\}:

Δ​V\(s∗\)2=ηGs∗⟹s∗=Δ​VηG\.\\frac\{\\Delta V\}\{\(s^\{\*\}\)^\{2\}\}=\\frac\{\\eta\_\{G\}\}\{s^\{\*\}\}\\quad\\Longrightarrow\\quad s^\{\*\}=\\frac\{\\Delta V\}\{\\eta\_\{G\}\}\.A more careful expansion retaining thess\-dependence ofGavgG\_\{\\mathrm\{avg\}\}yields the LambertWWform \([40](https://arxiv.org/html/2609.05650#A1.E40)\) via the substitutionu=Δ​V/su=\\Delta V/sand solvingu​eu=Δ​V/ηGu\\,e^\{u\}=\\Delta V/\\eta\_\{G\}\. ∎

### A\.5Explicit Window Bounds

###### Corollary 2\(Incoherence window\)\.

Define the curiosity window as the set\{σ:C^2​\(σ\)≥12​C^2​\(σ∗\)\}\\\{\\sigma:\\hat\{C\}\_\{2\}\(\\sigma\)\\geq\\tfrac\{1\}\{2\}\\hat\{C\}\_\{2\}\(\\sigma^\{\*\}\)\\\}\. Via the Laplace relation𝔼⁡\[I\]≈σ2/\[4​η​q​\(1−q\)\]\\mathbb\{E\}\[I\]\\approx\\sigma^\{2\}/\[4\\eta q\(1\-q\)\], the window translates toα≤𝔼⁡\[I\]≤β\\alpha\\leq\\mathbb\{E\}\[I\]\\leq\\betawith

α≈Δ​V2​η​q​\(1−q\)​log⁡\(2​π​Δ​V/\(η​σ02\)\),β≈Δ​V2​η​q​\(1−q\)​log⁡\(2​G∗/Gavgunif\)\\boxed\{\\;\\alpha\\approx\\frac\{\\Delta V\}\{2\\eta q\(1\-q\)\\,\\log\\\!\\bigl\(2\\pi\\Delta V/\(\\eta\\sigma\_\{0\}^\{2\}\)\\bigr\)\},\\qquad\\beta\\approx\\frac\{\\Delta V\}\{2\\eta q\(1\-q\)\\,\\log\\\!\\bigl\(2G^\{\*\}/G\_\{\\mathrm\{avg\}\}^\{\\mathrm\{unif\}\}\\bigr\)\}\\;\}\(42\)whereσ02=1/\|V′′​\(q\)\|\\sigma\_\{0\}^\{2\}=1/\|V^\{\\prime\\prime\}\(q\)\|is the curvature scale at the attractor andG∗=G⁡\(q\)=1G^\{\*\}=G\(q\)=1\.

### A\.6Asymmetric Basins and Robustness

###### Proposition 1\(Persistence under asymmetry\)\.

For arbitraryq∈\(12,1\)q\\in\(\\tfrac\{1\}\{2\},1\)the curiosity window persists\. The optimal noise shifts as

σasym∗=σsym∗⋅1\+4​Δ2\(1−4​Δ2\)​log⁡\(Δ​V/ηG\),\\sigma^\{\*\}\_\{\\mathrm\{asym\}\}=\\sigma^\{\*\}\_\{\\mathrm\{sym\}\}\\;\\cdot\\;\\sqrt\{1\+\\frac\{4\\Delta^\{2\}\}\{\(1\-4\\Delta^\{2\}\)\\,\\log\(\\Delta V/\\eta\_\{G\}\)\}\}\\,,\(43\)and the window width scales asσβ−σα∝Δ​V/log⁡\(Δ​V\)\\sigma\_\{\\beta\}\-\\sigma\_\{\\alpha\}\\propto\\sqrt\{\\Delta V\}\\,/\\,\\log\(\\Delta V\), which is sublinear in barrier height\.

###### Proof\.

The barrier heightΔ​V​\(Δ\)=−η2​log⁡\(1−4​Δ2\)\\Delta V\(\\Delta\)=\-\\frac\{\\eta\}\{2\}\\log\(1\-4\\Delta^\{2\}\)is monotone increasing in\|Δ\|\|\\Delta\|\. Higher barriers require more noise to escape, but the overlap penalty also grows\. The structural argument of Theorem[2](https://arxiv.org/html/2609.05650#Thmtheorem2)\(iii\)—monotone\-decreasing derivative oflog⁡C^2\\log\\hat\{C\}\_\{2\}—is unchanged, so existence and uniqueness ofσ∗\\sigma^\{\*\}persist\. The quantitative shift follows from substitutingΔ​V​\(Δ\)\\Delta V\(\\Delta\)into \([40](https://arxiv.org/html/2609.05650#A1.E40)\)\. ∎

## Appendix BProof of Theorem[1](https://arxiv.org/html/2609.05650#Thmtheorem1)

###### Proof\.

Part \(i\)\.If𝒯=0\\mathcal\{T\}=0, there are no transitions, which implies single\-basin trapping and thus𝒞=0\\mathcal\{C\}=0\. IfGGfalls below a coherence threshold, the system is fragmented and𝒞=0\\mathcal\{C\}=0\. Ifℬ=0\\mathcal\{B\}=0\(no coherently occupied basin\), thenGGmust be below threshold, so𝒞=0\\mathcal\{C\}=0again\.

Part \(ii\)\.We use a classical result from measurement theory\. Definef⁡\(g,b,τ\)=F⁡\(g,b,τ\)f\(g,b,\\tau\)=F\(g,b,\\tau\)on the positive orthantℝ\>03\\mathbb\{R\}\_\{\>0\}^\{3\}\. By the monotonicity assumptions,ffis strictly increasing in each coordinate\. By Part \(i\),ffvanishes whenever any coordinate vanishes\.

Consider the level sets\{\(g,b,τ\):f⁡\(g,b,τ\)=c\}\\\{\(g,b,\\tau\):f\(g,b,\\tau\)=c\\\}forc\>0c\>0\. By strict monotonicity, each level set is a smooth surface that can be written asτ=hc​\(g,b\)\\tau=h\_\{c\}\(g,b\)withhch\_\{c\}strictly decreasing in both arguments\. The boundary conditionf→0f\\to 0as any coordinate→0\\to 0forces these level sets to be asymptotic to the coordinate planes\.

Now impose*dimensional consistency*: sinceG∈\[0,1\]G\\in\[0,1\],ℬ∈\{0,…,K\}\\mathcal\{B\}\\in\\\{0,\\ldots,K\\\}, and𝒯∈\[0,1\]\\mathcal\{T\}\\in\[0,1\]are dimensionless quantities measured on different scales,FFshould be invariant under independent rescaling of each argument’s unit\. Formally, for anyλ1,λ2,λ3\>0\\lambda\_\{1\},\\lambda\_\{2\},\\lambda\_\{3\}\>0:

F⁡\(λ1​g,λ2​b,λ3​τ\)=Φ⁡\(λ1,λ2,λ3\)​F​\(g,b,τ\)F\(\\lambda\_\{1\}g,\\;\\lambda\_\{2\}b,\\;\\lambda\_\{3\}\\tau\)=\\Phi\(\\lambda\_\{1\},\\lambda\_\{2\},\\lambda\_\{3\}\)\\;F\(g,b,\\tau\)for some functionΦ\\Phi\. By the Aczél–Dhombres theorem on the multiplicative Cauchy functional equation, the only continuous solutions are power products:

F⁡\(g,b,τ\)=C​gα1​bα2​τα3F\(g,b,\\tau\)=C\\,g^\{\\alpha\_\{1\}\}\\,b^\{\\alpha\_\{2\}\}\\,\\tau^\{\\alpha\_\{3\}\}withαi\>0\\alpha\_\{i\}\>0\(strict monotonicity\) andC\>0C\>0\. Settingφ⁡\(x\)=C​xα1\\varphi\(x\)=C\\,x^\{\\alpha\_\{1\}\}\(absorbing exponents via a monotone transformation\), we obtain𝒞=φ⁡\(G⋅ℬ⋅𝒯\)\\mathcal\{C\}=\\varphi\(G\\cdot\\mathcal\{B\}\\cdot\\mathcal\{T\}\)\.

The simplest representative—and the one we adopt as canonical—isφ=id\\varphi=\\mathrm\{id\}:

C2=G⋅ℬ⋅𝒯\.C\_\{2\}\\;=\\;G\\;\\cdot\\;\\mathcal\{B\}\\;\\cdot\\;\\mathcal\{T\}\.\(44\)Any other admissible𝒞\\mathcal\{C\}is a monotone transformation ofC2C\_\{2\}and therefore induces the same ordering over system states\. ∎

Similar Articles

Large-scale study of curiosity-driven learning

OpenAI Blog

OpenAI presents a large-scale empirical study of curiosity-driven reinforcement learning without extrinsic rewards across 54 benchmark environments, showing strong performance and investigating the role of feature spaces in prediction-based reward signals.

Look Before You Leap: Autonomous Exploration for LLM Agents

Hugging Face Daily Papers

This paper identifies autonomous exploration as a critical capability for LLM agents and proposes the Explore-then-Act paradigm, which decouples information gathering from task execution to improve adaptability and real-world performance. It also introduces Exploration Checkpoint Coverage as a verifiable metric for evaluating exploration breadth.