Interoceptive Attention as Dynamic Homeostatic Prioritization in a Foraging Agent

arXiv cs.AI Papers

Summary

This paper studies how an agent with limited perceptual bandwidth should allocate interoceptive precision across bodily needs in a foraging task, showing that dynamically attending to the most-needed channel improves survival under a fixed precision budget.

arXiv:2608.04232v1 Announce Type: new Abstract: Biological systems must regulate competing needs under limited perceptual bandwidth, where sharpening one estimate costs the capacity to sharpen the others. Any fixed-budget system therefore has to decide where to allocate its perceptual precision. We study this in a foraging agent that must keep several bodily needs satisfied to survive, modelled with active inference. At each step it reads its own body-state beliefs, identifies the most-needed channel, and reallocates a fixed budget of interoceptive precision toward it, so that the same precision-shaped likelihood feeds both belief update and planning. In AffectWorld, a four-channel foraging gridworld, this selective allocation more than doubles learning-phase survival at matched budget against a uniform-precision agent ($0.414$ vs $0.199$ across 11 layouts, $n{=}32$ seeds each, paired cluster-bootstrap $p \leq 10^{-4}$). Two further results sharpen the mechanism. The benefit runs through planning as well as perception, since denying the shaped likelihood to the planner alone removes about half of it. It is also need-aligned, since aiming precision at the least-needed channel does worse than spreading it evenly. The attended channel additionally learns its own dynamics about twice as fast, and stays ahead even at matched observation count, a behavioural trace of the same precision routing, visible in learning speed, not survival.
Original Article
View Cached Full Text

Cached at: 08/06/26, 07:41 AM

# Interoceptive Attention as Dynamic Homeostatic Prioritization in a Foraging Agent
Source: [https://arxiv.org/html/2608.04232](https://arxiv.org/html/2608.04232)
11institutetext:Dept\. of Mathematics & Applied Mathematics, Univ\. of Cape Town, South Africa22institutetext:Neuroscience Institute, Univ\. of Cape Town, South Africa33institutetext:Laureate Institute for Brain Research, Tulsa, OK, USA44institutetext:MIND Institute and CSAM, University of the Witwatersrand, South Africa55institutetext:INRS, Montreal, Canada66institutetext:Delft University of Technology, Dept\. of Cognitive RoboticsNicolas KuskeEvert A\. BoonstraBruce A\. BassettCharel van HoofRowan HodsonBenjamin RosmanRyan SmithMark SolmsCo\-senior authors\.Jonathan P\. Shock††footnotemark:

###### Abstract

Biological systems must regulate competing needs under limited perceptual bandwidth, where sharpening one estimate costs the capacity to sharpen the others\. Any fixed\-budget system therefore has to decide where to allocate its perceptual precision\. We study this in a foraging agent that must keep several bodily needs satisfied to survive, modelled with active inference\. At each step it reads its own body\-state beliefs, identifies the most\-needed channel, and reallocates a fixed budget of interoceptive*precision*toward it, so that the same precision\-shaped likelihood feeds both belief update and planning\. InAffectWorld, a four\-channel foraging gridworld, this selective allocation more than doubles learning\-phase survival at matched budget against a uniform\-precision agent \(0\.4140\.414vs0\.1990\.199across 11 layouts,n=32n\{=\}32seeds each, paired cluster\-bootstrapp≤10−4p\\leq 10^\{\-4\}\)\. Two further results sharpen the mechanism\. The benefit runs through planning as well as perception, since denying the shaped likelihood to the planner alone removes about half of it\. It is also need\-aligned, since aiming precision at the least\-needed channel does worse than spreading it evenly\. The attended channel additionally learns its own dynamics about twice as fast, and stays ahead even at matched observation count, a behavioural trace of the same precision routing, visible in learning speed, not survival\.

## 1Introduction

A foraging animal must keep several bodily needs satisfied at once\. It has to eat, drink, and breathe, and letting any one of these lapse too far is fatal\. Faced with many needs and one body, the animal must continually decide what to do next, and behavioural ecology and the animat tradition have long studied this as a problem of*action selection*among competing needs\[[1](https://arxiv.org/html/2608.04232#bib.bib1),[2](https://arxiv.org/html/2608.04232#bib.bib2),[3](https://arxiv.org/html/2608.04232#bib.bib3),[4](https://arxiv.org/html/2608.04232#bib.bib4)\]\. A prior question has drawn less attention\. Before choosing an action, the animal must decide which of its internal signals to trust, because it cannot monitor every need equally well at every moment, and attending closely to one signal leaves less capacity for the rest\. This paper is about that perceptual choice: how an agent should allocate limited sensory precision across competing interoceptive channels, and what such allocation buys it\.

The quantity being allocated is*precision*: the weight an agent gives a signal when deciding how much to trust it\. In predictive coding and active inference, precision is the confidence assigned to a prediction error, and attention is modelled as control over it\[[5](https://arxiv.org/html/2608.04232#bib.bib5),[6](https://arxiv.org/html/2608.04232#bib.bib6),[7](https://arxiv.org/html/2608.04232#bib.bib7)\]\. Cognitive science studies the same trade\-off as bounded or resource\-rational attention\[[8](https://arxiv.org/html/2608.04232#bib.bib8)\], a limited processing budget spread across inputs\. Because that budget is fixed, trusting one signal more means trusting the others less\. In our model this allocation plays two roles, and we treat them separately: it changes how clearly the agent senses whichever need is most urgent, and it changes which actions the agent then judges worthwhile\.

This could be pursued in several modelling frameworks\. Reinforcement learning, for instance, can be given interoceptive state and a homeostatic reward\. We use active inference for two reasons\. Precision is already a first\-class quantity in it, so the mechanism we care about is native rather than added on\. And it has become a standard setting for biologically motivated, Bayesian models of decision making, including homeostatic and interoceptive control\[[9](https://arxiv.org/html/2608.04232#bib.bib9),[10](https://arxiv.org/html/2608.04232#bib.bib10),[11](https://arxiv.org/html/2608.04232#bib.bib11)\]\. In active inference\[[12](https://arxiv.org/html/2608.04232#bib.bib12),[13](https://arxiv.org/html/2608.04232#bib.bib13)\]the agent chooses actions to make its observations match its predictions, under a single objective defined over both its beliefs and its plans\. In our model the precision\-shaped likelihood enters both the belief update and the evaluation of plans \(Fig\.[1](https://arxiv.org/html/2608.04232#S1.F1)a\), so one precision choice acts on perception and decision at the same time\. Whether that dual action helps, and where, is a question our propagation analysis addresses\.

Our contribution sits between three lines of work\. Homeostatic reinforcement learning integrates bodily state through a hand\-designed reward\[[14](https://arxiv.org/html/2608.04232#bib.bib14),[15](https://arxiv.org/html/2608.04232#bib.bib15)\], but routes behaviour through reward rather than precision\. Active\-inference models of interoception use precision to regulate homeostatic control\[[9](https://arxiv.org/html/2608.04232#bib.bib9),[10](https://arxiv.org/html/2608.04232#bib.bib10)\], but not as a variable the agent dynamically reallocates across competing interoceptive channels\. Active\-inference treatments of attention as precision have mostly addressed a single exteroceptive target, such as visual search\[[7](https://arxiv.org/html/2608.04232#bib.bib7),[16](https://arxiv.org/html/2608.04232#bib.bib16)\], or cast the control of precision as an action within a hierarchical model\[[17](https://arxiv.org/html/2608.04232#bib.bib17)\], rather than spreading a fixed budget across several interoceptive needs\. We combine these: a fixed interoceptive precision budget, routed step by step by the agent’s own belief about which need is most urgent, and used at both the perception and planning stages\.

We ask a direct question: does pointing a limited precision budget at whichever need is currently most urgent help an agent survive and learn, compared with spreading precision evenly? We test this inAffectWorld, a foraging gridworld with several competing needs, and the answer is yes\. Selective allocation roughly doubles learning\-phase survival against a uniform\-precision agent at the same budget \(Fig\.[1](https://arxiv.org/html/2608.04232#S1.F1)b\)\. The*direction*of allocation is what matters\. An agent that instead sharpens its least\-needed channel does worse than uniform, so the gain comes from tracking need, not from unevenness on its own\. Two further results locate the effect\. The precision\-shaped signal helps at both the perception and planning stages, and denying it to the planner alone costs about half the benefit\. The attended channel also learns its own dynamics about twice as fast, staying ahead even at matched observation count, a behavioural difference between the two agents that gives another view of the mechanism at work\.

![Refer to caption](https://arxiv.org/html/2608.04232v1/x1.png)\(a\)
![Refer to caption](https://arxiv.org/html/2608.04232v1/x2.png)\(b\)

Figure 1:Architectural intervention and headline result\.\(a\)One shaped likelihood fieldA\(m\)A^\{\(m\)\}with two downstream consumers, the belief update and the EFE planner, under the budget∑mκm≤K=2\.60\\sum\_\{m\}\\kappa\_\{m\}\\leq K\{=\}2\.60\. Suffocation is tied to water\.\(b\)Across 11 layouts \(n=32n=32seeds\), selective allocation more than doubles \(∼2\.08×\{\\sim\}2\.08\{\\times\}\) uniform survival\. The anti\-aligned control reverses the gain\. Oracle ceiling: a planner given the true environment model\.
## 2Mechanism and Setting

![Refer to caption](https://arxiv.org/html/2608.04232v1/x3.png)Figure 2:Theκ\\kappa\-attention loop\.The agent observes its position, the local resource, and four interoceptive channels, updates its body\-state belief, and attends the most\-needed one\. That choice sets a fixed precision budgetKKwhose shaped likelihoodA\(m\)A^\{\(m\)\}enters both the belief update and the EFE planner \(dashed\) before the planner acts and closes the loop\.### 2\.1Active Inference under Fixed Interoceptive Precision

The agent is an active\-inference agent: it maintains beliefs about its bodily and world state and acts to make its observations match its predictions \(Fig\.[2](https://arxiv.org/html/2608.04232#S2.F2)\)\. Formally, it maintains a factorised partially observable Markov decision process \(POMDP\), a natural model class for an agent whose interoceptive observations are noisy reports of underlying physiological states, and selects policies by minimising expected free energy\[[12](https://arxiv.org/html/2608.04232#bib.bib12),[13](https://arxiv.org/html/2608.04232#bib.bib13)\]\. In its joint form

𝒢​\(π\)=𝔼q​\(s,o∣π\)​\[log⁡q​\(s∣π\)−log⁡P​\(s,o∣C\)\],\\mathcal\{G\}\(\\pi\)\\;=\\;\\mathbb\{E\}\_\{q\(s,o\\mid\\pi\)\}\\\!\\bigl\[\\log q\(s\\mid\\pi\)\-\\log P\(s,o\\mid C\)\\bigr\],
this decomposes into a*risk*term𝔼q​\(o∣π\)\[DKL\(q\(o∣π\)∥P\(o∣C\)\)\]\\mathbb\{E\}\_\{q\(o\\mid\\pi\)\}\\\!\\bigl\[D\_\{\\mathrm\{KL\}\}\\\!\\bigl\(q\(o\\mid\\pi\)\\,\\\|\\,P\(o\\mid C\)\\bigr\)\\bigr\]\(divergence between predicted and preferred observations\) and an*ambiguity*term𝔼q​\(s∣π\)​\[ℋ​\[P​\(o∣s\)\]\]\\mathbb\{E\}\_\{q\(s\\mid\\pi\)\}\\\!\\bigl\[\\mathcal\{H\}\[P\(o\\mid s\)\]\\bigr\]\(expected entropy of the observation likelihood\)\. In words, a policy scores well when it is expected to move the body toward its preferred replete states \(low risk\) while keeping observations informative about body state \(low ambiguity\)\. The preference distributionP​\(o∣C\)P\(o\\mid C\)is a soft prior concentrated near the full body\-state level \(s=smaxs\{=\}s\_\{\\max\}\) on each channel\. Policies are sampled fromP​\(π\)∝exp⁡\(−γ​𝒢​\(π\)\)P\(\\pi\)\{\\propto\}\\exp\(\-\\gamma\\mathcal\{G\}\(\\pi\)\)withγ=16\\gamma\{=\}16fixed throughout\. Precisionκ\\kappaenters both EFE terms through the likelihoodA\(m\)A^\{\(m\)\}defined in Sec\.[2\.2](https://arxiv.org/html/2608.04232#S2.SS2)\. The agent’s generative model has two parts: a fixed body\-transition priorBB\(matching biology’s access to body dynamics through proprioception and prior experience\) and a learned observation likelihoodAAmapping body and world states to observations\.AAis updated online via Dirichlet pseudo\-counts on each observation \(Sec\.[3\.5](https://arxiv.org/html/2608.04232#S3.SS5)\)\. Its Dirichlet prior has a concentrationα0\\alpha\_\{0\}, and a largerα0\\alpha\_\{0\}makes the model more rigid, so observations move the body\-state posterior less\. The agent therefore learns what each interoceptive reading reports about its true body state and what each exteroceptive reading reports about the grid\. Actions are the four cardinal moves plus a no\-op \(stay\-in\-place\)\. Resource consumption on a tile is automatic\.

The body produces a noisy categorical observationom\(i\)∈\{0,…,5\}o^\{\(i\)\}\_\{m\}\\in\\\{0,\\ldots,5\\\}on each ofM=4M\{=\}4interoceptive channelsm∈\{1,…,4\}m\\in\\\{1,\\ldots,4\\\}\(three active needs and one inert control\) at observation precisionκm\\kappa\_\{m\}\. Concretely,κm\\kappa\_\{m\}is the probability that the per\-step observation equals the underlying body\-state level on channelmm— the diagonal of the channel’s likelihood matrix isκm\\kappa\_\{m\}, with each “wrong” read drawn uniformly over the five other levels:

Ao,s\(m\)=\{κmif​o=s,\(1−κm\)/5otherwise,o,s∈\{0,…,5\},A^\{\(m\)\}\_\{o,\\,s\}\\;=\\;\\begin\{cases\}\\kappa\_\{m\}&\\text\{if \}o=s,\\\\\[2\.0pt\] \(1\-\\kappa\_\{m\}\)/5&\\text\{otherwise,\}\\end\{cases\}\\qquad o,\\,s\\in\\\{0,\\ldots,5\\\},so each column ofA\(m\)A^\{\(m\)\}is a valid categorical distribution andκm∈\[0,1\]\\kappa\_\{m\}\\in\[0,1\]is the per\-step probability that the observation reports the true level on channelmm\. Total interoceptive precision is constrained to a fixed soft budget,111We setK=2\.60K=2\.60so that the attended channel is informative while the others are not blind\. Sec\.[3\.4](https://arxiv.org/html/2608.04232#S3.SS4)sweepsKKover\{1\.5,…,4\.0\}\\\{1\.5,\\ldots,4\.0\\\}\.

∑m=14κm≤K,κm∈\[κfloor,1\],K=2\.60,κfloor=0\.05,\\sum\_\{m=1\}^\{4\}\\kappa\_\{m\}\\,\\leq\\,K,\\qquad\\kappa\_\{m\}\\in\[\\kappa^\{\\text\{floor\}\},\\,1\],\\qquad K=2\.60,\\quad\\kappa^\{\\text\{floor\}\}=0\.05,where the floor prevents any channel from being completely silenced and the upper boundκm≤1\\kappa\_\{m\}\\leq 1is required forAo,s\(m\)A^\{\(m\)\}\_\{o,s\}to remain a valid probability\. The allocation procedure clips per channel,κm←clip​\(κ~m,κfloor,1\)\\kappa\_\{m\}\\leftarrow\\mathrm\{clip\}\(\\tilde\{\\kappa\}\_\{m\},\\kappa^\{\\text\{floor\}\},1\), and any residual budget is absorbed into the constraint∑mκm≤K\\sum\_\{m\}\\kappa\_\{m\}\\leq K\. Uniform allocation givesκm=K/4=0\.65\\kappa\_\{m\}=K/4=0\.65on every channel — each channel reports correctly65%65\\%of the time\. Selective allocation givesκatt=0\.90\\kappa\_\{\\text\{att\}\}=0\.90\(90%90\\%correct on the attended channel\) andκun=\(K−κatt\)/3≈0\.567\\kappa\_\{\\text\{un\}\}=\(K\-\\kappa\_\{\\text\{att\}\}\)/3\\approx 0\.567\(56\.7%56\.7\\%correct on each of the other three\), preserving the sum\.

### 2\.2Theκ\\kappa\-Attention Mechanism

We use*attention*here in a narrow, operational sense\. It means the agent’s ongoing allocation of a fixed interoceptive precision budget across channels, set by its own belief about which need is most urgent, and it is a computational abstraction rather than a model of a specific neural system\. In the regime we study the body\-state posterior stays well\-calibrated to true need, so the allocation tracks genuine urgency\.

The selector reads only the agent’s own posterior beliefsq​\(sm\(i\)\)q\(s^\{\(i\)\}\_\{m\}\), never ground\-truth body counters\. The body observation model is itself learned online, so body beliefs remain uncertain throughout\. At each step the agent picks an attended channelmt⋆m^\{\\star\}\_\{t\}from its own posterior over body states,mt⋆=arg⁡maxm⁡𝔼q​\(sm\(i\)\)​\[needm\]m^\{\\star\}\_\{t\}=\\arg\\max\_\{m\}\\mathbb\{E\}\_\{q\(s^\{\(i\)\}\_\{m\}\)\}\[\\mathrm\{need\}\_\{m\}\], and reallocates the precision budget to give that channel the higher precisionκatt\\kappa\_\{\\text\{att\}\}, so the agent sharpens its interoceptive sense on whichever need it currently believes is most pressing\. The scalarneedm​\(s\)=\(smax−s\)/smax\\mathrm\{need\}\_\{m\}\(s\)=\(s\_\{\\max\}\-s\)/s\_\{\\max\}for the three active channels \(full body\-state level→0\\to 0, empty→1\\to 1\), and returns a constant0on the inert control channel\. Ties are broken by lowest channel index\. The default rule \(need\-aligned, picking the most\-needed channel\) is used in all main results\. Alternative criteria \(action\-aware, explorative, hysteresis, anti\-aligned\) appear as ablations in Sec\.[3\.2](https://arxiv.org/html/2608.04232#S3.SS2), along with a ground\-truth selector \(reading true body state\) as a directional ceiling on a single\-layout pilot \(n=3n=3\)\.

Theκ\\kappaallocation could act at three sites in the agent: \(i\)*belief update*\(A\(m\)A^\{\(m\)\}enters per\-step state inference and breaks posterior symmetry in a need\-correlated way\), \(ii\)*policy evaluation*\(the sameA\(m\)A^\{\(m\)\}enters EFE and sharpens the predicted\-vs\-preferred KL on that channel\), and \(iii\)*likelihood learning*\(cleaner observations on the attended channel produce more concentrated Dirichlet updates of theA\(m\)A^\{\(m\)\}pseudo\-counts\)\. We treat \(iii\) as an observed consequence of \(i\), since the update rule itself is unchanged across all agents, and test the contribution of \(i\) versus \(ii\) in Sec\.[3\.3](https://arxiv.org/html/2608.04232#S3.SS3)\.

### 2\.3Environment and Agents

##### Environment\.

AffectWorldis a6×66\{\\times\}6gridworld with two food and two water tiles per layout\. Agents receive two exteroceptive observations \(position and resource\) and three interoceptive observations, one per active body channel \(hunger, thirst, suffocation\)\. Episodes run up to 60 steps, terminating early on death\. We use 12 distinct environment layouts across easy, medium and far tiers \(Fig\.[3](https://arxiv.org/html/2608.04232#S2.F3)\)\. The headline results use 11 \(the excludedL01is a mechanistic stress\-test\)\.

##### Body Channels\.

Hunger, thirst and suffocation are three distinct needs, each starting the trial in the*full*state\. Hunger and thirst are relieved by the food and water tiles respectively, decay one unit/step, and are lethal at zero\. Suffocation depletes on water tiles, recovers elsewhere, and is non\-lethal: zero suffocation lowers the value the agent assigns to staying on water, but does not terminate the trial\. This gives a need the agent should prioritise during water\-seeking without introducing a second mortality route alongside hunger and thirst\. The fourth channel is an inert control: its underlying signal is constant, so it carries no need information and the need\-aligned selector has no reason to attend it\. It anchors the primary attentive agent\-vs\-uniform agent comparison: a channel that reports nothing about any need cannot by itself create the survival gap between those two agents\.

![Refer to caption](https://arxiv.org/html/2608.04232v1/x4.png)Figure 3:Representative easy, medium, and far tier layouts ofAffectWorld, with start\-to\-resource distances annotated\. The channels\-to\-resource mapping and budget split are summarised in the Sec\.[2\.3](https://arxiv.org/html/2608.04232#S2.SS3)prose\.
##### Agents\.

The three core agents share planner, priors, Dirichlet hyperparameters, body model and exteroceptive observation structure, and differ only inκ\\kappaallocation\. The planner minimises EFE over the full policy space at planning horizonH=3H\{=\}3throughout\. The uniform agent holdsκ=0\.65\\kappa=0\.65uniformly across all four channels\. This fixes the total precision budget atK=M×0\.65=2\.60K=M\\times 0\.65=2\.60forM=4M\{=\}4channels, and the attentive and reversed agents redistribute this same total, so all three agents are compared at an identical budget\. The attentive agent runs dynamicκ\\kappawithκatt=0\.90\\kappa\_\{\\text\{att\}\}=0\.90on whichever channel the body\-state belief currently flags as most\-needed andκun=0\.567\\kappa\_\{\\text\{un\}\}=0\.567on the other three, preserving the fixed budget\. The anti\-aligned control uses the sameκ\\kappasplit and same budget but selects the*least*\-needed channel, a direction\-flipped control\.

##### Metrics and Grouping\.

A*trial*is one episode \(up to 60 steps\), scored11if the agent survives to the6060\-step horizon and0if it dies first\. A*run*is one seed’s sequence of100100trials on a fixed layout, over which the likelihoodAAis learned online\. Each \(layout, agent\) cell usesn=32n\{=\}32seeds unless a caption states otherwise\.*Learning\-phase survival*is the mean trial outcome over a run’s100100trials, pooled across seeds and layouts\.*Plateau survival*restricts that mean to trials1010–2020, once both agents reach steady state\. The headline result pools the 11 layouts across easy, medium and far tiers\. The*easy\-tier*grouping is the five remaining easy\-tier layouts\.

## 3Results

### 3\.1Does Selective Precision Beat Uniform?

Table 1:Learning\-phase survival across 11 layouts \(L01excluded\),n=32n\{=\}32seeds per cell\. Cluster\-bootstrap95%95\\,\\%CIs and raw paired\-bootstrappp\-values \(Holm\-corrected in Sec\.[5](https://arxiv.org/html/2608.04232#S5)\)\.333Code, layout banks, and analysis pipelines for this paper are available at[https://github\.com/sgrimbly/attention\-aif\-sab2026\-snapshot](https://github.com/sgrimbly/attention-aif-sab2026-snapshot), where the supplementary appendices provide fuller experiments, results, and interpretations\.Replacing uniformκ\\kappawith a body\-state\-driven selector at the same budget more than doubles learning\-phase survival \(∼2\.08×\{\\sim\}2\.08\\times, paired\-bootstrapp≤10−4p\\leq 10^\{\-4\}, Table[3](https://arxiv.org/html/2608.04232#footnote3)\)\. The anti\-aligned control, which uses the same budget and allocation magnitudes but flips the selector direction, performs significantly*worse*than baseline\. The task is hard in this setting\. Even the oracle planner, given the true environment dynamics, reaches only∼0\.66\{\\sim\}0\.66survival, so we treat it as a performance ceiling rather than a competing baseline\.

### 3\.2Does the Direction of Allocation Matter?

![Refer to caption](https://arxiv.org/html/2608.04232v1/x5.png)Figure 4:Every need\-aligned selector beats the uniform agent and only the anti\-aligned rule fails, so the advantage comes from the direction of allocation, not from unevenness alone \(33layouts×\\times88seeds×\\times3030trials, plateau survival, bootstrap95%95\\,\\%CIs\)\.We probe selector choice and allocation magnitude jointly \(Fig\.[4](https://arxiv.org/html/2608.04232#S3.F4)\)\. Four need\-aligned criteria plus a ground\-truth selector substantially beat the uniform agent\. Only the direction\-flipped anti\-aligned rule fails\. The mechanism is robust to the specific criterion as long as the selector is direction\-aligned\.

Direction is what does the work, not allocation magnitude alone\. Sweepingκatt\\kappa\_\{\\text\{att\}\}jointly with the selector at fixedKK, the need\-aligned rule climbs monotonically with asymmetry while the anti\-aligned rule stays flat near0\.360\.36\. The two converge to baseline at uniform allocation \(κatt=0\.65\\kappa\_\{\\text\{att\}\}=0\.65\) and diverge thereafter, reaching a\+44\+44pp gap at the defaultκatt=0\.90\\kappa\_\{\\text\{att\}\}=0\.90and\+46\+46pp atκatt=0\.99\\kappa\_\{\\text\{att\}\}=0\.99\.

### 3\.3Where Does the Precision Signal Act?

A single precision\-shaped likelihood matrixA\(m\)A^\{\(m\)\}is read by two parts of the agent, the per\-step belief update and the EFE planner\. To test how much of the gain rides on the planner also seeing the shaped likelihood, we run an inference\-only ablation variant that gives the planner an*unshaped*likelihood, withA\(m\)A^\{\(m\)\}rebuilt using the uniform allocationκm=K/M\\kappa\_\{m\}=K/Mon every channel so that likelihood precision is decoupled from posterior urgency at the planning stage\. A drop here is the gain that the planner’s shaped likelihood was supplying\.

At the default prior concentration \(α0=0\.1\\alpha\_\{0\}\{=\}0\.1\) the inference\-only ablation loses2020pp relative to the full attentive agent\. The inversion begins atα0≥1\\alpha\_\{0\}\{\\geq\}1, and byα0=10\\alpha\_\{0\}\{=\}10the loss reaches8888pp, with the inference\-only ablation collapsing to the level of the direction\-flipped anti\-aligned control\. When the body\-state posterior is dominated by a rigid prior, only the planner’s use of the shaped likelihood can push policies away from the prior’s preferred actions\. The inference site alone cannot\. This refines the active\-inference picture of attention as policy\-precision modulation\[[7](https://arxiv.org/html/2608.04232#bib.bib7),[6](https://arxiv.org/html/2608.04232#bib.bib6),[10](https://arxiv.org/html/2608.04232#bib.bib10)\]\. State\-dependent precision modulation of the observation likelihood, formalised for exteroceptive visual search\[[16](https://arxiv.org/html/2608.04232#bib.bib16)\], here acts on multi\-channel interoceptive precision under a fixed budget and propagates from inference into the EFE planner\.

The mirror ablation \(*planning\-only*, with inference\-stage shaping disabled and the planner keeping the shapedAA\) on the same five easy\-tier layouts reaches pooled survival0\.8650\.865, not significantly above the full model’s easy\-tier0\.6960\.696given the smaller planning\-only sample \(overlapping CIs\)\. At loose priors the planning\-stage pathway therefore carries the gain on its own, and the inference\-stage shaping is at most neutral here\. Combined with the inference\-only ablation collapse at rigid priors, this localises the dominant pathway in the planner, with inference\-stage shaping becoming necessary only once the prior is strong enough to override state inference\.

### 3\.4Is the Advantage Robust to Budget and Prior Rigidity?

The benefit holds across the three parameters that most plausibly drive it: prior rigidityα0\\alpha\_\{0\}, attended\-channel precisionκatt\\kappa\_\{\\text\{att\}\}\(Sec\.[3\.2](https://arxiv.org/html/2608.04232#S3.SS2)\), and budgetKK\. Across Dirichlet prior concentrationα0∈\{10−3,…,102\}\\alpha\_\{0\}\\in\\\{10^\{\-3\},\\ldots,10^\{2\}\\\}, attentive agent maintains∼0\.85\{\\sim\}0\.85plateau survival while uniform agent collapses atα0≥10\\alpha\_\{0\}\\geq 10\.

The budget sweep is more nuanced\. AcrossK∈\{1\.5,…,4\.0\}K\\in\\\{1\.5,\\ldots,4\.0\\\}, attentive agent beats uniform agent by3232–5656pp at every testedKK\. Both peak near the canonicalK=2\.60K\{=\}2\.60and decline at highKK\. Selective allocation givesκatt=0\.90\\kappa\_\{\\text\{att\}\}\{=\}0\.90on the attended channel andκun=\(K−κatt\)/3\\kappa\_\{\\text\{un\}\}\{=\}\(K\-\\kappa\_\{\\text\{att\}\}\)/3on the other three\. AtK=4K\{=\}4the cap binds, so the direction claim is cleanly identified at the canonicalK=2\.60K\{=\}2\.60where no channel saturates \(Sec\.[3\.2](https://arxiv.org/html/2608.04232#S3.SS2)\)\. The baseline’s decline at highKKis a planner\-overcommitment effect that selective allocation avoids by keeping the policy posterior softer\.

To see what the dynamic selector adds, we compare it against a fixed\-channel control that alwaysκ\\kappa\-shapes one channel\. On easy\-tier layouts, where hunger dominates the need landscape, always\-attend\-hunger ties dynamic attentive agent \(both0\.830\.83\)\. The dynamic selector pulls ahead in two regimes\. Under rigid priors \(α0≥10\\alpha\_\{0\}\\geq 10\), fixed\-channel modes pool to≈0\.65\{\\approx\}0\.65while attentive agent reaches≈0\.90\{\\approx\}0\.90\. On a forced\-multi\-need setting where food\- and water\-need alternate in dominance, dynamic beats always\-attend\-hunger by pooled\+23\.5\+23\.5pp, since no single fixed channel tracks the switching need\. The per\-observation learning difference on the currently\-most\-needed channel \(Sec\.[3\.5](https://arxiv.org/html/2608.04232#S3.SS5)\) is what the dynamic selector buys in both regimes\.

The advantage also survives environmental non\-stationarity: a tile\-mutation sweep confirms it in sign across mutation rates0\.020\.02–0\.100\.10, though it weakens at high rates\.

### 3\.5Does the Attended Channel Also Learn Faster?

Selective allocation also makes the agent learn the body channel it attends to faster\. Over 50 trials, attentive agent’s hunger model converges about2\.4×2\.4\\timesfaster than uniform agent’s, and this is a per\-observation effect rather than the result of a surviving agent gathering more data\. Plotting hunger\-model accuracy against cumulative observation count rather than trial index, attentive agent sits∼0\.31\{\\sim\}0\.31above uniform agent at every matched observation level\. The anti\-aligned control anti\-aligned control sits∼0\.22\{\\sim\}0\.22above, so any non\-uniform allocation buys some acceleration, but attentive agent still beats anti\-aligned control by a further∼0\.09\{\\sim\}0\.09, so direction matters per observation as well as per trial\.

This gap is what we would expect ifκ\\kappasharpens the Dirichlet update at the update step itself, so changing precision shows up not just in survival but in how fast each channel is learned\. We treat this as an interesting behavioural difference between the two agents rather than a decisive test of mechanism\. A preference\-reweighting scheme that changed only which channel the planner pursues, without sharpening the per\-step likelihood, might behave differently here, and settling that needs a matched comparator we leave to future work\.

## 4Discussion

Selective interoceptive precision is, on its own, enough to prioritise between competing needs, and the attentive agent’s roughly twofold survival advantage over the uniform agent at matched budget establishes that\. The propagation, direction, and per\-observation learning results then locate where the signal acts and are consistent with it leaving a behavioural trace beyond survival \(Sec\.[3\.5](https://arxiv.org/html/2608.04232#S3.SS5)\)\. The robustness of that advantage comes from what drives the allocation\. Becauseκ\\kappais set by the body\-state belief, which updates independently of the world model, the prioritisation signal survives both rigid\-prior world\-model collapse and spatial\-map non\-stationarity\. The body\-signal\-free uniform agent tolerates neither\. Theκ\\kappa\-routing here is hand\-specified, and whether a learned router converges to the same need\-aligned policy is the natural follow\-up\.

The work draws on two literatures at once, attention\-as\-precision in active inference\[[7](https://arxiv.org/html/2608.04232#bib.bib7),[13](https://arxiv.org/html/2608.04232#bib.bib13),[16](https://arxiv.org/html/2608.04232#bib.bib16)\]and active\-interoceptive inference\[[10](https://arxiv.org/html/2608.04232#bib.bib10),[18](https://arxiv.org/html/2608.04232#bib.bib18),[9](https://arxiv.org/html/2608.04232#bib.bib9),[11](https://arxiv.org/html/2608.04232#bib.bib11),[19](https://arxiv.org/html/2608.04232#bib.bib19)\]\. Within the first, the shaped likelihood matters at the planning stage and not only at perception, since feeding it into the EFE planner accounts for about half the benefit at the default prior and almost all of it under rigid priors\. A complementary line extends active\-inference agents at the planning layer through active learning\[[20](https://arxiv.org/html/2608.04232#bib.bib20)\]\. We make the parallel extension at the precision\-allocation layer\.

Homeostatic RL\[[14](https://arxiv.org/html/2608.04232#bib.bib14),[15](https://arxiv.org/html/2608.04232#bib.bib15),[21](https://arxiv.org/html/2608.04232#bib.bib21)\]integrates interoceptive state through a hand\-designed homeostatic reward, a connection Keramati and Gutkin drew explicitly to active inference and interoceptive surprise\. Our contribution routes behaviour through per\-step observation precision rather than through that reward\.

We offer one biological reading tentatively, as a way of thinking about the model rather than a claim about the brain\. It is tempting to relate theκ\\kappaallocation to the idea that interoceptive processing is gain\-modulated by bodily urgency\[[10](https://arxiv.org/html/2608.04232#bib.bib10),[22](https://arxiv.org/html/2608.04232#bib.bib22),[23](https://arxiv.org/html/2608.04232#bib.bib23)\]\. If so, blunting that gain might impair need prioritisation most while an agent is still learning an unfamiliar environment, where the attentive agent–uniform agent gap of Table[3](https://arxiv.org/html/2608.04232#footnote3)is largest\. We intend this as a hypothesis\-generating analogy, not a specific neuroanatomical or clinical prediction\.

## 5Limitations and Scope

Our experiments vary the routing with the selector fixed\. The claim is therefore that selective interoceptive precision is*one sufficient implementation*of the prioritisation mechanism, not that it is uniquely identified relative to alternative actuation sites for the same selector\. A matched preference\-reweighting comparator is the natural next experiment\. It would test whether the learning difference of Sec\.[3\.5](https://arxiv.org/html/2608.04232#S3.SS5)is driven by per\-observation Dirichlet acceleration rather than by planning\-action selection\. The selector also assumes the body\-state posterior is well\-calibrated to true need\. Under strong miscalibration the same mechanism could misroute precision, which our fixed\-selector design does not test\.

The empirical scope is bounded \(layout bank, depletion\-rate range, planning horizon, channel count, discrete grid\)\.AffectWorldis a deliberately minimal gridworld, so the biological readings above are offered as model\-generated hypotheses about mechanism, not as claims that scale to the full complexity of biological interoception or to embodied intelligence more broadly\. The attentive agent\-vs\-uniform agent contrasts pass atα=0\.01\\alpha\{=\}0\.01with a comfortable margin\.

## 6Conclusion

We asked whether an agent should point a limited perceptual budget at whichever need is most urgent, and what such routing buys it\. The answer is that selective interoceptive precision, driven by the agent’s own belief about which need is most pressing, is by itself enough to prioritise between competing needs, roughly doubling learning\-phase survival over a uniform\-precision agent at the same budget\. Two features make the result specific rather than generic\. The gain depends on allocating precision at a genuine need, since reversing the direction does worse than spreading precision evenly\. It works through the planner as much as through perception, because the shaped likelihood must reach policy evaluation for most of the benefit to appear\.

\{credits\}

#### 6\.0\.1Acknowledgements

Funded by the Oppenheimer Memorial Trust \(UCT Neuroscience Institute, 475201 NSI1006\) and Conscium, Ltd \(UCT Psychology, PSY526 428430\)\. EAB: NWO grant 019\.223SG\.002\. RS: Laureate Institute for Brain Research\. Computations: UCT ICTS HPC\. SG used Claude \(Anthropic\) for coding and manuscript editing\.

#### 6\.0\.2Disclosure of Interests

The authors have no competing interests to declare\.

## References

- \[1\]McNamara, J\.M\., Houston, A\.I\.: The common currency for behavioral decisions\. The American Naturalist127\(3\), 358–378 \(1986\)
- \[2\]Maes, P\.: A bottom\-up mechanism for behavior selection in an artificial creature\. In: From Animals to Animats: Proceedings of the First International Conference on Simulation of Adaptive Behavior \(SAB\)\. pp\. 238–246\. MIT Press \(1991\)
- \[3\]Cañamero, L\.: Designing emotions for activity selection in autonomous agents\. In: Trappl, R\., Petta, P\., Payr, S\. \(eds\.\) Emotions in Humans and Artifacts, pp\. 115–148\. MIT Press, Cambridge, MA \(2003\)
- \[4\]Lewis, M\., Cañamero, L\.: Hedonic quality or reward? a study of basic pleasure in homeostasis and decision making of a motivated autonomous robot\. Adaptive Behavior24\(5\), 267–291 \(2016\)
- \[5\]Bastos, A\.M\., Usrey, W\.M\., Adams, R\.A\., Mangun, G\.R\., Fries, P\., Friston, K\.J\.: Canonical microcircuits for predictive coding\. Neuron76\(4\), 695–711 \(2012\)
- \[6\]Feldman, H\., Friston, K\.J\.: Attention, uncertainty, and free\-energy\. Frontiers in Human Neuroscience4, 215 \(2010\)
- \[7\]Parr, T\., Friston, K\.J\.: Uncertainty, epistemics and active inference\. Journal of the Royal Society Interface14\(136\), 20170376 \(2017\)
- \[8\]Lieder, F\., Griffiths, T\.L\.: Resource\-rational analysis: Understanding human cognition as the optimal use of limited computational resources\. Behavioral and Brain Sciences43, e1 \(2020\)
- \[9\]Pezzulo, G\., Rigoli, F\., Friston, K\.J\.: Active inference, homeostatic regulation and adaptive behavioural control\. Progress in Neurobiology134, 17–35 \(2015\)
- \[10\]Seth, A\.K\., Friston, K\.J\.: Active interoceptive inference and the emotional brain\. Philosophical Transactions of the Royal Society B: Biological Sciences371\(1708\), 20160007 \(2016\)
- \[11\]Tschantz, A\., Barca, L\., Maisto, D\., Buckley, C\.L\., Seth, A\.K\., Pezzulo, G\.: Simulating homeostatic, allostatic and goal\-directed forms of interoceptive control using active inference\. Biological Psychology169, 108266 \(2022\)
- \[12\]Friston, K\., FitzGerald, T\., Rigoli, F\., Schwartenbeck, P\., Pezzulo, G\.: Active inference: a process theory\. Neural Computation29\(1\), 1–49 \(2017\)
- \[13\]Parr, T\., Pezzulo, G\., Friston, K\.J\.: Active Inference: The Free Energy Principle in Mind, Brain, and Behavior\. MIT Press \(2022\)
- \[14\]Keramati, M\., Gutkin, B\.: Homeostatic reinforcement learning for integrating reward collection and physiological stability\. eLife3, e04811 \(2014\)
- \[15\]Yoshida, N\., Arikawa, E\., Kanazawa, H\., Kuniyoshi, Y\.: Modeling long\-term nutritional behaviors using deep homeostatic reinforcement learning\. PNAS Nexus3\(12\), pgae540 \(2024\)
- \[16\]Mirza, M\.B\., Adams, R\.A\., Friston, K\.J\., Parr, T\.: Introducing a Bayesian model of selective attention based on active inference\. Scientific Reports9, 13915 \(2019\)
- \[17\]Whyte, C\.J\., Hohwy, J\., Smith, R\.: An active inference model of conscious access: How cognitive action selection reconciles the results of report and no\-report paradigms\. Current Research in Neurobiology3, 100036 \(2022\)
- \[18\]Solms, M\.: The hard problem of consciousness and the free energy principle\. Frontiers in Psychology9, 2714 \(2018\)
- \[19\]Cea, I\.: From insentient allostasis to adaptive bodily selfhood: Conscious vs unconscious instrumental interoceptive inference\. Adaptive Behavior \(2026\), online first
- \[20\]Hodson, R\., Grimbly, S\.J\., Boonstra, E\.A\., Bassett, B\.A\., van Hoof, C\., Rosman, B\., Solms, M\., Hakimi, N\., Shock, J\.P\., Smith, R\.: Sophisticated learning: A novel algorithm for active learning during model\-based planning \(2026\), arXiv preprint; original submission 2023, 10\-author revision 2026
- \[21\]Yoshida, N\., Daikoku, T\., Nagai, Y\., Kuniyoshi, Y\.: Emergence of integrated behaviors through direct optimization for homeostasis\. Neural Networks177, 106379 \(2024\)
- \[22\]Fermin, A\.S\.R\., Friston, K\., Yamawaki, S\.: An insula hierarchical network architecture for active interoceptive inference\. Royal Society Open Science9\(6\), 220226 \(2022\)
- \[23\]Livneh, Y\., et al\.: Homeostatic circuits selectively gate food cue responses in insular cortex\. Nature546, 611–616 \(2017\)

## Supplementary Material

These supplementary Sections A–E were not part of the 12\-page SAB 2026 camera\-ready paper\. In the combined arXiv document, that paper is reproduced without alteration and followed by these sections\. They report implementation details and analyses for the same agents, layouts and binary 60\-step survival outcome\. We first specify the planner and precision mechanism, then report layout and statistical sensitivity, selector and propagation controls, parameter robustness, food\-to\-poison non\-stationarity, and learning at matched observation count\. The agent, environment and layout banks are available at[https://github\.com/sgrimbly/attention\-aif\-sab2026\-snapshot](https://github.com/sgrimbly/attention-aif-sab2026-snapshot)\.

### A\. Implementation Details

##### Planner objective\.

The reported experiments used the pragmatic policy score, with state information gain disabled\. The resulting objective was the cross\-entropy

𝒢​\(π\)\\displaystyle\\mathcal\{G\}\(\\pi\)=−∑τ∑m𝔼q​\(oτ\(m\)∣π\)​\[log⁡P​\(oτ\(m\)∣C\)\]\\displaystyle\\;=\\;\-\\sum\_\{\\tau\}\\sum\_\{m\}\\mathbb\{E\}\_\{q\(o^\{\(m\)\}\_\{\\tau\}\\mid\\pi\)\}\\\!\\left\[\\log P\\\!\\left\(o^\{\(m\)\}\_\{\\tau\}\\mid C\\right\)\\right\]=KL\[q\(o∣π\)∥P\(o∣C\)\]⏟risk\+H​\[q​\(o∣π\)\]⏟κ​\-dependent entropy\.\\displaystyle\\;=\\;\\underbrace\{\\mathrm\{KL\}\\\!\\left\[q\(o\\mid\\pi\)\\,\\\|\\,P\(o\\mid C\)\\right\]\}\_\{\\text\{risk\}\}\\;\+\\;\\underbrace\{H\\\!\\left\[q\(o\\mid\\pi\)\\right\]\}\_\{\\kappa\\text\{\-dependent entropy\}\}\.In other words, interoceptive precision still reached the planner through the predicted\-observation distributionq​\(o∣π\)=A\(m\)​q​\(s∣π\)q\(o\\mid\\pi\)=A^\{\(m\)\}q\(s\\mid\\pi\), without a separate state\-information\-gain term\.

##### Where precision acts\.

For channelmm,κm\\kappa\_\{m\}sets both the probability that the body emits the true interoceptive level and the precision of the agent’s corresponding likelihoodA\(m\)A^\{\(m\)\}\. It therefore changes the sampled observation before inference, while the same likelihood shapes belief update and policy evaluation\. In this sense, selective precision is a limited sensing capacity directed towards one channel at a time\.

##### Policy precision\.

The reported runs used the planning library’s default policy precisionγ=1\\gamma=1inP​\(π\)∝e−γ​𝒢​\(π\)P\(\\pi\)\\propto e^\{\-\\gamma\\mathcal\{G\}\(\\pi\)\}\. All agents shared this value, so the reported contrasts are at matchedγ\\gamma\.

Table[2](https://arxiv.org/html/2608.04232#Pt0.Ax1.T2)collects the notation used below\. The four agent shorthands are uniform agent \(uniformκ\\kappa\), attentive agent \(dynamicκ\\kappatowards the most\-needed channel\), anti\-aligned control \(dynamicκ\\kappatowards the least\-needed channel\), and inference\-only ablation \(the shaped likelihood withheld from the planner\)\.

Table 2:Notation used in the paper and supplement\.

### B\. Layout and Statistical Sensitivity

We group the 12 layouts by the start\-to\-farther\-resource distance\. This is11–22cells in the easy tier,33–44cells in the medium tier, and66cells in the far tier\.L01is not geometrically unique\. It is separated because it is the only easy\-tier layout on which the uniform agent collapses under a learned world model while recovering under the true model\. This is a post\-hoc behavioural criterion, and excludingL01is conservative because its inclusion increases the attentive\-over\-uniform gap\.

![Refer to caption](https://arxiv.org/html/2608.04232v1/x6.png)Figure 5:L01horizon sensitivity\.\(a\)The layout and start position\.\(b\)With the true world model, increasing the planning horizon fromH=3H\{=\}3toH=4H\{=\}4rescues the uniform agent\.\(c\)Under learning, the same increase does not rescue it, while the attentive agent already survives atH=3H\{=\}3\. TheH=3H\{=\}3cells usen=8n\{=\}8batched seeds over 50 trials; theH=4H\{=\}4cells use one seed over 10 trials and are descriptive only\.On the full 12\-layout panel the attentive agent survives at0\.3780\.378against the uniform agent’s0\.1740\.174, a\+20\+20percentage\-point difference atp≤10−4p\\leq 10^\{\-4\}\. Since theH=4H\{=\}4cells use one seed, they do not estimate an effect size\. Instead, they suggest that theL01failure depends on the start position, shallow planning and a learned world model together\.

The headline confidence intervals and tests resample paired \(layout, seed\) clusters rather than individual trials, using10,00010\{,\}000bootstrap resamples\. Holm–Bonferroni correction is applied to the three primary survival contrasts in Table[3](https://arxiv.org/html/2608.04232#Pt0.Ax1.T3)\. The remaining sweeps in this supplement are exploratory\.

Table 3:Holm–Bonferroni\-corrected paired\-bootstrappp\-values for the three primary survival contrasts\. Values at the bootstrap resolution floor are reported as≤3×10−4\\leq 3\\times 10^\{\-4\}\.![Refer to caption](https://arxiv.org/html/2608.04232v1/x7.png)Figure 6:Easy\-tier survival by layout\.Per\-layout learning\-phase survival across the six easy\-tier layouts,n=32n\{=\}32batched seeds and 100 trials per cell\. The pooled bars excludeL01\. The attentive agent leads the uniform agent on five layouts\. OnL02, where both are near ceiling, their estimates differ by one percentage point\.Table 4:Easy\-tier sub\-panel excludingL01, with five layouts,n=32n\{=\}32batched seeds and16,00016\{,\}000trial outcomes per agent\.The ratio of survival rates depends on which layouts are pooled, so the absolute difference is easier to interpret\. The far tier contributes a common zero for every agent, while the uniform agent is also at zero throughout the medium tier\. The attentive\-over\-uniform difference is\+29\+29percentage points on the easy\-tier sub\-panel and\+21\.5\+21\.5percentage points across the 11\-layout headline panel\.

### C\. Selector and Propagation Controls

The attentive agent and anti\-aligned control use the same precision multiset\{0\.90,0\.567,0\.567,0\.567\}\\\{0\.90,0\.567,0\.567,0\.567\\\}and therefore the same total budget\. They differ only in which channel receives the larger value\. The anti\-aligned control performs worse than the uniform agent on the 11\-layout panel \(rawp=0\.004p\{=\}0\.004\), so uneven allocation does not explain the benefit on its own\. Panel\(c\)of Fig\.[8](https://arxiv.org/html/2608.04232#Pt0.Ax1.F8)provides the corresponding control over allocation magnitude\. The selectors coincide at the uniform allocationκatt=0\.65\\kappa\_\{\\mathrm\{att\}\}=0\.65and separate as the allocation becomes sharper\.

![Refer to caption](https://arxiv.org/html/2608.04232v1/x8.png)Figure 7:Where the precision signal acts\.Plateau survival \(trials 10–20\) across six Dirichlet prior concentrationsα0\\alpha\_\{0\}, over three easy\-tier layouts andn=8n\{=\}8batched seeds per cell\. The inference\-only ablation withholds the shaped likelihood from the planner\. Its gap from the full attentive agent widens as the body\-state prior becomes more rigid\.Atα0=10\\alpha\_\{0\}=10, plateau survival is0\.9050\.905for the full attentive agent and0\.0230\.023for the inference\-only ablation, close to the uniform agent at0\.0000\.000\. The anti\-aligned control remains intermediate at0\.4920\.492\. At the defaultα0=0\.1\\alpha\_\{0\}=0\.1, inference alone retains more of the benefit\. As the prior becomes rigid, the shaped likelihood must reach the planner for the advantage to remain\.

![Refer to caption](https://arxiv.org/html/2608.04232v1/x9.png)Figure 8:Robustness to prior, budget, and allocation\.Plateau survival across three parameter sweeps, with three layouts andn=8n\{=\}8batched seeds per cell\.\(a\)The attentive agent remains stable as Dirichlet prior concentration increases while the uniform agent collapses\.\(b\)The attentive agent exceeds the uniform agent at every tested precision budget; the high\-KKcells enter a saturation regime\.\(c\)Need\-aligned and anti\-aligned selectors coincide at uniform allocation and separate asκatt\\kappa\_\{\\mathrm\{att\}\}increases\. Panel\(a\)uses trials 10–20; panels\(b\)and\(c\)use trials 10–30\. Ribbons are cluster\-bootstrap 95% CIs over \(layout, seed\)\.The budget sweep coversK∈\{1\.5,2\.0,2\.6,3\.0,3\.5,4\.0\}K\\in\\\{1\.5,2\.0,2\.6,3\.0,3\.5,4\.0\\\}and does not plot the exact allocation\-degeneracy pointK=M​κatt=3\.6K=M\\kappa\_\{\\mathrm\{att\}\}=3\.6\. At that value, the need\-aligned and anti\-aligned variants both assignκ=0\.90\\kappa=0\.90to every channel, so selector direction no longer changes the allocation\. The uniform curve uses the paper’s original baseline implementation, so panel\(b\)is not a matched\-architecture convergence test\. AtK=4K=4, unattended channels clip at1\.01\.0, making the nominally attended channel the least precise\. These cells therefore describe saturation rather than a reversal of the mechanism\.

### D\. Food\-to\-Poison Non\-stationarity

A fixed resource map is the easiest case for a learned spatial model\. We relax this assumption by allowing food to turn into poison in place during a trial\.

![Refer to caption](https://arxiv.org/html/2608.04232v1/x10.png)Figure 9:Food\-to\-poison non\-stationarity\.At each step, a food tile in the agent’s destination cell turns to poison in place with probabilityρ\\rho\. Water tiles never change and no tile moves\. The attentive agent leads the uniform agent at all three mutation rates, while survival and the absolute gap decline asρ\\rhoincreases\. Bars show marginal cluster\-bootstrap 95% CIs over \(layout, seed\); Table[5](https://arxiv.org/html/2608.04232#Pt0.Ax1.T5)reports paired tests\.Table 5:Food\-to\-poison sweep over three layouts,n=8n\{=\}8batched seeds and 50 trials per cell\. CIs are cluster\-bootstrap intervals over \(layout, seed\), andpp\-values are paired\-bootstrap values with Holm–Bonferroni correction across the three mutation rates\.The conversion occurs before the tile takes effect, so the agent consumes poison on the step that triggers the change\. The tile then stays poisoned for the rest of the trial, and layouts reset between trials\. Sinceρ\\rhois conditional on entering a food cell, agents that forage more often are exposed more often\. The attentive advantage remains positive across the sweep, but decreases as the resource map becomes less reliable\.

### E\. Learning at Matched Observation Count

For the channel\-wise learning analysis, we use three easy\-tier layouts \(L00,L02andL04\),n=8n\{=\}8batched seeds andα0=0\.1\\alpha\_\{0\}=0\.1\.

![[Uncaptioned image]](https://arxiv.org/html/2608.04232v1/x11.png)

Figure 10:Channel\-wise likelihood learning\.Per\-channel Dirichlet learning over 50 trials\.\(a\)The attentive agent’s hunger model is more accurate from the first trial \(0\.480\.48against0\.250\.25\), and both agents approach their plateaus within about five trials\. This is a level difference rather than a distinct convergence rate\.\(b\)Thirst learning is similar under both agents\.\(c\)The inert control stays near uniform\.\(d\)Suffocation shows a smaller difference than hunger\.\(e\)Within the attentive agent, attendance fraction correlates with learning quality for hunger \(r=0\.47r\{=\}0\.47\) and thirst \(r=0\.38r\{=\}0\.38\), but not suffocation \(r=0\.03r\{=\}0\.03\)\.
Over plateau trials, the attentive agent allocates the largest mean share to thirst \(0\.490\.49\), followed by hunger \(0\.330\.33\), suffocation \(0\.170\.17\) and the inert control \(0\.020\.02\)\. Hunger nevertheless separates the agents most, since it is the channel the uniform agent learns least accurately\. The within\-channel correlations in Fig\.[10](https://arxiv.org/html/2608.04232#Pt0.Ax1.F10)e are positive for hunger and thirst and near zero for suffocation, which stays close to its setpoint over most of an episode\.

![[Uncaptioned image]](https://arxiv.org/html/2608.04232v1/x12.png)

Figure 11:Learning at matched observation count\.Hunger\-channel likelihood accuracy against cumulative observation count rather than trial index\.\(a\)Accuracy with cluster\-bootstrap 95% CIs\.\(b\)Meanκ\\kappareceived by the hunger channel\. Where all 24 \(layout, seed\) clusters contribute \(x≤750x\\leq 750\), the attentive agent, uniform agent and anti\-aligned control receive meanκ\\kappaof0\.7490\.749,0\.6500\.650and0\.5860\.586, and reach accuracy of0\.6290\.629,0\.2990\.299and0\.5240\.524, respectively\. Beyondx=750x=750, only longer\-surviving clusters remain and the region is shaded\.
The attentive\-over\-uniform advantage persists at matched observation count, so it is not a result of longer\-surviving agents collecting more data\. The anti\-aligned control remains second in accuracy despite receiving the lowest mean hunger precision\. Direction therefore separates the two dynamic agents more clearly in survival than in this channel\-specific learning measure\.

Similar Articles

Risk-Aware Decision Policies for Agents Under Noisy Perception

arXiv cs.LG

This paper presents an artificial life predator-prey model of foraging under noisy perception, showing that uncertainty-aware decision policies significantly improve survival compared to blindly trusting perceptual labels, and that agents transition from exploratory to conservative strategies as uncertainty increases.

TaskSense: Focusing on What Matters in World Models

arXiv cs.AI

TaskSense introduces a task-centric world modeling framework that uses stochastic spatial attention conditioned on previous latent states and an auxiliary inverse-dynamics objective to focus on control-relevant regions, improving robustness to visual distractions compared to DreamerV3.

Proactive Memory for Long-Horizon Agents (16 minute read)

TLDR AI

This paper introduces a proactive memory agent that operates alongside a standard action agent to selectively inject memory-grounded reminders during long-horizon tasks, mitigating behavioral state decay. Experiments on Terminal-Bench and τ²-Bench show significant improvements in pass@1, and the approach is demonstrated with both weak and strong action agents.