A Transdiagnostic Space of Disorder Like Phenotypes in Reinforcement Learning Agents
Summary
This paper introduces a dose-controllable method for inducing seven psychological disorders in reinforcement learning agents by manipulating cognitive appraisal signals in an appraisal-guided PPO agent. The disorders self-organize into a two-dimensional affective space, and the framework enables modeling of both disorder induction and treatment.
View Cached Full Text
Cached at: 07/10/26, 06:14 AM
# A Transdiagnostic Space of Disorder Like Phenotypes in Reinforcement Learning Agents
Source: [https://arxiv.org/html/2607.07753](https://arxiv.org/html/2607.07753)
###### Abstract
Modelling psychological disorders in artificial agents offers both a testbed for computational psychiatry and a lens on the failure modes of affective control\. Prior work induces one or two disorders in a reinforcement\-learning \(RL\) agent by hand\-tuned reward shaping, labels the behaviour post hoc, and reports single runs\. We recast disorder modelling as*dose\-controllable*manipulation of cognitive appraisal signals in an appraisal\-guided PPO agent, expressing seven disorders \(anxiety, mania, obsessive–compulsive checking, depression, impulsivity, addiction, and post\-traumatic stress\) each as a single knob grounded in a computational psychiatry account, with each symptom measured by a preregistered assay mapped to a recognised paradigm\. Across more than a thousand runs \(10 seeds, four controls, 95% confidence intervals\) every disorder shows a graded, monotone dose–response that no control reproduces\. Beyond these induced effects, three findings emerge that were not written into the reward: the disorders self\-organise into a two\-dimensional affective space in which mania mirrors anxiety; removing a knob remits reward distortion disorders \(mania, checking, addiction\) but not avoidance disorders \(anxiety, PTSD\), which instead recover under a graded exposure curriculum; and two simultaneous knobs interact nonadditively, yielding testable comorbidity predictions\. Appraisal weights thus parameterise a controllable space of affective phenotypes in which the same knobs that induce a disorder can model its treatment\. We also show that three disorder knobs \(depression, addiction, anxiety\) transfer to a three\-dimensional pixel environment \(MiniWorld\) with a standard convolutional agent and no appraisal critic, with cross\-assay dissociation confirmed across both domains, indicating the framework is not specific to grid worlds or to PPO’s appraisal critic\.
## Introduction
Reinforcement\-learning \(RL\) agents increasingly act in healthcare, transport, and human\-facing assistants, where their affective stability is a safety property rather than a curiosity: value estimation that over\-weights harm or reward produces avoidance, perseveration, freezing, or reckless behaviour that undermine reliability and user trust\. Psychological disorders are, in this framing, the characteristic failure modes of affective control, and an agent in which such failures can be induced, measured, and reversed under experimental control is a useful object of study\.
The same object is valuable to computational psychiatry, whose central aim is to relate transdiagnostic dimensions of symptomatology to the latent learning and decision computations that generate them\[[Montague et al\. 2012](https://arxiv.org/html/2607.07753#bib.bibx13),[Huys et al\. 2016](https://arxiv.org/html/2607.07753#bib.bibx5),[Maia and Frank 2011](https://arxiv.org/html/2607.07753#bib.bibx10),[Insel et al\. 2010](https://arxiv.org/html/2607.07753#bib.bibx29)\]\. Human and rodent studies fit abstract bandit or Markov\-decision tasks to behaviour, but cannot manipulate a disorder’s severity continuously or observe the full trajectory from health to pathology and back\. An artificial agent can\.
Appraisal theory\[[Lazarus 1991](https://arxiv.org/html/2607.07753#bib.bibx9),[Scherer 2001](https://arxiv.org/html/2607.07753#bib.bibx16)\]holds that emotion arises from domain\-independent evaluations \(relevance, certainty, novelty, congruence, coping potential, and anticipation\) of events with respect to an agent’s goals\. These dimensions overlap with intrinsic motivation signals used in RL\[[Sequeira et al\. 2011](https://arxiv.org/html/2607.07753#bib.bibx18)\], which makes appraisal a natural interface between affect and value learning\. Prior work uses appraisals to*elicit*emotions; we instead use them to*induce disorders*in a graded, controlled fashion\.
Existing RL disorder models share three weaknesses\. Manipulations are hand\-tuned rather than derived from a mechanistic account, so the mapping to a clinical construct is loose\. Behaviour is labelled after the fact, “this looks like OCD”, rather than measured against a criterion specified in advance, inviting confirmation bias\. And results are typically single runs without variance or controls\.
We address all three, but our central claim is stronger and is worth stating plainly to pre\-empt an obvious objection\. Shaping reward to induce a behaviour can make that behaviour trivial: penalise threat and of course the agent avoids\. We therefore distinguish*induced*effects, the dose–response curves, which merely*validate*that each knob is well\-behaved and graded, from*emergent*effects, which were not written into the reward and are the paper’s primary contribution\. Three results are emergent\. The seven disorders self\-organise into a two\-dimensional affective space in which mania is the exact mirror of anxiety, though the two are trained by opposite signs of a single knob and never compared during training\. Removing a disorder’s knob reveals a remit\-versus\-resist dissociation that no reward term encodes\. And placing two knobs together yields nonadditive interactions\. Our contributions are:
1. 1\.Emergent transdiagnostic structure: the seven disorders occupy predictable, mutually consistent regions of a reward–approach×\\timesthreat–avoidance space, with mania mirroring anxiety \(Fig\.[4](https://arxiv.org/html/2607.07753#Sx5.F4)\)\.
2. 2\.An emergent treatment dissociation, and its resolution: removing the pathological knob remits reward distortion disorders but not habit/avoidance disorders; a graded exposure\-with\-response\-prevention curriculum then recovers the resistant ones, mirroring why avoidance clinically requires active exposure \(Fig\.[5](https://arxiv.org/html/2607.07753#Sx5.F5), Table[4](https://arxiv.org/html/2607.07753#Sx5.T4)\)\.
3. 3\.Emergent comorbidity: two simultaneous knobs interact nonadditively, a prediction the framework makes and we test \(Fig\.[6](https://arxiv.org/html/2607.07753#Sx5.F6)\)\.
4. 4\.A grounded, controllable substrateenabling the above: AG\-PPO \(Fig\.[1](https://arxiv.org/html/2607.07753#Sx1.F1)\) with seven knobs each tied to a named computational psychiatry account \(Table[1](https://arxiv.org/html/2607.07753#Sx4.T1)\), preregistered assays, 10\-seed dose–response with confidence intervals, and four controls that reproduce no phenotype \(Tables[2](https://arxiv.org/html/2607.07753#Sx5.T2),[3](https://arxiv.org/html/2607.07753#Sx5.T3)\)\.
DynamicenvironmentConvencoderActorCriticAppraisalestimationNRE netactionvalueζ\\zetareward shapingρ\(ζ\)\\rho\(\\zeta\)Figure 1:AG\-PPO\. A shared convolutional encoder feeds an actor and a critic; the critic additionally receives the six\-dimensional appraisal vectorζ\\zeta\. A next\-reward \(NRE\) network supplies the anticipation appraisal\. Appraisals also shape the environment reward\. The disorder knob is a single term in the shaping function or the discount factor\.
## Related Work
##### Emotion in RL and affective computing\.
Computational models of emotion combine appraisal theory with RL to elicit affective responses, operationalising appraisal dimensions through temporal\-difference signals and intrinsic motivation features\[[Sequeira et al\. 2011](https://arxiv.org/html/2607.07753#bib.bibx18),[Moerland et al\. 2018](https://arxiv.org/html/2607.07753#bib.bibx12)\]\. Symbolic architectures such as the OCC model provide structured emotion representations but adapt poorly to unstructured tasks\. These lines target emotion*elicitation*; we instead use appraisals as controllable pathological knobs and evaluate the resulting behaviour against clinical paradigms\.
##### Computational psychiatry\.
A central aim is to relate transdiagnostic symptom dimensions to latent learning and decision computations\[[Montague et al\. 2012](https://arxiv.org/html/2607.07753#bib.bibx13),[Huys et al\. 2016](https://arxiv.org/html/2607.07753#bib.bibx5),[Maia and Frank 2011](https://arxiv.org/html/2607.07753#bib.bibx10)\]\. Canonical mechanistic accounts include reward prediction\-error models of addiction in which a drug signal fails to be compensated\[[Redish 2004](https://arxiv.org/html/2607.07753#bib.bibx15)\], incentive salience accounts in which wanting dissociates from liking\[[Robinson and Berridge 1993](https://arxiv.org/html/2607.07753#bib.bibx33)\], effort\-based decision deficits in depression\[[Treadway and Zald 2011](https://arxiv.org/html/2607.07753#bib.bibx19),[Beck 1979](https://arxiv.org/html/2607.07753#bib.bibx30)\], steep delay discounting in impulsivity\[[Ainslie 1975](https://arxiv.org/html/2607.07753#bib.bibx2)\], checking as diminishing reassurance in OCD\[[Rachman 2002](https://arxiv.org/html/2607.07753#bib.bibx14)\], orbitofrontal dysfunction as a neural substrate\[[Chamberlain et al\. 2008](https://arxiv.org/html/2607.07753#bib.bibx32)\], and impaired fear extinction in PTSD\[[Pitman et al\. 2012](https://arxiv.org/html/2607.07753#bib.bibx31)\]\. These are usually fit to behaviour in abstract tasks\. We realise them as behavioural phenotypes of a single spatial agent whose severity is tuned continuously, enabling dose–response and treatment analyses that fitting cannot provide\.
##### Relation to prior appraisal\-guided disorder modelling\.
Our prior work\[[Prasad, Jacob, and Ahamed 2024](https://arxiv.org/html/2607.07753#bib.bibx3)\]introduced the appraisal\-guided PPO architecture and the six appraisal equations used here, and induced two disorders \(anxiety and OCD\) via reward shaping, but labelled the behaviour post hoc and reported single runs on a single grid world\. We build on that formulation and generalise it substantially: seven grounded, dose\-controlled disorders instead of two; preregistered assays rather than post\-hoc labels; ten seeds with confidence intervals and four control conditions rather than single runs; a unifying two\-dimensional affective space; and a rescue/extinction experiment that models treatment\.
## Background
##### Proximal Policy Optimization\.
PPO\[[Schulman et al\. 2017](https://arxiv.org/html/2607.07753#bib.bibx17)\]optimises the clipped surrogate
LCLIP\(θ\)=𝔼^t\[min\(rtA^t,clip\(rt,1−ϵ,1\+ϵ\)A^t\)\],L^\{\\text\{CLIP\}\}\(\\theta\)=\\hat\{\\mathbb\{E\}\}\_\{t\}\\\!\\left\[\\min\\\!\\big\(r\_\{t\}\\hat\{A\}\_\{t\},\\,\\text\{clip\}\(r\_\{t\},1\{\-\}\\epsilon,1\{\+\}\\epsilon\)\\,\\hat\{A\}\_\{t\}\\big\)\\right\],\(1\)wherertr\_\{t\}is the importance ratio andA^t\\hat\{A\}\_\{t\}a generalised\-advantage estimate\. Its on\-policy, real\-time nature suits the nonstationary grid worlds we study, in which obstacles and goals move\.
##### Cognitive appraisals\.
At each step the agent computes six appraisalsζit∈\(0,1\)\\zeta^\{t\}\_\{i\}\\in\(0,1\): motivational relevance and goal congruence \(functions of agent–goal distance\), certainty and novelty \(entropy and uniform\-KL of the policy\), coping potential \(fraction of threats outside the agent’s view\), and anticipation \(complement of a next\-reward prediction error\)\. The formulas are given in the supplement\. These variables are re\-scaled to\(0,1\)\(0,1\)and enter both value estimation and reward shaping\.
## Method
##### AG\-PPO architecture\.
A convolutional encoder \(three layers, then a 256\-unit head\) maps the7×7×37\{\\times\}7\{\\times\}3egocentric symbolic observation to features shared by the actor and critic \(Fig\.[1](https://arxiv.org/html/2607.07753#Sx1.F1)\)\. The critic additionally consumes the six\-dimensional appraisal vector, so value estimation is appraisal\-informed; the actor produces action probabilities over three actions \(turn left, turn right, move forward\)\. A next\-reward network predictsrtr\_\{t\}from\(ot−1,at−1\)\(o\_\{t\-1\},a\_\{t\-1\}\)and supplies the anticipation appraisal\. Reward is shaped as
rt′=rt−∑iwigi\(ζt\)−ceff\[at=fwd\]\+bchk\+bdrug−bshk,r^\{\\prime\}\_\{t\}=r\_\{t\}\-\\textstyle\\sum\_\{i\}w\_\{i\}\\,g\_\{i\}\(\\zeta^\{t\}\)\-c\_\{\\text\{eff\}\}\[a\_\{t\}\{=\}\\text\{fwd\}\]\+b\_\{\\text\{chk\}\}\+b\_\{\\text\{drug\}\}\-b\_\{\\text\{shk\}\},\(2\)wherertr\_\{t\}is the environment reward at steptt,wiw\_\{i\}are appraisal\-shaping weights,ceffc\_\{\\text\{eff\}\}an effort cost,bchkb\_\{\\text\{chk\}\}a checking bonus,bdrugb\_\{\\text\{drug\}\}a drug bonus, andbshkb\_\{\\text\{shk\}\}a trauma shock\. Exactly one term is active per disorder; the discount factorγ\\gammais the impulsivity knob\. This isolates each disorder to a single, interpretable degree of freedom\.
##### The six appraisals and how they are used\.
Each step the agent computes six appraisalsζit∈\(0,1\)\\zeta^\{t\}\_\{i\}\\in\(0,1\)\. Let\(xa,ya\)\(x\_\{a\},y\_\{a\}\)be the agent position,\(xg,yg\)\(x\_\{g\},y\_\{g\}\)the goal,wwthe grid size,nnthe view size, andp=softmax\(logits\)p=\\mathrm\{softmax\}\(\\text\{logits\}\)the policy over actions\. Motivational relevance and goal congruence are geometric \(agent–goal distance\); certainty and novelty are read from the actor’s action distribution; coping potential is read from the fraction of threats in view; and anticipation is the complement of the next\-reward error from the NRE network:
ζMRt\\displaystyle\\zeta^\{t\}\_\{\\text\{MR\}\}=1−\(\|xa−xg\|\+\|ya−yg\|\)−12\(w−1\)\\displaystyle=1\-\\tfrac\{\(\|x\_\{a\}\{\-\}x\_\{g\}\|\+\|y\_\{a\}\{\-\}y\_\{g\}\|\)\-1\}\{2\(w\-1\)\}\(3\)ζGCt\\displaystyle\\zeta^\{t\}\_\{\\text\{GC\}\}=1−\(xa−xg\)2\+\(ya−yg\)2\(\(n−1\)/2\)2\+n2\\displaystyle=1\-\\tfrac\{\\sqrt\{\(x\_\{a\}\{\-\}x\_\{g\}\)^\{2\}\+\(y\_\{a\}\{\-\}y\_\{g\}\)^\{2\}\}\}\{\\sqrt\{\(\(n\{\-\}1\)/2\)^\{2\}\+n^\{2\}\}\}\(4\)ζCt\\displaystyle\\zeta^\{t\}\_\{\\text\{C\}\}=1−−∑plogp1−∑plogp\\displaystyle=1\-\\tfrac\{\-\\sum p\\log p\}\{1\-\\sum p\\log p\}\(5\)ζNt\\displaystyle\\zeta^\{t\}\_\{\\text\{N\}\}=KL\(U∥p\)1\+KL\(U∥p\)\\displaystyle=\\tfrac\{\\mathrm\{KL\}\(U\\\|p\)\}\{1\+\\mathrm\{KL\}\(U\\\|p\)\}\(6\)ζCPt\\displaystyle\\zeta^\{t\}\_\{\\text\{CP\}\}=1−kobsnobs\+ε\\displaystyle=1\-\\tfrac\{k\_\{\\text\{obs\}\}\}\{n\_\{\\text\{obs\}\}\+\\varepsilon\}\(7\)ζAt\\displaystyle\\zeta^\{t\}\_\{\\text\{A\}\}=1−\|rt−NRE\(ot−1,at−1\)\|\\displaystyle=1\-\\big\|r\_\{t\}\-\\mathrm\{NRE\}\(o\_\{t\-1\},a\_\{t\-1\}\)\\big\|\(8\)whereKL\\mathrm\{KL\}is the Kullback–Leibler divergence\[[Kullback and Leibler 1951](https://arxiv.org/html/2607.07753#bib.bibx8)\],UUis the uniform policy,kobsk\_\{\\text\{obs\}\}/nobsn\_\{\\text\{obs\}\}are the number of threats in view / in total, andε\\varepsilonavoids division by zero\. The appraisal vectorζt=\(ζMR,ζC,ζN,ζGC,ζCP,ζA\)\\zeta^\{t\}=\(\\zeta\_\{\\text\{MR\}\},\\zeta\_\{\\text\{C\}\},\\zeta\_\{\\text\{N\}\},\\zeta\_\{\\text\{GC\}\},\\zeta\_\{\\text\{CP\}\},\\zeta\_\{\\text\{A\}\}\)plays two roles\. First, it is concatenated to the encoder features and consumed by the critic, so value estimation is appraisal\-informed \(the AG\-PPO critic of Fig\.[1](https://arxiv.org/html/2607.07753#Sx1.F1)\)\. Second, it is the basis of the reward shaping in Eq\.[2](https://arxiv.org/html/2607.07753#Sx4.E2): the anxiety and mania knobs act on coping potentialζCP\\zeta\_\{\\text\{CP\}\}\(penalising low or high values respectively\), while the stress index used as a secondary assay is the weighted deviation∑i\(1−ζi\)wi\\sum\_\{i\}\(1\-\\zeta\_\{i\}\)w\_\{i\}\. The remaining knobs \(effort, checking, drug, shock, discount\) act on the environment reward directly rather than throughζ\\zeta, but the critic still observesζ\\zetathroughout\.
##### Disorder mechanisms\.
Table[1](https://arxiv.org/html/2607.07753#Sx4.T1)lists each knob and its grounding\.*Anxiety*penalises low coping potential \(a threat in view\), inducing avoidance;*mania*is its exact mirror, penalising high coping potential and thus seeking threat\.*Checking*grants a bonus for returning to a checkpoint whosekk\-th visit within an episode paysλk\\lambda^\{k\}of the base \(λ=0\.5\\lambda\{=\}0\.5\), modelling diminishing reassurance\[[Rachman 2002](https://arxiv.org/html/2607.07753#bib.bibx14)\]; this bounds the episodic total and converts an otherwise unbounded, all\-or\-nothing incentive into a graded number of checks\.*Depression*imposes a per\-step effort cost on movement\[[Treadway and Zald 2011](https://arxiv.org/html/2607.07753#bib.bibx19)\], reducing behavioural activation\.*Impulsivity*lowersγ\\gamma\(steeper delay discounting\[[Ainslie 1975](https://arxiv.org/html/2607.07753#bib.bibx2)\]\)\.*Addiction*adds a non\-habituating drug\-tile bonus whose accumulated value can exceed the one\-shot goal reward\[[Redish 2004](https://arxiv.org/html/2607.07753#bib.bibx15)\]\.*PTSD*applies a shock on a trauma tile that lies on the short route to the goal, so avoidance over\-generalises to a long detour\.
Table 1:Disorder mechanisms as dose\-controllable knobs and their grounding\. Each disorder activates a single term of Eq\.[2](https://arxiv.org/html/2607.07753#Sx4.E2)or the discount\.Algorithm 1AG\-PPO with a disorder knobκ\\kappa1:Input:policy
πθ\\pi\_\{\\theta\}, appraisal critic
VϕV\_\{\\phi\}, NRE
fψf\_\{\\psi\}, knob
κ\\kappa
2:foriteration
=1,2,…=1,2,\\dotsdo
3:for
t=1t=1to
TT\(rollout\)do
4:observe
sts\_\{t\}; sample
at∼πθ\(⋅∣st\)a\_\{t\}\\sim\\pi\_\{\\theta\}\(\\cdot\\mid s\_\{t\}\)
5:compute appraisals
ζt\\zeta^\{t\}\{Eqs\.[3](https://arxiv.org/html/2607.07753#Sx4.E3)–[8](https://arxiv.org/html/2607.07753#Sx4.E8)\}
6:execute
ata\_\{t\}; observe
rt,st\+1r\_\{t\},\\,s\_\{t\+1\}
7:
rt′←Shape\(rt,ζt,κ\)r^\{\\prime\}\_\{t\}\\leftarrow\\textsc\{Shape\}\(r\_\{t\},\\zeta^\{t\},\\kappa\)\{Eq\.[2](https://arxiv.org/html/2607.07753#Sx4.E2)\}
8:store
\(st,at,rt′,ζt\)\(s\_\{t\},a\_\{t\},r^\{\\prime\}\_\{t\},\\zeta^\{t\}\)
9:endfor
10:compute advantages via GAE using
Vϕ\(st,ζt\)V\_\{\\phi\}\(s\_\{t\},\\zeta^\{t\}\)
11:forepoch
=1=1to
KK; each minibatchdo
12:update
θ,ϕ\\theta,\\phion clipped surrogate
\+\+value loss
13:update
ψ\\psion
\(rt−fψ\(ot−1,at−1\)\)2\(r\_\{t\}\-f\_\{\\psi\}\(o\_\{t\-1\},a\_\{t\-1\}\)\)^\{2\}
14:endfor
15:endfor
##### Environments\.
We use four threat/spatial grid worlds and three purpose\-built environments \(rendered in appendix Fig\.[8](https://arxiv.org/html/2607.07753#A0.F8)\)\. Dynamic\-Obstacles, LavaGap, and LavaCrossing are standard MiniGrid tasks with moving obstacles or lava hazards\[[Chevalier\-Boisvert et al\. 2023](https://arxiv.org/html/2607.07753#bib.bibx20)\]\. Approach–Avoidance is a custom two\-goal conflict: a high\-value goal sits behind a lava\-lined corridor in which every approach step is threat\-adjacent, and a low\-value goal \(0\.15×0\.15\\times\) has an open approach at equal distance, so route choice reflects threat sensitivity\. Temporal\-Choice is a corridor with a near reward0\.50\.5and a far reward1\.01\.0, resolved by the agent’s discount \(crossoverγ≈0\.89\\gamma\{\\approx\}0\.89\)\. Addiction places a repeatable drug tile opposite the goal\. Trauma places a conditioned\-shock tile on the short route with a long safe detour available\. All use7×77\{\\times\}7egocentric symbolic observations\.
##### Pre\-registered assays\.
Before running experiments we fixed one primary symptom assay per disorder, each mapping to a recognised paradigm: risky\-goal choice and mean threat distance \(approach–avoidance conflict\) for anxiety and mania; checking rate for OCD; forward\-action fraction \(behavioural activation\) for depression; near\-versus\-far reward choice \(delay discounting\) for impulsivity; drug\-tile occupancy for addiction; and trauma distance for PTSD\. Secondary assays, thigmotaxis, freezing, stereotypy, turnarounds, visitation entropy, and a stress index, are reported in the supplement\.
Figure 2:Phenotype gallery: state\-occupancy over 40 episodes at the severe dose of each disorder, overlaid on the environment \(brighter is more visited\)\. The spatial signature of each disorder is directly legible: the anxiety avoidance band, the mania approach to lava, the OCD checking loop, the depressive stationarity, the impulsive near\-goal fixation, the addictive drug corner, and the PTSD detour\.
## Experiments
##### Protocol\.
Each configuration is run with 10 seeds; we report means with 95% confidence intervals\. Four controls anchor every comparison: standard PPO, a critic\-noise control \(a random vector replaces the appraisals\), PPO\+RND intrinsic motivation, and the appraisal critic without shaping \(theϵ=0\\epsilon\{=\}0point of every dose–response\)\. Because avoidance can trivially reduce success, symptom analysis is restricted, by a criterion fixed before analysis, to runs that solve the task; all counts are reported\. In total we ran 1,375 configurations on CPU\.
##### Dose–response\.
Table[2](https://arxiv.org/html/2607.07753#Sx5.T2)and Fig\.[3](https://arxiv.org/html/2607.07753#Sx5.F3)report the primary assay for each disorder across knob doses\. Every disorder is graded and monotone\. Anxiety abandons the high\-value risky goal as the coping penalty grows \(0\.73→0\.000\.73\\\!\\to\\\!0\.00\) while task success is preserved, the signature of avoidance rather than incompetence; mania hugs lava and its death rate rises to0\.700\.70; checking rises smoothly to0\.290\.29with success maintained until the most severe dose; behavioural activation collapses under effort cost; impulsive agents shift to the near reward asγ\\gammafalls; drug occupancy rises past a vulnerability threshold nearϵ=0\.05\\epsilon\{=\}0\.05, beyond which the agent forgoes the goal entirely; and trauma distance grows monotonically as the agent forgoes the short route\.
Figure 3:Dose–response of the primary symptom assay per disorder \(mean, shaded 95% CI, 10 seeds\); dashed lines are the PPO baseline\. Severity increases smoothly and monotonically with the knob dose in every case\.Table 2:Dose–response of the primary symptom assay per disorder \(mean±\\pm95% CI, 10 seeds; 30 for anxiety\)\.ϵ1<…<ϵ4\\epsilon\_\{1\}\\\!<\\\!\\dots\\\!<\\\!\\epsilon\_\{4\}are the four knob doses \(disorder\-specific; see supplement\); PPO is the untreated baseline\. Depression is reported on Approach–Avoidance, where the baseline solves the task cleanly\.
##### Controls do not produce phenotypes\.
Table[3](https://arxiv.org/html/2607.07753#Sx5.T3)contrasts the four controls on two environments\. None induces a disorder: on Approach–Avoidance every control retains a high risky\-goal choice and short threat distance, and on LavaGap checking and death remain near zero\. The appraisal critic matches PPO, confirming that the phenotypes require the specific shaping knob rather than appraisal input per se or a generic intrinsic bonus\. Notably, RND over\-explores into lava on LavaGap \(6% success\), a reminder that generic novelty seeking is not a disorder model\.
Table 3:Controls do not reproduce phenotypes \(mean, 10 seeds\)\. Left: Approach–Avoidance\. Right: LavaGap\.
##### A negative result and a redesign\.
Our first OCD mechanism penalised low certainty; it*reduced*checking rather than inducing it, because rewarding decisiveness suppresses re\-verification\. A checkpoint bonus without habituation was then bistable: ignored below a threshold, and an unbounded loop that captured the agent above it\. Introducing diminishing reassurance \(λk\\lambda^\{k\}decay\) bounded the bonus and produced the graded checking of Table[2](https://arxiv.org/html/2607.07753#Sx5.T2)\. We report both the failure and the fix, as they clarify what does and does not induce compulsion and illustrate the value of a mechanistic, rather than cosmetic, manipulation\.
##### A transdiagnostic affective space\.
Figure[4](https://arxiv.org/html/2607.07753#Sx5.F4)places each disorder in a plane spanned by reward–approach \(anhedonic withdrawal to disinhibited over\-pursuit\) and threat–avoidance \(threat\-seeking to hypervigilant avoidance\), using normalised behavioural deviations from each environment’s healthy baseline\. As knob dose increases, agents travel from the healthy origin into disorder\-specific regions: anxiety and PTSD into avoidance, depression into withdrawal, addiction and impulsivity into over\-pursuit, and mania into disinhibition, the reflection of anxiety across the origin\. OCD lies near the origin because compulsivity is a distinct out\-of\-plane dimension, which we annotate rather than force into two dimensions\. The spatial signatures underlying these coordinates are visible directly in the state\-occupancy maps of Fig\.[2](https://arxiv.org/html/2607.07753#Sx4.F2): the anxiety avoidance band, the mania approach to lava, the OCD checking loop, the depressive stationarity, the impulsive near\-goal fixation, the addictive drug corner, and the PTSD detour\.
Figure 4:The transdiagnostic affective space\. Each disorder is a dose\-trajectory from the healthy origin; marker size encodes dose\. Mania is the mirror of anxiety across the origin; OCD occupies a separate compulsivity dimension\.
##### Rescue as treatment\.
We warm\-start from the most severe trained model of each disorder and continue training for a fixed budget either with the pathological knob removed \(treatment\) or kept \(control\), logging the primary assay throughout\. Table[4](https://arxiv.org/html/2607.07753#Sx5.T4)and Fig\.[5](https://arxiv.org/html/2607.07753#Sx5.F5)show a dissociation\. Disorders driven by an ongoing reward distortion \(mania, checking, addiction\) remit immediately once the knob is removed, because the goal gradient re\-dominates\. Disorders that have reshaped the policy into a self\-reinforcing behavioural habit \(avoidance in anxiety and PTSD, passivity in depression, near\-reward fixation in impulsivity\) resist passive removal: the safe or myopic policy keeps working, so the agent never re\-encounters the disconfirming evidence that would extinguish it\. This mirrors the clinical observation that avoidance disorders require active exposure, not mere removal of the stressor\.
##### Exposure therapy for the resistant disorders\.
We therefore test the corresponding intervention on the two resistant avoidance disorders\. Warm\-starting from the severe anxiety and PTSD models, we apply a graded exposure curriculum with response prevention: a penalty on the avoidance route \(the safe alternative\) annealed to zero over training, forcing the agent to re\-confront the feared route and discover it is survivable\. Crucially, a simple reward for*reaching*the feared route fails, an avoidant policy never goes there to collect it, so exposure must act on the avoidance response itself, exactly the logic of exposure\-and\-response\-prevention therapy\[[Craske et al\. 2014](https://arxiv.org/html/2607.07753#bib.bibx4)\]\. Evaluated afterwards on the unmodified environment, both disorders recover: feared\-route choice rises from0\.200\.20under passive removal to0\.900\.90–0\.930\.93\(Table[4](https://arxiv.org/html/2607.07753#Sx5.T4)\), and the recovery persists after the exposure prompt has faded, genuine relearning rather than a maintained incentive\. The curriculum succeeds precisely where passive knob\-removal failed, closing the loop from induction to treatment\.
Figure 5:Rescue trajectories: primary symptom over continued training with the knob removed \(treatment\) vs\. kept \(control\), 5 seeds\. Reward\-distortion disorders remit; habit/avoidance disorders resist\.Table 4:Rescue/extinction \(5 seeds\)\.*Passive*: pathological knob removed\.*Control*: knob kept\.*Expo\.*: graded exposure curriculum \(avoidance disorders only; 15 seeds anxiety, 10 PTSD\)\. Reward\-distortion disorders remit under passive removal; avoidance disorders resist it but recover under exposure\.
##### Emergent comorbidity\.
Because each disorder is a separate term, two can be activated at once\. We sweep 2\-D dose grids for two pairs and ask whether the joint effect is the sum of the single\-knob effects; departure from additivity is an emergent interaction the reward never encodes\. Both pairs are strongly nonadditive \(Fig\.[6](https://arxiv.org/html/2607.07753#Sx5.F6)\)\. For*mania×\{\\times\}impulsivity*the interaction is striking: mania alone drives the death rate to0\.700\.70–0\.820\.82\(reckless approach to lava\), yet adding*any*impulsivity collapses it to0\.000\.00\(maximum interaction residual0\.820\.82\)\. Steep discounting protects the manic agent from its own risk\-taking, because dying in lava requires a committed multi\-step approach that a myopic agent will not undertake, a non\-obvious prediction that follows from the interaction of two mechanisms, not from either alone\. For*anxiety×\{\\times\}depression*, severe depression overrides the anxiety readout: once effort cost is high the agent stops acting altogether, so avoidance can no longer be expressed \(residual0\.500\.50\)\. Comorbidity is thus not simply additive in this model, matching the clinical intuition that co\-occurring conditions modify one another’s expression\.
Figure 6:Comorbidity: joint dose grids for two knob pairs \(mean readout per cell\)\. Both are strongly nonadditive; the reported interaction is the maximum residual against the additive prediction\. Impulsivity suppresses the lethality of mania \(right\)\.
##### Correspondence with known clinical phenomena\.
Because the disorder knobs are grounded but the*consequences*of learning under them are not designed in, several results reproduce documented effects without being fitted to them\. \(i\) The rescue dissociation \(avoidance in anxiety and PTSD resisting passive removal of the stressor\) parallels the clinical observation that avoidance is self\-maintaining and extinction\-resistant, the rationale for exposure\-based therapy\[[Craske et al\. 2014](https://arxiv.org/html/2607.07753#bib.bibx4),[Milad and Quirk 2012](https://arxiv.org/html/2607.07753#bib.bibx11)\]\. \(ii\) Addiction shows a vulnerability threshold \(drug occupancy jumps nearϵ=0\.05\\epsilon\{=\}0\.05\) and then escalation to the exclusion of natural reward, the qualitative signature of computational and animal models of addiction\[[Redish 2004](https://arxiv.org/html/2607.07753#bib.bibx15),[Ahmed and Koob 1998](https://arxiv.org/html/2607.07753#bib.bibx1)\]\. \(iii\) Impulsivity produces a discounting crossover, agents abandon the larger later reward as the discount steepens, the defining marker of impulsive choice\[[Ainslie 1975](https://arxiv.org/html/2607.07753#bib.bibx2),[Kirby, Petry, and Bickel 1999](https://arxiv.org/html/2607.07753#bib.bibx7)\]\. \(iv\) Depression manifests as reduced willingness to exert effort for reward rather than an inability to act, matching effort\-based accounts of anhedonia\[[Treadway and Zald 2011](https://arxiv.org/html/2607.07753#bib.bibx19)\]\. \(v\) Mania appears as reduced harm sensitivity with elevated reward pursuit, consistent with reward\-hypersensitivity models\[[Johnson et al\. 2012](https://arxiv.org/html/2607.07753#bib.bibx6)\]\. These correspondences are qualitative, we do not fit human data, but they distinguish the framework from a mere re\-description: the same simple manipulations recover effects that clinical theory independently predicts\.
##### Generalisation to 3D pixel observations\.
To test whether the disorder phenotypes are artefacts of the grid\-world setting or reflect deeper structural properties of the knobs, we transfer all three evaluated disorders \(depression, addiction, anxiety\) to a three\-dimensional first\-person environment \(MiniWorld\[[Gym\-MiniWorld 2023](https://arxiv.org/html/2607.07753#bib.bibx21)\]\) in which a standard Nature CNN agent\[[Mnih et al\. 2015](https://arxiv.org/html/2607.07753#bib.bibx25)\]receives raw pixel observations with no appraisal critic\. Three doses per disorder, 3 seeds each, 3 M steps per run \(36 runs total on a 4×\\timesGPU instance\)\.
Depressionreplicates exactly: forward\-action fraction collapses from0\.730\.73\(baseline\) to0\.720\.72,0\.00040\.0004, and0\.00010\.0001at dosesϵ=0\.01,0\.03,0\.1\\epsilon\\\!=\\\!0\.01,\\,0\.03,\\,0\.1respectively \(means over 3 seeds\), and drug occupancy remains exactly0\.000\.00at every dose—zero cross\-contamination, matching the grid\-world dissociation\.
Addictionshows compulsive drug\-seeking that generalises cleanly: drug occupancy rises from0\.000\.00\(baseline\) to0\.790\.79\(ϵ=0\.1\\epsilon\\\!=\\\!0\.1\) and0\.780\.78\(ϵ=0\.3\\epsilon\\\!=\\\!0\.3\)\. At the highest dose \(ϵ=0\.6\\epsilon\\\!=\\\!0\.6\) drug occupancy falls slightly to0\.740\.74while forward fraction rises to0\.650\.65—the agent moves*more*, not less\. We interpret this as a motivational saturation effect: the drug signal dominates so strongly that the policy becomes a persistent, undirected seeking behaviour rather than a goal\-directed route to the drug tile; locomotion increases but efficient drug\-finding deteriorates\. This pattern has a clinical analogue in severe addiction, where compulsive motivation drives frantic seeking activity that paradoxically reduces the probability of obtaining the substance\[[Ahmed and Koob 1998](https://arxiv.org/html/2607.07753#bib.bibx1)\]\. The non\-monotonicity at the top dose is therefore not a failure of the mechanism but an emergent consequence of crossing a motivational threshold beyond which the seeking behaviour itself is impaired—consistent with escalation models in which intake regulation breaks down at extreme doses\.
Anxietyconfirms mechanism specificity: the CNN baseline already routes entirely around the threat tile \(safe\-route fraction1\.001\.00at baseline\), and the anxiety knob*preserves*this avoidance at all three doses \(ϵ=0\.1,0\.3,0\.6\\epsilon\\\!=\\\!0\.1,\\,0\.3,\\,0\.6\) without degrading forward locomotion \(0\.730\.73–0\.710\.71\)\. The critical dissociation check passes: drug occupancy remains0\.000\.00across all anxiety doses, confirming the knob acts only on the intended threat\-avoidance channel\.
Cross\-assay dissociationholds across both domains and all three disorders: the depression knob never raises drug occupancy or safe\-route fraction above0\.000\.00; the addiction knob does not collapse locomotion; the anxiety knob does not raise drug occupancy or suppress locomotion\. The fact that a CNN agent with no access to appraisal signals replicates the same phenotypic signatures confirms that the disorder properties are properties of the*reward structure*, not of the particular observation encoding or policy architecture\.
Combining the 3D generalisation result with the full seven\-disorder 2D battery, Figure[7](https://arxiv.org/html/2607.07753#Sx5.F7)plots the block\-diagonal dissociation matrix across all disorders and all primary assays\. Each row \(one disorder at high dose\) lights up only its own column \(the designated assay\), with off\-diagonal cells near zero\. This pattern holds despite the assays being measured in five distinct environments with different dynamics, observation spaces, and episode structures—confirming that the mechanism specificity is a property of the knob design, not of shared environmental confounds\.
Figure 7:Block\-diagonal cross\-assay dissociation across all seven disorders\. Rows are disorder conditions \(at high dose\); columns are the primary assay metric of each disorder\. Values are normalised effects relative to the healthy baseline; blue boxes mark the expected diagonal\. Off\-diagonal cells are near zero, confirming that each knob selectively impairs its designated symptom without bleeding into other disorder signatures\.
## Discussion
Appraisal weights form a compact, controllable basis for a space of affective phenotypes: seven grounded knobs reproduce qualitatively distinct disorders and place them in mutually consistent positions, with severity tunable continuously\. The mania–anxiety symmetry supports a dimensional rather than categorical view of affect, consistent with transdiagnostic frameworks in which conditions are points along shared axes rather than discrete kinds\. The rescue dissociation offers a mechanistic reading of treatment resistance: pathologies that persist as behavioural habits are self\-reinforcing precisely because avoidance and myopia prevent the experiences that would disconfirm them, the RL analogue of why exposure, not avoidance of stressors, extinguishes clinical avoidance\. That this falls out of a value\-learning agent, without any bespoke “memory” of the trauma, is itself informative\.
## Limitations and Conclusion
We have shown that appraisal\-guided RL yields a controllable, grounded, dose\-dependent space of disorder\-like phenotypes in which induction, structure, and treatment can all be studied in a single agent, and that the seven disorders organise, dissociate under treatment, and interact in ways that are emergent rather than designed\. Three limitations bound these claims\. First, our disorders are computational analogies, not clinical equivalences: the assays map to recognised paradigms and reproduce documented effects, but validating the labels against human or animal data remains open\. Second, behavioural coordinates in the affective space are normalised deviations, and a two\-dimensional projection omits dimensions such as compulsivity, which we mark explicitly\. Third, two appraisal signals—certainty \(policy entropy\) and anticipation \(NRE error\)—are themselves functions of the evolving policy and NRE network, making the shaped reward nonstationary in the same way as curiosity\-driven methods such as ICM and RND; standard PPO convergence guarantees do not strictly apply, though the disorder phenotypes stabilise empirically within the training budget\. Our generalisation experiment \(Sec\. 5\) begins to address the concern that the findings are specific to grid worlds or to the appraisal critic, by transferring three knobs to a pixel\-based arcade agent; extending the appraisal\-based knobs to such domains, together with comorbidity beyond two knobs, active\-exposure therapy at scale, and quantitative fits to behavioural data, are the natural next steps\. We see the controllability of the framework, one grounded knob per disorder with severity tunable continuously, as its most useful property: it turns disorder modelling from post\-hoc description into a manipulable, falsifiable science of affective phenotypes\.
## References
- \[Ahmed and Koob 1998\]Ahmed, S\. H\.; and Koob, G\. F\. 1998\. Transition from moderate to excessive drug intake: change in hedonic set point\.*Science*282\(5387\):298–300\.
- \[Ainslie 1975\]Ainslie, G\. 1975\. Specious reward: A behavioral theory of impulsiveness and impulse control\.*Psychological Bulletin*82\(4\):463–496\.
- \[Prasad, Jacob, and Ahamed 2024\]Prasad, H\.; Jacob, C\.; and Ahamed, I\. 2024\. Appraisal\-Guided Proximal Policy Optimization: Modeling Psychological Disorders in Dynamic Grid World\.*arXiv:2407\.20383*\.
- \[Craske et al\. 2014\]Craske, M\. G\.; Treanor, M\.; Conway, C\. C\.; Zbozinek, T\.; and Vervliet, B\. 2014\. Maximizing exposure therapy: An inhibitory learning approach\.*Behaviour Research and Therapy*58:10–23\.
- \[Huys et al\. 2016\]Huys, Q\. J\. M\.; Maia, T\. V\.; and Frank, M\. J\. 2016\. Computational psychiatry as a bridge from neuroscience to clinical applications\.*Nature Neuroscience*19\(3\):404–413\.
- \[Johnson et al\. 2012\]Johnson, S\. L\.; Edge, M\. D\.; Holmes, M\. K\.; and Carver, C\. S\. 2012\. The behavioral activation system and mania\.*Annual Review of Clinical Psychology*8:243–267\.
- \[Kirby, Petry, and Bickel 1999\]Kirby, K\. N\.; Petry, N\. M\.; and Bickel, W\. K\. 1999\. Heroin addicts have higher discount rates for delayed rewards than non\-drug\-using controls\.*Journal of Experimental Psychology: General*128\(1\):78–87\.
- \[Kullback and Leibler 1951\]Kullback, S\.; and Leibler, R\. A\. 1951\. On information and sufficiency\.*Annals of Mathematical Statistics*22\(1\):79–86\.
- \[Lazarus 1991\]Lazarus, R\. S\. 1991\.*Emotion and Adaptation*\. Oxford University Press\.
- \[Maia and Frank 2011\]Maia, T\. V\.; and Frank, M\. J\. 2011\. From reinforcement learning models to psychiatric and neurological disorders\.*Nature Neuroscience*14\(2\):154–162\.
- \[Milad and Quirk 2012\]Milad, M\. R\.; and Quirk, G\. J\. 2012\. Fear extinction as a model for translational neuroscience: Ten years of progress\.*Annual Review of Psychology*63:129–151\.
- \[Moerland et al\. 2018\]Moerland, T\. M\.; Broekens, J\.; and Jonker, C\. M\. 2018\. Emotion in reinforcement learning agents and robots: A survey\.*Machine Learning*107\(2\):443–480\.
- \[Montague et al\. 2012\]Montague, P\. R\.; Dolan, R\. J\.; Friston, K\. J\.; and Dayan, P\. 2012\. Computational psychiatry\.*Trends in Cognitive Sciences*16\(1\):72–80\.
- \[Rachman 2002\]Rachman, S\. 2002\. A cognitive theory of compulsive checking\.*Behaviour Research and Therapy*40\(6\):625–639\.
- \[Redish 2004\]Redish, A\. D\. 2004\. Addiction as a computational process gone awry\.*Science*306\(5703\):1944–1947\.
- \[Scherer 2001\]Scherer, K\. R\. 2001\. Appraisal considered as a process of multilevel sequential checking\. In*Appraisal Processes in Emotion*, 92–120\. Oxford University Press\.
- \[Schulman et al\. 2017\]Schulman, J\.; Wolski, F\.; Dhariwal, P\.; Radford, A\.; and Klimov, O\. 2017\. Proximal policy optimization algorithms\.*arXiv:1707\.06347*\.
- \[Sequeira et al\. 2011\]Sequeira, P\.; Melo, F\. S\.; and Paiva, A\. 2011\. Emotion\-based intrinsic motivation for reinforcement learning agents\. In*ACII*, 326–336\.
- \[Treadway and Zald 2011\]Treadway, M\. T\.; and Zald, D\. H\. 2011\. Reconsidering anhedonia in depression: Lessons from translational neuroscience\.*Neuroscience & Biobehavioral Reviews*35\(3\):537–555\.
- \[Chevalier\-Boisvert et al\. 2023\]Chevalier\-Boisvert, M\.; Dai, B\.; Towers, M\.; de Lazcano, R\.; Willems, L\.; Lahlou, S\.; Pal, S\.; Castro, P\. S\.; and Terry, J\. 2023\. Minigrid & Miniworld: Modular & Customizable Reinforcement Learning Environments for Goal\-Oriented Tasks\.*arXiv:2306\.13831*\.
- \[Gym\-MiniWorld 2023\]Chevalier\-Boisvert, M\. 2018–2023\. MiniWorld: Minimalistic 3D Interior Environment Simulator\.*GitHub: Farama\-Foundation/miniworld*\.
- \[Brockman et al\. 2016\]Brockman, G\.; Cheung, V\.; Pettersson, L\.; Schneider, J\.; Schulman, J\.; Tang, J\.; and Zaremba, W\. 2016\. OpenAI Gym\.*arXiv:1606\.01540*\.
- \[Huang et al\. 2022\]Huang, S\.; Dossa, R\. F\. J\.; Ye, C\.; Braga, J\.; Chakraborty, D\.; Mehta, K\.; and Araújo, J\. G\. 2022\. CleanRL: High\-quality single\-file implementations of deep reinforcement learning algorithms\.*JMLR*23\(274\):1–18\.
- \[Ziegler et al\. 2019\]Ziegler, D\. M\.; Stiennon, N\.; Wu, J\.; Brown, T\. B\.; Radford, A\.; Amodei, D\.; Christiano, P\.; and Irving, G\. 2019\. Fine\-tuning language models from human preferences\.*arXiv:1909\.08593*\.
- \[Mnih et al\. 2015\]Mnih, V\.; Kavukcuoglu, K\.; Silver, D\.; Rusu, A\. A\.; Veness, J\.; Bellemare, M\. G\.; Graves, A\.; Riedmiller, M\.; Fidjeland, A\. K\.; Ostrovski, G\.; et al\. 2015\. Human\-level control through deep reinforcement learning\.*Nature*518\(7540\):529–533\.
- \[Sutton and Barto 2018\]Sutton, R\. S\.; and Barto, A\. G\. 2018\.*Reinforcement Learning: An Introduction*, 2nd ed\. MIT Press\.
- \[LeCun et al\. 2015\]LeCun, Y\.; Bengio, Y\.; and Hinton, G\. 2015\. Deep learning\.*Nature*521\(7553\):436–444\.
- \[Williams 1992\]Williams, R\. J\. 1992\. Simple statistical gradient\-following algorithms for connectionist reinforcement learning\.*Machine Learning*8\(3\):229–256\.
- \[Insel et al\. 2010\]Insel, T\.; Cuthbert, B\.; Garvey, M\.; Heinssen, R\.; Pine, D\. S\.; Quinn, K\.; Sanislow, C\.; and Wang, P\. 2010\. Research domain criteria \(RDoC\): Toward a new classification framework for research on mental disorders\.*American Journal of Psychiatry*167\(7\):748–751\.
- \[Beck 1979\]Beck, A\. T\. 1979\.*Cognitive Therapy of Depression*\. Guilford Press\.
- \[Pitman et al\. 2012\]Pitman, R\. K\.; Rasmusson, A\. M\.; Koenen, K\. C\.; Shin, L\. M\.; Orr, S\. P\.; Gilbertson, M\. W\.; Milad, M\. R\.; and Liberzon, I\. 2012\. Biological studies of post\-traumatic stress disorder\.*Nature Reviews Neuroscience*13\(11\):769–787\.
- \[Chamberlain et al\. 2008\]Chamberlain, S\. R\.; Menzies, L\.; Hampshire, A\.; Suckling, J\.; Fineberg, N\. A\.; del Campo, N\.; Aitken, M\.; Craig, K\.; Owen, A\. M\.; Bullmore, E\. T\.; Robbins, T\. W\.; and Sahakian, B\. J\. 2008\. Orbitofrontal dysfunction in patients with obsessive\-compulsive disorder and their unaffected relatives\.*Science*321\(5887\):421–422\.
- \[Robinson and Berridge 1993\]Robinson, T\. E\.; and Berridge, K\. C\. 1993\. The neural basis of drug craving: An incentive salience theory of addiction\.*Brain Research Reviews*18\(3\):247–291\.
This appendix provides the full experimental record behind the main paper\. It first explains how the cognitive appraisals were formulated and how the experiments were set up and calibrated, then gives every environment and assay in detail, and finally reports the complete result tables \(dose\-response across all environments, secondary assays, the full control battery, the affective\-space coordinates, and the full rescue, exposure, and comorbidity statistics\) with a walkthrough of how to read each one\. All numbers are computed from the same 1,375\-run corpus by a single analysis script, and figures reuse the main\-paper plots where possible to avoid duplication\.
### Formulation of the Cognitive Appraisals
Appraisal theory holds that emotion arises from a small set of domain\-independent evaluations of an event with respect to an agent’s goals\[[Lazarus 1991](https://arxiv.org/html/2607.07753#bib.bibx9),[Scherer 2001](https://arxiv.org/html/2607.07753#bib.bibx16)\]\. We operationalise six of these dimensions so that each is \(i\) computable online from quantities the agent already has, \(ii\) bounded to\(0,1\)\(0,1\)so they compose into a single vector, and \(iii\) mapped to a recognised appraisal construct\. Two are primary \(goal\-relevance\) appraisals, one is a secondary \(coping\) appraisal, two are information\-theoretic readings of the agent’s own policy, and one is predictive\.
- •Motivational relevanceζMR\\zeta\_\{\\text\{MR\}\}andgoal congruenceζGC\\zeta\_\{\\text\{GC\}\}are primary appraisals of how much the current state bears on the goal\. We use the complement of the \(Manhattan\) goal distance for relevance and of the \(Euclidean\) goal distance for congruence; the two distances give a coarse and a fine reading of proximity, and both are normalised by the largest attainable distance so they saturate near the goal\.
- •Coping potentialζCP\\zeta\_\{\\text\{CP\}\}is a secondary appraisal of perceived control\. We read it from threat visibility: the fraction of known threats*outside*the agent’s egocentric view, so that a threat entering view lowers coping potential\. This is the appraisal the anxiety and mania knobs act on \(penalising low or high coping potential respectively\)\.
- •CertaintyζC\\zeta\_\{\\text\{C\}\}andnoveltyζN\\zeta\_\{\\text\{N\}\}are read from the actor’s action distributionpp\. Certainty is the complement of the \(squashed\) policy entropy, so a confident, low\-entropy policy appraises high certainty\. Novelty is the \(squashed\) KL divergence ofppfrom the uniform policy, capturing how far the current decision departs from indifference\. Both use the squashingx↦x/\(1\+x\)x\\mapsto x/\(1\+x\)to map an unbounded information quantity into\(0,1\)\(0,1\)\.
- •AnticipationζA\\zeta\_\{\\text\{A\}\}is the complement of the next\-reward\-estimation \(NRE\) error: a small three\-layer network predictsrtr\_\{t\}from\(ot−1,at−1\)\(o\_\{t\-1\},a\_\{t\-1\}\), and accurate prediction yields high anticipation\. This gives the agent a forward\-looking appraisal grounded in its own reward model\.
The exact formulas are given in the main paper \(Eqs\. 3–8\)\. The six values are concatenated to the critic \(so value estimation is appraisal\-informed\) and, for anxiety and mania, drive the reward\-shaping term; the stress index reported as a secondary assay is the weighted deviation∑i\(1−ζi\)wi\\sum\_\{i\}\(1\-\\zeta\_\{i\}\)w\_\{i\}\.
### Disorder Knobs, Dose Grids, and Hyperparameters
Table[5](https://arxiv.org/html/2607.07753#A0.T5)gives each knob variable and its dose grid; Table[6](https://arxiv.org/html/2607.07753#A0.T6)gives the shared learning hyperparameters\. The checking bonus isbchk=ϵλkb\_\{\\text\{chk\}\}=\\epsilon\\lambda^\{k\}on thekk\-th checkpoint return within an episode \(λ=0\.5\\lambda=0\.5\), which bounds the episodic total byϵ/\(1−λ\)\\epsilon/\(1\-\\lambda\)and converts an all\-or\-nothing incentive into a graded number of checks \(Sec\. A\.10 explains why this habituation is necessary\)\.
Table 5:Knob variable and dose grid per disorder\. The impulsivity dose is a discount reduction,γ=1−ϵ\\gamma=1\-\\epsilon\.Table 6:PPO / AG\-PPO hyperparameters \(shared across all runs unless a disorder knob overridesγ\\gamma\)\.
### Experimental Setup and Protocol
The agent is a CleanRL\-style\[[Huang et al\. 2022](https://arxiv.org/html/2607.07753#bib.bibx23)\]PPO implementation with the appraisal\-informed critic and NRE network described above, running on 8 synchronous vectorised environments with a rollout of 128 steps\. Training uses the hyperparameters of Table[6](https://arxiv.org/html/2607.07753#A0.T6); the per\-environment budget is 600k steps, raised to 1M for LavaCrossing, which mixes lava and needs longer to solve\.
*Evaluation\.*After training, each model is evaluated for 40 episodes under the stochastic policy on held\-out seeds disjoint from training\. Every assay in this paper is computed from these evaluation episodes; primary and secondary assays are averaged over the 40 episodes and then over seeds, and we report means with95%95\\%confidence intervals across seeds\.
*Corpus\.*The full corpus is 1,375 runs: \(i\) dose\-response, seven disorders across the environments in which their primary assay is defined, with four control conditions per threat environment, 10 seeds each and 30 for anxiety; \(ii\) rescue and its matched control, 5 seeds per disorder; \(iii\) the exposure curriculum, 15 seeds for anxiety and 10 for PTSD; and \(iv\) two comorbidity dose grids of 90 runs each\. Runs are independent and were executed in parallel on a 208\-vCPU cloud machine and locally; the sweep is resumable and every run writes a self\-describing result file, from which all tables and figures are regenerated by one script\.
*Convergence criterion\.*Because avoidance can trivially reduce task success, symptom analysis is restricted, by a criterion fixed before analysis, to runs that solve the task \(success≥50%\\geq 50\\%\); the number of contributing seeds is reported alongside every aggregate\.
### Hyperparameter and Mechanism Calibration
Two kinds of calibration were needed: learning hyperparameters \(so that baseline PPO solves every environment, giving each disorder a healthy reference\), and the scale of each new reward mechanism \(so that the knob produces a graded, rather than degenerate, dose\-response\)\.
*Learning hyperparameters\.*The common PPO defaults\(lr=2\.5×10−4,entropy=0\.01\)\(\\text\{lr\}=2\.5\\times 10^\{\-4\},\\text\{entropy\}=0\.01\)leave the agent in a freeze\-in\-place local optimum on Dynamic\-Obstacles: it rotates without advancing and times out, so the baseline never solves the task\. A small grid overlr∈\{2\.5×10−4,1×10−3\}\\text\{lr\}\\in\\\{2\.5\\times 10^\{\-4\},1\\times 10^\{\-3\}\\\}andentropy∈\{0\.01,0\.03,0\.05\}\\text\{entropy\}\\in\\\{0\.01,0\.03,0\.05\\\}showed that\(1×10−3,0\.03\)\(1\\times 10^\{\-3\},0\.03\)solves all four threat environments; we adopted it everywhere\. LavaCrossing additionally required the larger 1M\-step budget\.
*Checking mechanism\.*An early OCD mechanism that penalised low certainty reduced checking rather than inducing it \(Sec\. A\.10\)\. A checkpoint\-return bonus without habituation was bistable: below a threshold it was ignored, above it the bonus became an unbounded loop that captured the agent\. We calibrated the habituation factorλ=0\.5\\lambda=0\.5and the dose grid\{0\.1,0\.2,0\.4,0\.6\}\\\{0\.1,0\.2,0\.4,0\.6\\\}by sweeping candidate values and selecting the range that produced rising checking with maintained task success\.
*Addiction and exposure scales\.*The drug bonus was calibrated by sweeping\{0\.01,0\.02,0\.05,0\.1\}\\\{0\.01,0\.02,0\.05,0\.1\\\}and locating the vulnerability threshold at which drug occupancy overtakes the goal \(nearϵ=0\.05\\epsilon=0\.05\)\. For the exposure curriculum we first tried rewarding the agent for reaching the feared route; this failed because an avoidant policy never reaches it, so we switched to response prevention \(a penalty on the avoidance route, annealed to zero\), with the penalty coefficient set to0\.50\.5after a short sweep\.
### Environments in Detail
The seven environments are rendered in Fig\.[8](https://arxiv.org/html/2607.07753#A0.F8)\. The four threat/spatial grids are Dynamic\-Obstacles \(moving obstacles\), LavaGap and LavaCrossing \(lava hazards\), and the custom Approach\-Avoidance conflict, which places a high\-value goal behind a lava\-lined corridor where every approach step is threat\-adjacent and a low\-value goal \(worth0\.15×0\.15\\times\) with an open approach at equal distance, so route choice isolates threat sensitivity\. Temporal\-Choice is a corridor with a near reward0\.50\.5at distance∼3\{\\sim\}3and a far reward1\.01\.0at distance∼9\{\\sim\}9; the discount crossover isγ≈0\.89\\gamma\\approx 0\.89\. Addiction places a repeatable drug tile in one corner with the goal in the opposite corner\. Trauma places a conditioned\-shock tile on the short route to the goal, with a long safe detour available\. All observations are7×7×37\\times 7\\times 3egocentric symbolic tensors with three actions\.
Figure 8:The seven environments: four threat/spatial grids and three custom environments realising delay discounting, drug self\-administration, and conditioned\-trauma avoidance\.
### Full Behavioural Assay Battery
Primary assays are defined in the main text\. Secondary assays, computed per episode and averaged over the 40 evaluation episodes: thigmotaxis index \(fraction of steps wall\-adjacent\), edge occupancy, mean and near threat distance, freezing fraction \(runs of≥3\\geq 3consecutive turn actions\), turnaround rate, action stereotypy \(repeated action trigrams\), revisit and checking rates, visitation entropy \(normalised\), drug occupancy, trauma distance, and the stress index∑i\(1−ζi\)wi\\sum\_\{i\}\(1\-\\zeta\_\{i\}\)w\_\{i\}with weights\(0\.25,0\.05,0\.1,0\.2,0\.35,0\.05\)\(0\.25,0\.05,0\.1,0\.2,0\.35,0\.05\)carried over from the base model for continuity\.
### Full Dose\-Response Across Environments
Table[7](https://arxiv.org/html/2607.07753#A0.T7)reports every disorder’s primary assay in each environment where that assay is defined, expanding Table 2 of the main paper\. To read it: each row is a disorder\-environment pair, the PPO column is the untreated baseline, andϵ1\\epsilon\_\{1\}toϵ4\\epsilon\_\{4\}are the four doses of Table[5](https://arxiv.org/html/2607.07753#A0.T5); a monotone increase \(or decrease, for depression’s forward\-action fraction\) from PPO throughϵ4\\epsilon\_\{4\}is the dose\-response\. Mania, OCD, and depression were induced on all four threat grids and are monotone on each, confirming the effect is not specific to one layout\. Anxiety’s primary assay \(risky\-goal choice\) is only defined on Approach\-Avoidance, which has the two\-goal conflict; on the other threat grids anxiety instead elevates thigmotaxis and threat distance \(Table[8](https://arxiv.org/html/2607.07753#A0.T8)\)\.
Table 7:Full dose–response of each disorder’s primary assay across every environment it was run in \(mean±\\pm95% CI\)\. Blank where a knob/env pair was not run\. This expands Table 2 of the main paper\.
### Secondary Behavioural Assays
Table[8](https://arxiv.org/html/2607.07753#A0.T8)reports the secondary assays at each disorder’s severe dose\. These reveal the behavioural texture behind the primary readout: each column is one assay, and the final row is the PPO baseline for reference\. Anxiety and mania saturate thigmotaxis; depression drives freezing to near unity; OCD elevates stereotypy and revisits; PTSD produces long detours with high visitation entropy\.
Table 8:Secondary behavioural assays at the severe dose of each disorder \(mean, on the environment used in the main text\)\. Values reveal the behavioural texture beyond the primary assay\.
### Complete Control Battery
Table[10](https://arxiv.org/html/2607.07753#A0.T10)extends the two\-environment control comparison of the main paper to all four threat environments\. Each block is one environment; the columns are the four controls \(standard PPO, a critic\-noise control, PPO\+RND, and the appraisal critic without shaping\)\. No control reproduces a disorder phenotype: success stays high and threat distance stays baseline\-like\. The one exception is RND, which fails on both lava environments by over\-exploring into the hazard, a reminder that generic novelty\-seeking is not a model of any disorder\.
Table 9:All four controls across the four threat/spatial environments \(mean success and mean threat distance, 10 seeds\)\. No control produces a disorder phenotype; RND fails on the lava environments by over\-exploring\.
Table 10:Affective\-space coordinates \(normalised behavioural deviations from the healthy baseline\) at the severe dose\. Reward–approachxxand threat–avoidanceyyas plotted in main\-paper Fig\. 4\.
### The Negative Result and Mechanism Redesign
Our first OCD mechanism penalised low certainty\. It reduced checking rather than inducing it, because rewarding decisiveness suppresses re\-verification: at the strongest certainty penalty the checking rate fell below baseline\. A checkpoint bonus without habituation was then bistable\. Below a threshold the agent ignored it and solved the task normally; above the threshold the bonus became an unbounded loop that captured the agent entirely, collapsing task success to zero with no graded regime in between\. Introducing diminishing reassurance \(λk\\lambda^\{k\}decay,λ=0\.5\\lambda=0\.5\) bounded the episodic bonus and produced the graded checking of the main paper\. We report both failures because they clarify what does and does not induce compulsion: a mechanistic account \(diminishing reassurance\) succeeds where a cosmetic reward does not\.
### Affective\-Space Construction
The two\-dimensional affective space places each disorder by normalised behavioural deviations from its environment’s healthy baseline\. The reward\-approach axisxxcombines reward\-pursuit markers \(forward\-action fraction for depression, drug occupancy for addiction, near\-reward choice for impulsivity, risk\-taking for mania\); the threat\-avoidance axisyyuses threat and trauma distance \(positive for anxiety and PTSD, negative for mania\)\. Each axis is normalised to\[−1,1\]\[\-1,1\]by its maximum absolute deviation\. Table[10](https://arxiv.org/html/2607.07753#A0.T10)lists the severe\-dose coordinates plotted in main\-paper Fig\.[4](https://arxiv.org/html/2607.07753#Sx5.F4)\. OCD lies near the origin because its compulsivity is an out\-of\-plane dimension, indicated separately in the figure\.
### Rescue and Extinction, Full Statistics
Table[11](https://arxiv.org/html/2607.07753#A0.T11)gives the treated\-versus\-control endpoints with confidence intervals for all seven disorders; the full recovery trajectories are in main\-paper Fig\.[5](https://arxiv.org/html/2607.07753#Sx5.F5)\. To read the table:*Treated*removes the pathological knob,*Control*keeps it with matched extra training, andnnis the number of contributing seeds\. The dissociation is clean: mania, OCD, and addiction remit under passive removal \(treated near zero, control retaining the symptom\), whereas anxiety, PTSD, depression, and impulsivity resist, because the safe or myopic policy keeps working and the agent never re\-encounters the disconfirming evidence\. The contrast between passive\-remitting and exposure\-resistant disorders has clinical significance: it parallels the established finding that extinction\-based therapies succeed for OCD and addiction but fail for avoidance\-maintained conditions unless the exposure itself is engineered to override the avoidance response\. The model captures this mechanistically—no special treatment logic is required; the dissociation arises from the agent’s own policy structure and the topology of the environment\.
Table 11:Rescue/extinction, full statistics \(mean±\\pm95% CI\)\.*Treated*: knob removed\.*Control*: knob kept, matched extra training\.nnis the number of seeds\.
### Exposure Therapy, Full Statistics
Table[12](https://arxiv.org/html/2607.07753#A0.T12)reports the graded exposure result for the two resistant avoidance disorders\. Exposure uses response prevention: a penalty on the avoidance route \(the safe alternative\) annealed to zero over training\. A simple reward for reaching the feared route fails, because an avoidant policy never goes there to collect it, so the intervention must act on the avoidance response itself\. Evaluated afterwards on the unmodified environment, anxiety recovers feared\-route choice to0\.930\.93\(15 seeds\) and PTSD to0\.900\.90\(10 seeds\), from0\.200\.20under passive removal, and the recovery persists after the exposure prompt has faded\.
Table 12:Graded\-exposure curriculum for the resistant avoidance disorders \(mean±\\pm95% CI\)\. Feared\-route choice after exposure, compared with passive removal and the severe untreated model\.
### Comorbidity, Full Grids
Tables[13](https://arxiv.org/html/2607.07753#A0.T13)and[14](https://arxiv.org/html/2607.07753#A0.T14)give the complete joint dose grids, visualised in main\-paper Fig\.[6](https://arxiv.org/html/2607.07753#Sx5.F6)\. Each cell is the mean readout for a \(row\-dose, column\-dose\) pair; strong departure from the sum of the single\-knob effects along the edges is the emergent interaction\. For mania and impulsivity the interaction is striking: mania alone drives the death rate to0\.700\.70to0\.820\.82, but adding any impulsivity collapses it to zero, because dying in lava requires a committed multi\-step approach that a myopic agent will not undertake\. For anxiety and depression, severe depression overrides the anxiety readout: once effort cost is high the agent stops acting altogether, so avoidance can no longer be expressed\.
Table 13:Comorbidity grid for anxiety×\\timesdepression \(mean of the readout, 10 seeds per cell\)\. Rows: depression dose; columns: anxiety dose\. Strong nonadditivity\.Table 14:Comorbidity grid for mania×\\timesimpulsivity \(mean of the readout, 10 seeds per cell\)\. Rows: impulsivity dose; columns: mania dose\. Strong nonadditivity\.
### Per\-Disorder Occupancy and Qualitative Observations
Main\-paper Fig\.[2](https://arxiv.org/html/2607.07753#Sx4.F2)shows state\-occupancy at each disorder’s severe dose\. The spatial signatures are legible and distinct: anxiety traces a wall\-hugging avoidance band; mania concentrates against the lava; OCD forms a looping checking pattern near the checkpoint; depression is near\-stationary at the start; impulsivity fixates on the near reward; addiction pins to the drug corner; PTSD makes a wide detour around the trauma tile\. These qualitative readings match the quantitative assays \(thigmotaxis, checking rate, forward fraction, drug occupancy, trauma distance\) reported above, and correspond one\-to\-one to the occupancy maps of main\-paper Fig\.[2](https://arxiv.org/html/2607.07753#Sx4.F2)\.
The signatures are disorder\-specific rather than task\-specific: the four control agents \(PPO, critic\-noise, PPO\+RND, appraisal critic without shaping\) run on the same environments and produce uniform or random occupancy patterns with no spatial clustering\. The phenotype gallery therefore provides face validity at a glance: an anxiety\-like agent crowds the periphery, a PTSD\-like agent takes the long safe detour, a depression\-like agent stalls at the start, and an OCD\-like agent loops near the checkpoint even when the task reward lies beyond it\. These patterns are emergent—no explicit spatial objective was specified; the knob acts only on the scalar reward signal, and the spatial structure arises from the agent’s learned policy\.
### 3D Pixel Generalisation: Setup and Full Results
#### Environment\.
We use MiniWorld \(https://github\.com/Farama\-Foundation/miniworld\), a first\-person 3D environment rendered with OpenGL/EGL\. Each episode places the agent in a room with a goal object \(green box\) and, depending on mode, a drug object \(purple box, addiction mode\) or a risky goal and three threat boxes \(anxiety mode\)\. Observations are60×80×360\{\\times\}80\{\\times\}3RGB pixels\. Actions are discrete \(turn left, turn right, move forward\)\. Episode length is capped at 250 steps; reaching the goal gives\+1\+1\(depression/addiction modes\) or\+0\.3\+0\.3\(anxiety safe goal\)\. Reaching the drug object triggers the addiction shaping bonus but gives no task reward\.
#### Agent and knobs\.
Standard Nature CNN \(three conv layers: 32/64/64 filters,8/4/38/4/3kernels, stride4/2/14/2/1\) followed by a linear layer \(512 units\) and separate actor/critic heads\.*No appraisal critic\.*Disorder knobs are computed from ground\-truth environment state and applied as reward shaping only:
Depression:r~t=rt−ε⋅𝟏\[action=forward\]\\displaystyle\\tilde\{r\}\_\{t\}=r\_\{t\}\-\\varepsilon\\cdot\\mathbf\{1\}\[\\text\{action\}=\\text\{forward\}\]\(9\)Addiction:r~t=rt\+ε⋅𝟏\[on\_drug\]\\displaystyle\\tilde\{r\}\_\{t\}=r\_\{t\}\+\\varepsilon\\cdot\\mathbf\{1\}\[\\text\{on\\\_drug\}\]\(10\)Anxiety:r~t=rt−ε⋅\(1−coping\_potential\)\\displaystyle\\tilde\{r\}\_\{t\}=r\_\{t\}\-\\varepsilon\\cdot\(1\-\\text\{coping\\\_potential\}\)\(11\)where coping\_potential is the fraction of threat boxes*not*in the agent’s45∘45^\{\\circ\}forward field of view\. PPO hyperparameters: learning rate3×10−43\{\\times\}10^\{\-4\}, clipεclip=0\.2\\varepsilon\_\{\\text\{clip\}\}\{=\}0\.2, 8 parallel SyncVectorEnv workers, 128\-step rollouts, 4 epochs, GAEλ=0\.95\\lambda\{=\}0\.95,γ=0\.99\\gamma\{=\}0\.99, entropy coefficient0\.010\.01\.
#### Compute\.
36 runs \(3 disorders×\\times4 conditions×\\times3 seeds\), 3 M steps each, on a 4×\\timesRTX 5090 instance \(≈\\approx2\.5 hours wall\-clock\)\. EGL offscreen rendering \(PYGLET\_HEADLESS=1\) with SyncVectorEnv \(AsyncVectorEnv causes heap corruption on EGL fork\)\. PyTorch nightly cu128 required for sm\_120 \(Blackwell\) GPUs\.
#### Full results\.
Table 15:3D MiniWorld generalisation results: mean over 3 seeds, 3 M steps\. Baseline is disorder\-free; the primary assay for each disorder is in bold\.Cross\-assay dissociation is confirmed: drug\_occ and safe\_frac remain0\.000\.00across all depression doses; safe\_frac remains0\.000\.00and forward fraction is unaffected across all addiction doses; drug\_occ remains0\.000\.00across all anxiety doses\.
Figure 9:Grouped bar chart of all three assay metrics across all 10 conditions\. Each disorder region lights up only its own assay; off\-diagonal bars are zero, confirming block\-diagonal mechanism specificity in the 3D pixel domain\.
### Reproducibility
Grid\-world experiments use symbolic MiniGrid observations and run on CPU\. The full corpus is 1,375 runs \(main paper\) plus 36 MiniWorld runs \(3D generalisation\)\. Random seeds, dose grids, and the convergence criterion are fixed in configuration\. Every table and figure is regenerated from the per\-run result files by a single analysis script, and each figure derives from the same evaluation protocol \(40 episodes, stochastic policy, held\-out seeds\)\.Similar Articles
Modeling Pathology-Like Behavioral Patterns in Language Models Through Behavioral Fine-Tuning
This paper introduces a behavioral induction framework that fine-tunes language models on structured decision-making tasks to induce stable, context-general shifts in generative distributions, modeling pathology-like behavioral patterns such as depression and paranoia.
A Proactive Multi-Agent Dialogue Framework for Assessing Social Language Disorder Traits in Autism
This paper introduces TPA (Think, Plan, Ask), a proactive multi-agent dialogue framework using LLMs to systematically surface latent social language disorder traits in autism by selecting clinically grounded questioning strategies. It achieves 82.1% trait coverage, outperforming real clinical dialogues by clinicians.
TD-DPO: Difference-Aware Preference Optimization for Mitigating Sycophancy in Clinical Autism Intervention Dialogue
This paper proposes TD-DPO, a token-level difference-aware preference optimization method to mitigate sycophancy in LLMs for clinical autism intervention dialogue, achieving a better trade-off between sycophancy reduction and intervention ability retention.
Human-AI Agent Interaction as a Neuroplastic Training Environment
This paper proposes that the iterative loop of human-AI agent interaction (request, response, appraisal, revision) is a high-frequency neuroplastic training environment that can reinforce negative psychological patterns through repetition, and suggests it can be leveraged for beneficial cognitive training.
Learning Safe Agent Behaviour from Human Preferences and Justifications via World Models
This paper introduces DROPJ, a human-centred method for safely training and deploying agent policies by learning a world model from real-world trajectories, then eliciting human preferences with justifications to train a reward model for model predictive control. Experiments show that using human-generated simulated trajectories and justifications improves safety and reduces computational cost.