在部分可观测条件下进行模仿学习的最小递归行为记忆

arXiv cs.LG 论文

摘要

本文通过信息论方法和实验验证,刻画了在部分可观测条件下模仿专家所需的最小递归行为记忆。

arXiv:2609.25757v1 Announce Type: new Abstract: What is the least recurrent memory needed to reproduce a specified expert under partial observability? The instantaneous requirement is the conditional entropy of the expert's behavioral quotient, but recurrence must also preserve distinctions that future observations will not restore before use. We characterize this minimal recurrent behavioral memory by a compatibility relation: under transitivity its classes attain the exact minimum, while the general case is an entropy minimization over closed compatible state assignments, with exact certificates on finite instances. A sole-carrier measurement protocol separates behavioral sufficiency, excess code rate, and information carried by observations or other memory paths; experimental bit requirements refer to the induced symbolic behavioral model under the stated occupancy. Across manipulation tasks, learned code rates remain near zero- and two-bit requirements as hidden modes grow to $512$, and anticipatory memory follows a $2\to1\to0$ requirement despite zero instantaneous demand during waiting. Learning this representation remains difficult: event-agnostic future-behavior supervision yields $36/40$ sufficient seeds with one frozen configuration and improves the longest-horizon pixel setting from $0/8$ to $6/8$ sufficient held-out seeds (closed-loop success from $0.08$ to $0.57$). On unmodified community benchmarks, the protocol certifies delay-independent requirements, which sufficient codes match at mid-delay. The supervision aids commitment but can induce predictive surplus; annealing it lets imitation and rate training reduce that surplus, separating the information-theoretic target from the ability to learn it.
查看原文
查看缓存全文

缓存时间: 2026/09/23 09:36

# Minimal Recurrent Behavioral Memoryfor Imitation under Partial Observability
Source: [https://arxiv.org/html/2609.25757](https://arxiv.org/html/2609.25757)
###### Abstract

What is the least recurrent memory needed to reproduce a specified expert under partial observability? The instantaneous requirement is the conditional entropy of the expert’s behavioral quotient, but recurrence must also preserve distinctions that future observations will not restore before use\. We characterize this minimal recurrent behavioral memory by a compatibility relation: under transitivity its classes attain the exact minimum, while the general case is an entropy minimization over closed compatible state assignments, with exact certificates on finite instances\. A sole\-carrier measurement protocol separates behavioral sufficiency, excess code rate, and information carried by observations or other memory paths; experimental bit requirements refer to the induced symbolic behavioral model under the stated occupancy\. Across manipulation tasks, learned code rates remain near zero\- and two\-bit requirements as hidden modes grow to512512, and anticipatory memory follows a2→1→02\\to 1\\to 0requirement despite zero instantaneous demand during waiting\. Learning this representation remains difficult: event\-agnostic future\-behavior supervision yields36/4036/40sufficient seeds with one frozen configuration and improves the longest\-horizon pixel setting from0/80/8to6/86/8sufficient held\-out seeds \(closed\-loop success from0\.080\.08to0\.570\.57\)\. On unmodified community benchmarks, the protocol certifies delay\-independent requirements, which sufficient codes match at mid\-delay\. The supervision aids commitment but can induce predictive surplus; annealing it lets imitation and rate training reduce that surplus, separating the information\-theoretic target from the ability to learn it\.

## 1Introduction

Imitating a specified expert under partial observability raises three questions:*Which distinctions determine its behavior now? Which must persist because future observations will not restore them before use? Can a compact recurrent learner acquire that representation?*These concern current behavior, recurrent memory, and learning, respectively\.

Consider a robot that observes a grasp side and a placement slot, then waits\. Its waiting action depends on neither cue, but it must retain both for later decisions\. After grasping, only the slot remains necessary\. If the slot will be displayed again before placement, it need not be carried across the wait\.Future observations act as side information:the policy receives them as it acts, so memory need only bridge intervals without a new reveal\.

The behavioral quotientGE,tG\_\{E,t\}groups states with the same observation and current expert action distribution\. The recurrent targetΓt\\Gamma\_\{t\}additionally distinguishes histories that require different behavior after a common reachable continuation \(Figure[1](https://arxiv.org/html/2609.25757#S1.F1)\)\. It retains distinctions only while future observations cannot restore them before use\. Under transitive compatibility, we prove that its conditional entropy is the exact minimum among zero\-distortion recurrent realizations\. Without transitivity, closed compatible state assignments replace the class variable; small instances admit exact search, and matching bounds certify the minimum on our finite corridor instances\.

Existing representations answer different questions\. A recurrent architecture specifies where memory resides, while a control\-state target supports reward or dynamics prediction\([Subramanian et al\., 2022](https://arxiv.org/html/2609.25757#bib.bib24)\)\. Predictive states preserve future\-process distributions\([Shalizi and Crutchfield, 2001](https://arxiv.org/html/2609.25757#bib.bib4);[Littman et al\., 2001](https://arxiv.org/html/2609.25757#bib.bib5)\); system identification seeks hidden parameters\([Kumar et al\., 2021](https://arxiv.org/html/2609.25757#bib.bib15)\)\. These targets can distinguish histories that the specified expert treats identically\. Even future\-action prediction can retain a re\-displayed cue if its decoder receives no future observations\. Our target instead preserves the expert’s responses*as those observations arrive*\.

Figure 1:Current behavior, recurrent memory, and its learned realization\.In the transitive regime, history induces a current behavioral classGE,tG\_\{E,t\}and a recurrent targetΓt\\Gamma\_\{t\}that accounts for future observations as side information;CtC\_\{t\}is trained to realize this target\. Bottom: A′has three distinct rate levels,2→1→02\\to 1\\to 0\. The grasp and place columns show the rates at the decision,*before*the corresponding class is consumed\. The green note refers to the re\-reveal toy \(§[4\.3](https://arxiv.org/html/2609.25757#S4.SS3)\)\.The primary contribution is this representation target and its characterization\. Measuring it requires a sole temporal carrier and explicit observation and occupancy conventions\([Dann et al\., 2016](https://arxiv.org/html/2609.25757#bib.bib22), cf\.\)\. We check sufficiency before minimality and probe continuous bypasses, body memory, and current observations\. Exact experimental bits concern the induced symbolic model, not hardware storage or exact continuous control\.

Learning experiments test attainability: code rates remain near00and22bits as hidden\-mode counts grow to512512, while system\-identification rates grow with larger codebooks\. Plain training becomes less reliable with temporal distance or load\([Ni et al\., 2023](https://arxiv.org/html/2609.25757#bib.bib28), cf\.\)\. Event\-agnostic future\-behavior supervision improves acquisition from states and pixels in the tested settings, but can retain surplus information and still fails at longer delays\. The architecture uses standard components\([Tishby et al\., 2000](https://arxiv.org/html/2609.25757#bib.bib12);[Lee et al\., 2024](https://arxiv.org/html/2609.25757#bib.bib14)\); supervision is a learning surrogate, separate from the minimal\-memory characterization\.

## 2Minimal recurrent behavioral memory

### 2\.1The instantaneous problem

Consider a controlled process with stateStS\_\{t\}, observationOt=𝒪⁡\(St\)O\_\{t\}=\\mathcal\{O\}\(S\_\{t\}\), historyHt=\(O1:t,A1:t−1\)H\_\{t\}=\(O\_\{1:t\},A\_\{1:t\-1\}\), and expertπE\(⋅∣St\)\\pi\_\{E\}\(\\cdot\\mid S\_\{t\}\)\. All distributions below use the expert occupancydE,td\_\{E,t\}on a finite horizon\. We use countable reachable histories, as in the enumerated models; Appendix[A](https://arxiv.org/html/2609.25757#A1)states the conventions\.

###### Definition 1\(Behavioral quotient\)\.

Statess,s′s,s^\{\\prime\}are equivalent when𝒪⁡\(s\)=𝒪⁡\(s′\)\\mathcal\{O\}\(s\)=\\mathcal\{O\}\(s^\{\\prime\}\)andπE\(⋅∣s\)=πE\(⋅∣s′\)\\pi\_\{E\}\(\\cdot\\mid s\)=\\pi\_\{E\}\(\\cdot\\mid s^\{\\prime\}\)\. Their equivalence class isGE,t=gE​\(St\)G\_\{E,t\}=g\_\{E\}\(S\_\{t\}\)\.

We assume \(A1\) finitely many classes per observation fiber andH⁡\(GE,t∣Ot\)<∞H\(G\_\{E,t\}\\mid O\_\{t\}\)<\\infty; \(A2\)H⁡\(GE,t∣Ot,Ht\)=0H\(G\_\{E,t\}\\mid O\_\{t\},H\_\{t\}\)=0, so history determines the expert’s current behavior; and \(A3\) a nonnegative distortionδ\\deltathat vanishes exactly when action distributions agree\. Assumption \(A2\) excludes expert\-private information unavailable in history\([Yu, 2026](https://arxiv.org/html/2609.25757#bib.bib17)\)\.

LetRE​\(D\)R\_\{E\}\(D\)minimizeI⁡\(C;Ht∣Ot\)I\(C;H\_\{t\}\\mid O\_\{t\}\)over history encoders and decoders with expected distortion at mostDD\. Compressing history reduces to compressing the behavioral quotient:

RE​\(D\)=RG\|O​\(D\),RE​\(0\)=H⁡\(GE,t∣Ot\)\.R\_\{E\}\(D\)=R\_\{G\\mid O\}\(D\),\\qquad R\_\{E\}\(0\)=H\(G\_\{E,t\}\\mid O\_\{t\}\)\.\(1\)Every zero\-distortion code determinesGE,tG\_\{E,t\}together withOtO\_\{t\}, and encodingGE,tG\_\{E,t\}attains the bound \(Appendix[A\.1](https://arxiv.org/html/2609.25757#A1.SS1)\)\. IfMMuniform hidden modes formRRequal behavioral classes,log2⁡\(M/R\)\\log\_\{2\}\(M/R\)bits are unnecessary for the current action\. This elementary reduction answers the instantaneous question\. It does not ensure a representation that can be updated recursively\.

### 2\.2From current behavior to an exact recurrent target

Why recurrence changes the requirement\.In A′, an early cue reveals independent binary classes\(β1,β2\)\(\\beta\_\{1\},\\beta\_\{2\}\)for grasp and placement\. Immediately afterwardsH⁡\(GE,t∣Ot\)=0H\(G\_\{E,t\}\\mid O\_\{t\}\)=0, since every expert waits in the same way\. Nevertheless, a recurrent policy must retain all four combinations until the first decision:22bits\. After the grasp, the required memory falls to11bit; after placement it falls to00\. This is the distinction that the recurrent definition must capture\.

###### Definition 2\(Recurrent realization\)\.

A realization has a discrete stateCt=Ft​\(Ct−1,Ot,At−1\)C\_\{t\}=F\_\{t\}\(C\_\{t\-1\},O\_\{t\},A\_\{t\-1\}\)with deterministic updates and fixedC0C\_\{0\}, and a policyπ^\(⋅∣Ot,Ct\)\\hat\{\\pi\}\(\\cdot\\mid O\_\{t\},C\_\{t\}\)\. The code is its sole internal temporal carrier\. It has zero distortion ifπ^\(⋅∣Ot,Ct\)=PE\(At∣Ht\)\\hat\{\\pi\}\(\\cdot\\mid O\_\{t\},C\_\{t\}\)=P\_\{E\}\(A\_\{t\}\\mid H\_\{t\}\)almost surely at every step\.

DefineRtmem​\(0\)=infH⁡\(Ct∣Ot\)R\_\{t\}^\{\\rm mem\}\(0\)=\\inf H\(C\_\{t\}\\mid O\_\{t\}\)over zero\-distortion realizations on the*full horizon*, allowing countable state spaces\. The experiments use finite codebooks\. WriteUt​\(h\)=supp⁡PE​\(At,Ot\+1∣Ht=h\)U\_\{t\}\(h\)=\\operatorname\{supp\}P\_\{E\}\(A\_\{t\},O\_\{t\+1\}\\mid H\_\{t\}=h\)andh​uhufor a history extended byu=\(at,ot\+1\)u=\(a\_\{t\},o\_\{t\+1\}\)\.

###### Definition 3\(Behavioral memory compatibility\)\.

AtTT,h∼Th′h\\sim\_\{T\}h^\{\\prime\}iff\(OT,GE,T\)​\(h\)=\(OT,GE,T\)​\(h′\)\(O\_\{T\},G\_\{E,T\}\)\(h\)=\(O\_\{T\},G\_\{E,T\}\)\(h^\{\\prime\}\)\. Recursively,h∼th′h\\sim\_\{t\}h^\{\\prime\}iff\(Ot,GE,t\)​\(h\)=\(Ot,GE,t\)​\(h′\)\(O\_\{t\},G\_\{E,t\}\)\(h\)=\(O\_\{t\},G\_\{E,t\}\)\(h^\{\\prime\}\)and

hu∼t\+1h′ufor everyu∈Ut\(h\)∩Ut\(h′\)\.hu\\sim\_\{t\+1\}h^\{\\prime\}u\\quad\\text\{for every \}u\\in U\_\{t\}\(h\)\\cap U\_\{t\}\(h^\{\\prime\}\)\.

Compatible histories demand identical current behavior and remain compatible after every continuation reachable from both\. Comparing only common continuations credits future observations with the distinctions they will supply\.

###### Theorem 1\(Minimal recurrent behavioral memory\)\.

Under \(A1\)–\(A3\), if∼t\\sim\_\{t\}is transitive at every step, its classesΓt=\[Ht\]∼t\\Gamma\_\{t\}=\[H\_\{t\}\]\_\{\\sim\_\{t\}\}admit a deterministic updateΓt\+1=Φt​\(Γt,Ot\+1,At\)\\Gamma\_\{t\+1\}=\\Phi\_\{t\}\(\\Gamma\_\{t\},O\_\{t\+1\},A\_\{t\}\)and

Rtmem​\(0\)=H⁡\(Γt∣Ot\)≥H⁡\(GE,t∣Ot\)\.R\_\{t\}^\{\\rm mem\}\(0\)=H\(\\Gamma\_\{t\}\\mid O\_\{t\}\)\\geq H\(G\_\{E,t\}\\mid O\_\{t\}\)\.\(2\)Every zero\-distortion realization determinesΓt\\Gamma\_\{t\}from\(Ot,Ct\)\(O\_\{t\},C\_\{t\}\); the realizationCt=ΓtC\_\{t\}=\\Gamma\_\{t\}attains the minimum simultaneously at all steps\.

*Proof idea\.*Histories sharing\(Ot,Ct\)\(O\_\{t\},C\_\{t\}\)receive the same action distribution and, after a common continuation, the same updated code\. Backward induction makes each such cell compatible\. Under transitivity it lies in oneΓt\\Gamma\_\{t\}class, giving the lower bound\. Conversely, compatibility makes the successor class independent of the representative history, so the classes themselves define a sufficient realization\. Appendix[A](https://arxiv.org/html/2609.25757#A1)gives the proof\.

The*anticipatory memory*isΔt=H⁡\(Γt∣Ot\)−H⁡\(GE,t∣Ot\)≥0\\Delta\_\{t\}=H\(\\Gamma\_\{t\}\\mid O\_\{t\}\)\-H\(G\_\{E,t\}\\mid O\_\{t\}\)\\geq 0\. It quantifies information that must persist despite being unnecessary for the current action\. A sufficient finite codebook needsK≥maxt,o⁡\|Γt\|oK\\geq\\max\_\{t,o\}\|\\Gamma\_\{t\}\|\_\{o\}: in A′the joint memory has four states although each individual decision is binary\. A distinction can leave memory after its last use, or earlier if a future observation restores it before its next use\.

Transitivity is directly checkable\. A convenient sufficient condition, continuation\-support homogeneity \(A4\), requires equalUt​\(h\)U\_\{t\}\(h\)whenever\(Ot,GE,t\)​\(h\)\(O\_\{t\},G\_\{E,t\}\)\(h\)agrees\. It holds in A′and cue–corridor tasks; the solver certifies it there\. It is not necessary: Task A is transitive although a remembered mass changes the support of the next transport observation \(Appendix[A\.2](https://arxiv.org/html/2609.25757#A1.SS2)\)\.

### 2\.3Why the general case is harder

Future observations can distinguish histories that currently share a code\. If a class will be re\-observed before use, histories differing in that class can have disjoint continuation supports and be compatible without storing it\. But compatibility need not be transitive:h1∼h2h\_\{1\}\\sim h\_\{2\}andh2∼h3h\_\{2\}\\sim h\_\{3\}can hold on disjoint futures whileh1,h3h\_\{1\},h\_\{3\}conflict on a common future\. Taking a transitive closure would then merge histories that must remain distinct\.

The correct general object is a*closed compatible state assignment*: a partition of histories into pairwise\-compatible cells whose successors under each common continuation remain in one cell\. We prove that these assignments are exactly the joint\(Ot,Ct\)\(O\_\{t\},C\_\{t\}\)partitions of zero\-distortion recurrent realizations, with

Rtmem​\(0\)=infclosed compatible assignments​\(𝒫s\)H⁡\(Pt∣Ot\),R\_\{t\}^\{\\rm mem\}\(0\)=\\inf\_\{\\text\{closed compatible assignments \}\(\\mathcal\{P\}\_\{s\}\)\}H\(P\_\{t\}\\mid O\_\{t\}\),wherePtP\_\{t\}is the cell index \(Proposition[2](https://arxiv.org/html/2609.25757#Thmproposition2)\)\. This is a recurrent zero\-error problem with decoder side information\([Witsenhausen, 1976](https://arxiv.org/html/2609.25757#bib.bib9);[Alon and Orlitsky, 1996](https://arxiv.org/html/2609.25757#bib.bib10);[Koulgi et al\., 2003](https://arxiv.org/html/2609.25757#bib.bib49)\), related to incompletely specified machines\([Paull and Unger, 1959](https://arxiv.org/html/2609.25757#bib.bib7)\); the objective here is occupancy\-weighted conditional entropy\.

A stronger relation∼ts\\sim\_\{t\}^\{s\}also requires equal continuation supports and recursively applies that requirement\. It always defines a realizable equivalenceΓts\\Gamma\_\{t\}^\{s\}\.

###### Theorem 2\(General bounds\)\.

Under \(A1\)–\(A3\), every joint\(Ot,Ct\)\(O\_\{t\},C\_\{t\}\)cell of a zero\-distortion realization is compatible, and

H⁡\(GE,t∣Ot\)≤Rtmem​\(0\)≤H⁡\(Γts∣Ot\)\.H\(G\_\{E,t\}\\mid O\_\{t\}\)\\leq R\_\{t\}^\{\\rm mem\}\(0\)\\leq H\(\\Gamma\_\{t\}^\{s\}\\mid O\_\{t\}\)\.Both inequalities can be strict\.

Three equiprobable histories can have minimumh2​\(1/3\)h\_\{2\}\(1/3\)strictly inside\[0,log2⁡3\]\[0,\\log\_\{2\}3\]\. On Task A,Γts\\Gamma\_\{t\}^\{s\}charges a mass\-half bit because observation supports differ, whileΓt\\Gamma\_\{t\}requires zero: the distinction changes no expert action\.

Computation and certificates\.Backward refinement computes compatibility, checks transitivity, and returns exact class entropies when the condition holds\. A partition dynamic program solves small general instances\. For the larger non\-transitive corridor instancesW∈\{3,5,…,13\}W\\in\\\{3,5,\\dots,13\\\}, an incompatibility\-based entropy lower bound matches a certified closed compatible assignment at every step, establishing exact finite\-instance rates of22and11bits in the two halls \(Appendix[A\.3](https://arxiv.org/html/2609.25757#A1.SS3)\)\. This is a finite\-instance certificate, not a general polynomial\-time algorithm\. For nonzero distortion, the instantaneous rate–distortion function remains a lower bound; we do not characterize the full recurrent frontier \(Appendix[A\.4](https://arxiv.org/html/2609.25757#A1.SS4)\)\.

## 3When does a bit count as memory?

Our measurement principle separates three questions\.

Sufficiency:does\(Ct,O¯t\)\(C\_\{t\},\\bar\{O\}\_\{t\}\)determine the required behavioral memory?

Minimality:conditional on sufficiency, how much code rate exceeds the requirement?

Attribution:is the information inCtC\_\{t\}, rather than in another temporal carrier or in observations under a different occupancy?

sufficiency first, minimality second\\boxed\{\\text\{sufficiency first, minimality second\}\}

What the reported bits mean\.We distinguish the theoretical*behavioral memory rate*H⁡\(Γt∣O¯t\)H\(\\Gamma\_\{t\}\\mid\\bar\{O\}\_\{t\}\), the empirical*learned code rate*H^​\(Ct∣O¯t\)\\hat\{H\}\(C\_\{t\}\\mid\\bar\{O\}\_\{t\}\),*codebook capacity*log2⁡K\\log\_\{2\}K, and*storage footprint*in bytes\. The first is exact for the induced finite symbolic behavioral model under expert occupancy\. The observation conventionO¯t=ft​\(Ot\)\\bar\{O\}\_\{t\}=f\_\{t\}\(O\_\{t\}\)credits behaviorally decodable classes to the current observation\. Except for a disclosed, behaviorally irrelevant Task A sag symbol, it is a frozen coarsening, soH⁡\(Ct∣O¯t\)≥H⁡\(Ct∣Ot\)H\(C\_\{t\}\\mid\\bar\{O\}\_\{t\}\)\\geq H\(C\_\{t\}\\mid O\_\{t\}\)\(Appendix[B\.1](https://arxiv.org/html/2609.25757#A2.SS1)\)\. These rates quantify internal information burden under the convention, not exact continuous\-control requirements\.

Sufficiency and excess rate\.For transitive models, defineSΓ=1−H^​\(Γt∣Ct,O¯t\)/H^​\(Γt∣O¯t\)S\_\{\\Gamma\}=1\-\\hat\{H\}\(\\Gamma\_\{t\}\\mid C\_\{t\},\\bar\{O\}\_\{t\}\)/\\hat\{H\}\(\\Gamma\_\{t\}\\mid\\bar\{O\}\_\{t\}\)where the denominator is positive, andSGS\_\{G\}analogously\. A seed passes our empirical gate whenSΓ\>0\.9S\_\{\\Gamma\}\>0\.9at every positive\-requirement step of both waiting gaps andSG\>0\.9S\_\{G\}\>0\.9at the first grasp and first placement decisions; Task A uses its placement gate\. We call these seeds*sufficient*as shorthand for passing this diagnostic, not for exact zero distortion\. Non\-transitive corridor comparisons use the certified rate bound rather than assume a canonicalΓt\\Gamma\_\{t\}\. Under exact sufficiency,

H⁡\(Ct∣O¯t\)−H⁡\(Γt∣O¯t\)=H⁡\(Ct∣Γt,O¯t\)\.H\(C\_\{t\}\\mid\\bar\{O\}\_\{t\}\)\-H\(\\Gamma\_\{t\}\\mid\\bar\{O\}\_\{t\}\)=H\(C\_\{t\}\\mid\\Gamma\_\{t\},\\bar\{O\}\_\{t\}\)\.Thus matching the lower bound is evidence of absent redundancy only after sufficiency is established\. Threshold sensitivity and the full\-versus\-relaxed gate comparison are in Appendix[B\.5](https://arxiv.org/html/2609.25757#A2.SS5)\.

Attribution controls\.A quantized readout of a persistent GRU state reports only1\.061\.06bits on A′despite a two\-bit requirement: the unmeasured state carries history\. A policy can also encode a class in its own pose; body\-memory probes detect this\. We therefore report expert\-occupancy diagnostics separately from closed\-loop success under policy occupancy, with raw\-observation leak probes and matched\-side\-information controls \(Appendix[B\.4](https://arxiv.org/html/2609.25757#A2.SS4)\)\. Low code rate or high task success alone does not establish minimal internal memory\.

### 3\.1A measurable recurrent realization

We implement a discrete sole\-carrier recurrent policy \(our DIACRITIC realization\),Ct=Q⁡\(F⁡\(Ct−1,Ot,At−1\)\)C\_\{t\}=Q\(F\(C\_\{t\-1\},O\_\{t\},A\_\{t\-1\}\)\), using a residual proposal, hard nearest\-code quantization, and straight\-through gradients\. The action head receives\(Ot,Ct\)\(O\_\{t\},C\_\{t\}\); no continuous hidden state persists\. With a conditional priorrηr\_\{\\eta\}, the base objective is

ℒbase=1T​∑t\[ℓ⁡\(π^​\(AtE∣Ot,Ct\)\)\+β⁡\(−log⁡rη​\(Ct∣Ot\)\)\]\+ℒVQ\.\\mathcal\{L\}\_\{\\rm base\}=\\tfrac\{1\}\{T\}\\sum\_\{t\}\\\!\\left\[\\ell\\big\(\\hat\{\\pi\}\(A\_\{t\}^\{E\}\\mid O\_\{t\},C\_\{t\}\)\\big\)\+\\beta\\big\(\-\\log r\_\{\\eta\}\(C\_\{t\}\\mid O\_\{t\}\)\\big\)\\right\]\+\\mathcal\{L\}\_\{\\rm VQ\}\.\(3\)The expected rate penalty equalsH\(Ct∣Ot\)\+𝔼OtKL\(p\(⋅∣Ot\)∥rη\(⋅∣Ot\)\)H\(C\_\{t\}\\mid O\_\{t\}\)\+\\mathbb\{E\}\_\{O\_\{t\}\}\\mathrm\{KL\}\(p\(\\cdot\\mid O\_\{t\}\)\\\|r\_\{\\eta\}\(\\cdot\\mid O\_\{t\}\)\)\. It is a variational upper bound; reported rates are plug\-in entropies of hard codes\. These components are standard\([Tishby et al\., 2000](https://arxiv.org/html/2609.25757#bib.bib12);[Peng et al\., 2019](https://arxiv.org/html/2609.25757#bib.bib13);[Agustsson et al\., 2017](https://arxiv.org/html/2609.25757#bib.bib18);[Lee et al\., 2024](https://arxiv.org/html/2609.25757#bib.bib14)\)\. Future\-behavior supervision, defined in §[5](https://arxiv.org/html/2609.25757#S5), changes the training signal while leaving the evaluated carrier unchanged\. Appendix[B\.2](https://arxiv.org/html/2609.25757#A2.SS2)specifies the realization and auxiliary variants\.

## 4Does learned memory match the behavioral requirement?

### 4\.1Evaluation protocol

The main tasks are Task A \(physical hidden mass, zero required memory during grasp\), readout\-2 \(nonzero memory at growing hidden\-mode count\), and A′\(two delayed decisions\)\. The latter two use a probe channel and kinematic attachment; pixel A′retains the probe and phase inputs\. We certify the induced symbolic models and evaluate the same code diagnostics across tasks\. Unless stated otherwise, a cell has eight seeds and each policy has128128closed\-loop rollouts\. Rate minimality is assessed on sufficient seeds; counts and all\-seed control success are reported alongside it\.

Two protocols answer different questions\. For the plain learner’s*distortion\-constrained*comparison,β\\betais selected by closed\-loop fidelity before inspecting rates \(A′:β≤10−3\\beta\\leq 10^\{\-3\}\)\. An*attainability*control uses a continuous scaffold annealed to zero before evaluation \(N≈4N\\approx 4k,β=×10−3\\beta=2\\\!\\times\\\!10^\{\-3\}\); it demonstrates a reachable boundary, not a point on the plain learner’s frontier\. The supervised learning experiments use a separately frozen configuration\. Full settings, fresh\-recording checks, and estimator diagnostics are in Appendices[B\.1](https://arxiv.org/html/2609.25757#A2.SS1)–[B\.7](https://arxiv.org/html/2609.25757#A2.SS7)\.

Auditing observation side information\.Injecting a placement\-class leak into A′raises success from0\.590\.59to0\.960\.96and slot accuracy from0\.680\.68to0\.980\.98, while action error falls\. A pre\-specified low\-rate rule flags only1/961/96runs: codes can retain other content\. An exploratory rule pairing correct placement with insufficient second\-gap memory flags11/1511/15and8/148/14policies at22and44mm, versus0/1660/166correct clean unsupervised runs\. This auxiliary signal is nonmonotone \(4/164/16at88mm\)\. Direct observation probes test whetherOtO\_\{t\}reveals class information uncredited toO¯t\\bar\{O\}\_\{t\}\(Appendix[B\.6](https://arxiv.org/html/2609.25757#A2.SS6)\)\.

### 4\.2World complexity: zero and nonzero requirements

Figure 2:Learned code rate at fixed behavioral requirement\.\(a, b\) World complexity: Task A has a zero\-bit grasp requirement and readout\-2 a two\-bit first\-gap requirement asMMgrows\. Points are seed means with standard deviations over sufficient seeds \(counts shown\); system\-identification rates use all eight seeds and the largest tested codebook\. \(c\)*Mid\-delay rate*on the unmodified bsuite memory chain: every sufficient seed matches the requirement at this measurement point across the tested delays; labels count sufficient seeds of eight per learner, and hollow markers indicate none\. Dashed lines are certified requirements, not fits\.Task A\.An early probe identifies one ofM=4M=4–512512hidden masses\. The mass changes the arm’s dynamics and later observation supports but never the expert’s action choice; an independent slot is revealed shortly before placement\. Compatibility is transitive, so the exact grasp\-phase requirement is zero even though the strong congruence charges one bit\. The plainK=16K=16carrier is sufficient on8/88/8seeds at everyMM, carries0\.000\.00–0\.100\.10bit during grasp with at most0\.040\.04bit about mass, and attains0\.910\.91–0\.990\.99closed\-loop success\. AtM=512M=512its grasp rate is0\.030\.03bit\.

Readout\-2\.A binary display revealsθ∈\[M\]\\theta\\in\[M\]; the expert grasps according to its quartile and places according to its half\. The first\-gap requirement is exactly two bits atM=32,128,512M=32,128,512, while full\-history information during transport grows from66to1010bits\. Plain training is sufficient on8/8,7/8,8/88/8,7/8,8/8seeds; task\-informed future\-behavior supervision gives8/88/8throughout\. Learned first\-gap rates remain near22bits, with closed\-loop success0\.850\.85–0\.960\.96\. Verdicts reproduce on independent recordings\. The separation therefore also holds with a nonzero behavioral memory requirement\.

What the identification comparison establishes\.A system\-identification head forces the hidden parameter through the same regularized carrier\. With codebook capacity increased toK=128K=128–10241024, its learned rate grows from2\.252\.25to8\.718\.71bits on Task A and from4\.824\.82to7\.497\.49on readout\-2\. These are learned rates, not claims of perfect parameter recovery\. A separate identification carrier restores Task A success to0\.900\.90–0\.970\.97while carrying an additional4\.14\.1–6\.06\.0bits\. This control separates the information burden of identification from interference when both objectives share a carrier \(Appendix[C\.3](https://arxiv.org/html/2609.25757#A3.SS3)\)\.

External validation on community benchmarks\.On the bsuite memory chain and Passive T\-maze, used unmodified\([Cherepanov et al\., 2026](https://arxiv.org/html/2609.25757#bib.bib38);[Ni et al\., 2023](https://arxiv.org/html/2609.25757#bib.bib28)\), the solver certifies the induced symbolic models\. The memory\-chain requirement is zero during the two cue\-visible steps,bbbits from the first cue\-free step until the query, and one bit once the query index appears, independently of delay\. Of192192runs,8787pass the gate, including closed\-loop success≥0\.99\\geq 0\.99, and match the certified rate*at mid\-delay*\(Figure[2](https://arxiv.org/html/2609.25757#S4.F2)c\)\.

This agreement is constrained by the task structure: at mid\-delay the code is a deterministic function of the context, so exact sufficiency forces its rate to equal the context entropy\. These tasks validate protocol transfer and acquisition of the required information; Task A and readout test compression in the presence of additional world information\. Reliability decreases with delay: none of3232runs with two or three bits at delay100100passes \(Appendix[E\.5](https://arxiv.org/html/2609.25757#A5.SS5)\)\.

### 4\.3Anticipation, use, and re\-observation

Figure 3:A recurrent realization tracks the behavioral requirement\.A′at gap 20, using the matched\-side\-information hierarchical policy with task\-informed supervision \(β=10−3\\beta=10^\{\-3\}, eight seeds\)\. Thin blue curves show every seed’s learned code rate; black and orange curves show the certified recurrent and instantaneous requirements\. Both waiting gaps require memory despite zero instantaneous demand\. The theory and policy share the same symbolic side information\.A′isolates anticipation: its required memory follows2→1→02\\to 1\\to 0, whileH⁡\(GE,t∣O¯t\)=0H\(G\_\{E,t\}\\mid\\bar\{O\}\_\{t\}\)=0in both gaps\. Sufficient plain seeds carry2\.002\.00bits in the first gap\. The informative result is their negligible surplus, not the lower bound already implied by exact sufficiency\. A matched\-side\-information hierarchical control restricts the transition, prior, and behavioral head to\(Ct,O¯t\)\(C\_\{t\},\\bar\{O\}\_\{t\}\); a memory\-free controller executes the chosen action usingOtO\_\{t\}\. With task\-informed supervision, all3232runs across two horizons and two rate weights pass, carrying2\.032\.03–2\.092\.09bits before grasp and1\.001\.00afterwards, with closed\-loop success0\.960\.96–1\.001\.00\(Appendix[B\.4](https://arxiv.org/html/2609.25757#A2.SS4)\)\.

When the requirement falls\.Information expires after its last behavioral use or when future observations will restore it before use\. The re\-reveal toy exhibits the latter: its exact requirement is11bit before the first decision and00afterwards, while an open\-loop future\-action target retains2\.062\.06and1\.291\.29bits\. A fully physical weighing task also has a certified2→02\\to 0requirement because the next lift re\-reveals mass; its offline attainability and closed\-loop failure are documented in Appendix[E\.3](https://arxiv.org/html/2609.25757#A5.SS3)\.

Reaching the lower boundary\.Structural expiration does not guarantee learned minimality\. On A′, among seeds sufficient under policy occupancy, plain training atβ=10−3\\beta=10^\{\-3\}achieves0\.970\.97success with post\-use rate1\.381\.38; stronger pressure reduces it to1\.261\.26while success falls to0\.700\.70\. The scaffold control instead reaches0\.990\.99–1\.051\.05bit on five of six seeds sufficient under policy occupancy\. In the signpost corridor \(Appendix[C\.4](https://arxiv.org/html/2609.25757#A3.SS4)\), sufficient codes in theM=16M=16profile average1\.21\.2–1\.81\.8bits against the one\-bit requirement\. These observations establish attainability in specific settings and optimization slack elsewhere; the weight sweep is not an exact rate–distortion frontier \(Appendix[D\.6](https://arxiv.org/html/2609.25757#A4.SS6)\)\.

## 5Can a compact learner find the minimal memory?

We compare learned code rates with certified requirements, then ask how reliably training reaches a sufficient representation\. The theorem specifies the target, not a learning guarantee\.

### 5\.1A temporal learning difficulty

At a fixed certified requirement, code sufficiency is negatively associated with delay across320320sole\-carrier runs on A′\(−0\.146\-0\.146/step, CI\[−0\.18,−0\.12\]\[\-0\.18,\-0\.12\]\)\. This pooled association is descriptive because configurations differ\. Controlled1616\-seed sweeps also give negative within\-learner slopes; larger memory load reduces reliability \(Appendices[D\.7](https://arxiv.org/html/2609.25757#A4.SS7),[D\.8](https://arxiv.org/html/2609.25757#A4.SS8)\)\. Full\-history attention recovers the required decisions in each tested setting under teacher forcing and reaches0\.880\.88–1\.001\.00closed\-loop success, showing attainability with unrestricted history\. Its persistent history excludes its code readout from the memory\-rate comparison\.

### 5\.2Event\-agnostic future\-behavior supervision

To help a compact code acquire the required memory, we samplej∼U​\{1,…,T−1\}j\\sim U\\\{1,\\ldots,T\-1\\\}at each steptt\. A training\-only head predicts a causal forecaster’s future\-action estimatea~t,jE​\(Ht\)\\tilde\{a\}^\{E\}\_\{t,j\}\(H\_\{t\}\):

ℒfuture=𝔼t,j​\[𝟏t\+j≤T​‖qψ​\(Ot,Ct,j\)−a~t,jE​\(Ht\)‖2\]\.\\mathcal\{L\}\_\{\\rm future\}=\\mathbb\{E\}\_\{t,j\}\\\!\\left\[\\mathbf\{1\}\_\{t\+j\\leq T\}\\left\\\|q\_\{\\psi\}\(O\_\{t\},C\_\{t\},j\)\-\\tilde\{a\}^\{E\}\_\{t,j\}\(H\_\{t\}\)\\right\\\|^\{2\}\\right\]\.\(4\)The forecaster learns from recorded histories and provides conditional predictions even before cue visibility\. No latent labels or event annotations enter the objective; only offsets beyond the episode are masked\. Both forecaster and auxiliary head are discarded at evaluation\. We optimizeℒbase\+λ⁡\(k\)​ℒfuture\\mathcal\{L\}\_\{\\rm base\}\+\\lambda\(k\)\\mathcal\{L\}\_\{\\rm future\}:λ=1\\lambda=1through60%60\\%of training, falls to zero by80%80\\%, and leaves imitation and rate alone for the remaining steps\.

Table 1:Learning near the certified two\-bit requirement\.First\-gap rates are reported below; counts give full\-gate sufficient seeds out of eight, and the last row gives all\-seed closed\-loop success\. The event\-agnostic configuration is frozen; all4040verdicts reproduce on independent recordings\. The task\-informed reference uses event structure\.Event\-agnostic supervision yields first\-gap rates of2\.002\.00–2\.042\.04bits on sufficient seeds, against a two\-bit requirement\. The frozen configuration \(K=16K=16,β=10−3\\beta=10^\{\-3\},N=448N=448,10410^\{4\}steps\) yields36/4036/40such seeds \(Table[1](https://arxiv.org/html/2609.25757#S5.T1)\); every verdict reproduces on fresh recordings\. A more informed reference predicts only the first grasp and placement actions and masks them after use\. It reaches37/4037/40, with lower post\-use rates, but requires the event structure\. These are two surrogates with different information requirements, not different definitions ofΓt\\Gamma\_\{t\}\.

What the controls suggest\.Matching a full\-history teacher’s internal state yields only00–2/82/8sufficient seeds: a state used for the current action need not expose a pending class that the teacher can retrieve later\. Longer training, curriculum, and code refinement leave the longest A′horizon at00–1/81/8\. Improvement from future\-behavior supervision at the same capacity therefore points to the long\-range training signal as an important bottleneck, without identifying a unique optimization mechanism\. Direct recorded\-action targets work for the task\-informed reference \(7/87/8\); they do not establish teacher\-free performance for the frozen event\-agnostic configuration \(Appendix[D\.1](https://arxiv.org/html/2609.25757#A4.SS1)\)\.

State count and learning reliability\.Readout\-3 requires three bits and at least eight states, yet sufficient\-seed counts are0,1,5,15,230,1,5,15,23out of2424forK=8,10,12,16,24K=8,10,12,16,24, respectively; sufficient seeds average2\.992\.99–3\.143\.14bits \(Appendix[D\.5](https://arxiv.org/html/2609.25757#A4.SS5)\)\. This separates the required state count from the capacity that supports reliable learning under the tested configuration\. Small continuous RNNs also fail on the external T\-maze, although these controls do not isolate quantization from network size and optimization \(Appendix[E\.4](https://arxiv.org/html/2609.25757#A5.SS4)\)\.

### 5\.3Pixels and predictive surplus

Table 2:Pixel code rates versus the certified requirement\.A′at gap 20 requires2→12\\to 1bits\. Sixteen seeds per row; rates average sufficient seeds, success all seeds\. Inputs retain probe and phase channels\. The60%60\\%schedule was selected on seeds 0–7 and tested unchanged on seeds 8–15\.With pixels, sufficient codes approach the certified rates \(Table[2](https://arxiv.org/html/2609.25757#S5.T2)\)\. On held\-out seeds, event\-agnostic supervision yields6/86/8sufficient policies and0\.570\.57closed\-loop success, against0/80/8and0\.080\.08for plain training; a schedule fixed before the pixel runs gives12/1612/16overall \(Appendix[D\.1](https://arxiv.org/html/2609.25757#A4.SS1)\)\. Grasp\-side accuracy is high for all learners; the difference concerns the later placement\. Pixel control precision still limits success\.

Acquisition and minimality respond differently to supervision\.Without annealing, event\-agnostic supervision still acquires the memory at gap 20 \(7/87/8\), but the second\-gap code rate remains2\.142\.14bits from states and2\.012\.01from pixels against a one\-bit requirement\. The expired grasp class contributes at most0\.020\.02bit with or without annealing\.

Conditioned on the pending placement class, the state\-input code carries0\.560\.56bit about the initialxx\-position quartile, compared with1\.161\.16bits of total surplus\. Removing positional jitter and retraining yields8/88/8sufficient seeds and0\.940\.94closed\-loop success, while reducing un\-annealed surplus from1\.161\.16to0\.230\.23bit\. This supports initial\-position variability as an important driver\. With the original jitter, annealing reduces information aboutxxto0\.010\.01–0\.030\.03bit and surplus to0\.170\.17–0\.210\.21bit \(Appendix[D\.1](https://arxiv.org/html/2609.25757#A4.SS1)\)\.

The external memory chain also separates acquisition from forgetting\. After the query appears, the requirement falls to about one bit, but sufficient runs with two\- and three\-bit contexts retain mean rates of1\.911\.91and2\.692\.69bits, despite their exact mid\-delay rates \(Appendix[E\.5](https://arxiv.org/html/2609.25757#A5.SS5)\)\. The query is shown only at the final step, so the rate term can reward forgetting at that step alone\.

An open\-loop action target can reward continuous\-control details outside the symbolic target, as well as distinctions that future observations will supply\. The tested A′controls separate two effects: future supervision helps acquire required memory, while removing it lets the rate objective reduce predictive surplus\.

## 6Related representation targets

Predictive states\.Causal states and predictive\-state representations preserve future observation distributions\([Shalizi and Crutchfield, 2001](https://arxiv.org/html/2609.25757#bib.bib4);[Still et al\., 2010](https://arxiv.org/html/2609.25757#bib.bib25);[Littman et al\., 2001](https://arxiv.org/html/2609.25757#bib.bib5);[Gangwani et al\., 2020](https://arxiv.org/html/2609.25757#bib.bib50);[Ni et al\., 2024](https://arxiv.org/html/2609.25757#bib.bib30)\); theϵ\\epsilon\-transducer preserves conditional input–output behavior\([Barnett and Crutchfield, 2015](https://arxiv.org/html/2609.25757#bib.bib51)\)\. Our fixed\-expert target compares responses on common reachable continuations, treating future observations as side information\. Predicting the full process can require additional distinctions \(Appendix[F](https://arxiv.org/html/2609.25757#A6)\)\.

States for control or identification\.Bisimulation, approximate information states, and stable quotients preserve reward, dynamics, or Markov structure\([Givan et al\., 2003](https://arxiv.org/html/2609.25757#bib.bib31);[Ferns et al\., 2004](https://arxiv.org/html/2609.25757#bib.bib32);[Castro, 2020](https://arxiv.org/html/2609.25757#bib.bib2);[Subramanian et al\., 2022](https://arxiv.org/html/2609.25757#bib.bib24);[Zhang et al\., 2026](https://arxiv.org/html/2609.25757#bib.bib6)\)\. Inverse representations target control\-endogenous distinctions\([Mhammedi et al\., 2023](https://arxiv.org/html/2609.25757#bib.bib21);[Lamb et al\., 2023](https://arxiv.org/html/2609.25757#bib.bib16);[Wu et al\., 2024](https://arxiv.org/html/2609.25757#bib.bib44)\); identification targets hidden parameters\([Kumar et al\., 2021](https://arxiv.org/html/2609.25757#bib.bib15);[Liang et al\., 2024](https://arxiv.org/html/2609.25757#bib.bib47)\)\. A specified expert can ignore these\. Our matched objectives compare targets, rather than full published algorithms \(Appendix[C\.3](https://arxiv.org/html/2609.25757#A3.SS3)\)\.

Information\-limited memory\.Bounded\-memory control uses sequential rate–distortion objectives\([Fox and Tishby, 2012](https://arxiv.org/html/2609.25757#bib.bib53);[Fox and Tishby, 2016](https://arxiv.org/html/2609.25757#bib.bib54)\); decision\-centric memory retains distinctions supporting good decisions\([Zou et al\., 2026](https://arxiv.org/html/2609.25757#bib.bib34);[Walsh, 2026](https://arxiv.org/html/2609.25757#bib.bib43);[Yamin et al\., 2026](https://arxiv.org/html/2609.25757#bib.bib27)\)\. Our fixed\-expert characterization combines recurrent realizability, future observations as side information, and entropy minimality\. Memory Lens bounds action\-relevant information\([Dann et al\., 2016](https://arxiv.org/html/2609.25757#bib.bib22)\); our target includes the anticipatory gap\. Complementary work studies spurious history dependence\([de Haan et al\., 2019](https://arxiv.org/html/2609.25757#bib.bib42);[Wen et al\., 2020](https://arxiv.org/html/2609.25757#bib.bib37);[Swamy et al\., 2022](https://arxiv.org/html/2609.25757#bib.bib29)\), memory architectures and benchmarks\([Yue et al\., 2025](https://arxiv.org/html/2609.25757#bib.bib26);[Wang et al\., 2026](https://arxiv.org/html/2609.25757#bib.bib20);[Shah et al\., 2026](https://arxiv.org/html/2609.25757#bib.bib19);[Villasevil et al\., 2025](https://arxiv.org/html/2609.25757#bib.bib36);[Gao et al\., 2026](https://arxiv.org/html/2609.25757#bib.bib46);[Morad et al\., 2023](https://arxiv.org/html/2609.25757#bib.bib39);[Pleines et al\., 2025](https://arxiv.org/html/2609.25757#bib.bib40);[Chen et al\., 2026](https://arxiv.org/html/2609.25757#bib.bib45);[Dai et al\., 2026](https://arxiv.org/html/2609.25757#bib.bib41)\), and the separation of memory from credit\-assignment length\([Ni et al\., 2023](https://arxiv.org/html/2609.25757#bib.bib28)\)or distillation\([Parisotto and Salakhutdinov, 2021](https://arxiv.org/html/2609.25757#bib.bib52);[Weinzaepfel et al\., 2026](https://arxiv.org/html/2609.25757#bib.bib33)\)\. These motivate our learning controls\.

## 7Scope and conclusion

The exact quotient formula requires history\-determined expert behavior and transitive compatibility; general exact computation remains limited to small instances or matching finite\-instance bounds\. Experimental rates concern symbolic behavior under the stated observation convention and occupancy\. Reliable closed\-loop evidence is strongest through two bits\. On physical weighing, expert\-led observation generation restores grasp\-side accuracy from0\.300\.30–0\.320\.32to0\.930\.93–0\.990\.99, supporting retention and use of this distinction while subsequent control remains imprecise \(Appendix[E\.3](https://arxiv.org/html/2609.25757#A5.SS3)\)\.

Event\-agnostic prediction need not preserve a cue whose future action depends on observations yet to arrive\. It extends the compact carrier’s learning horizon on the external Passive T\-maze but still fails at length100100; a128128\-dimensional GRU is more reliable \(Appendix[E\.4](https://arxiv.org/html/2609.25757#A5.SS4)\)\. Compactness serves bounded, auditable memory here, without an established architectural advantage\.

The central result is the minimal recurrent behavioral memory of a specified expert\. Future observations determine which historical distinctions must persist; the measurement protocol makes this requirement testable; and the learning experiments show both realizations near the boundary and a gap between information sufficiency and its acquisition\.

### Reproducibility statement

The theoretical quantities of the enumerated benchmark models are certified by refinement or matching bounds, with closed\-form validation checks reported in Appendix[A\.6](https://arxiv.org/html/2609.25757#A1.SS6); Appendix[A](https://arxiv.org/html/2609.25757#A1)states the assumptions and complete proofs\. Appendix[B\.1](https://arxiv.org/html/2609.25757#A2.SS1)specifies the benchmarks, the binned\-observation convention, the leak and body\-memory probes, and the acceptance checks for each dataset\. Section[4\.1](https://arxiv.org/html/2609.25757#S4.SS1)gives the frozen configuration for each family and the two evaluation protocols\. Section[5](https://arxiv.org/html/2609.25757#S5)and Appendix[D\.7](https://arxiv.org/html/2609.25757#A4.SS7)specify the row set and statistical inference used in the pooled regression, and Appendix[B\.10](https://arxiv.org/html/2609.25757#A2.SS10)reports compute\. Code for the environments, training, evaluation, probes and analysis, per\-run result files for the learning comparisons, and the solver as a standalone package are available at[https://github\.com/XianyaoLi/DIACRITIC](https://github.com/XianyaoLi/DIACRITIC)\. Given an adapter that resets an environment for each hidden value, supplies an oracle expert and maps raw observations to symbolsO¯t\\bar\{O\}\_\{t\}, the solver enumerates the hidden values, builds the induced finite symbolic model, checks \(A2\), \(A4\) and transitivity at every step, and returnsH⁡\(GE,t∣O¯t\)H\(G\_\{E,t\}\\mid\\bar\{O\}\_\{t\}\),H⁡\(Γt∣O¯t\)H\(\\Gamma\_\{t\}\\mid\\bar\{O\}\_\{t\}\)andH⁡\(Γts∣O¯t\)H\(\\Gamma^\{s\}\_\{t\}\\mid\\bar\{O\}\_\{t\}\); when transitivity fails it reports only the bracket\[H⁡\(GE,t∣O¯t\),H⁡\(Γts∣O¯t\)\]\[H\(G\_\{E,t\}\\mid\\bar\{O\}\_\{t\}\),\\,H\(\\Gamma^\{s\}\_\{t\}\\mid\\bar\{O\}\_\{t\}\)\]of Theorem[2](https://arxiv.org/html/2609.25757#Thmtheorem2), not a certified minimum\. On a CPU, a single command \(python \-m certify \-\-paper\) reproduces the1313memory\-chain and Passive T\-maze certifications of §[4\.2](https://arxiv.org/html/2609.25757#S4.SS2)against unmodified MIKASA\-Base code \(commitac81b6f\), together with the non\-transitive instance of Appendix[A\.3](https://arxiv.org/html/2609.25757#A1.SS3)and the \(A4\)\-failing re\-reveal instance of Appendix[A\.6](https://arxiv.org/html/2609.25757#A1.SS6)\. The corridor certificates of Appendix[A\.3](https://arxiv.org/html/2609.25757#A1.SS3)are reproduced by two separate scripts, and the pooled regression of §[5](https://arxiv.org/html/2609.25757#S5)by one\. The solver requires enumerable hidden variables and a deterministic map from hidden values to symbolic observation sequences\. Recorded demonstrations, trained models, and hard\-code arrays are omitted because of their size; the recording and training scripts regenerate them\. Estimator and code\-content diagnostics require these artifacts in addition to the bundled ledgers\.

## References

- Agarwalet al\.\(2026\)A\. Agarwal, A\. Wei, T\. Kargin, M\. Zeng, C\. Becker, A\. K\. Dayi, P\. Parrilo, A\. Ozdaglar, and R\. TedrakeTraining and evaluating diffusion policies with long context lengths\.External Links:2606\.16447,[Link](https://arxiv.org/abs/2606.16447)Cited by:[§B\.1](https://arxiv.org/html/2609.25757#A2.SS1.p3.1)\.
- Agustssonet al\.\(2017\)E\. Agustsson, F\. Mentzer, M\. Tschannen, L\. Cavigelli, R\. Timofte, L\. Benini, and L\. V\. GoolSoft\-to\-hard vector quantization for end\-to\-end learning compressible representations\.InAdvances in Neural Information Processing Systems,I\. Guyon, U\. V\. Luxburg, S\. Bengio, H\. Wallach, R\. Fergus, S\. Vishwanathan, and R\. Garnett \(Eds\.\),Vol\.30,pp\.\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2017/file/86b122d4358357d834a87ce618a55de0-Paper.pdf)Cited by:[§3\.1](https://arxiv.org/html/2609.25757#S3.SS1.p1.2)\.
- Alon and Orlitsky \(1996\)N\. Alon and A\. OrlitskySource coding and graph entropies\.IEEE Transactions on Information Theory42\(5\),pp\. 1329–1339\.Cited by:[§A\.3](https://arxiv.org/html/2609.25757#A1.SS3.SSS0.Px1.p1.1),[§2\.3](https://arxiv.org/html/2609.25757#S2.SS3.p2.2)\.
- Barnett and Crutchfield \(2015\)N\. Barnett and J\. P\. CrutchfieldComputational mechanics of input–output processes: structured transformations and theϵ\\epsilon\-transducer\.Journal of Statistical Physics161\(2\),pp\. 404–451\.Cited by:[§6](https://arxiv.org/html/2609.25757#S6.p1.1)\.
- Castro \(2020\)P\. S\. CastroScalable methods for computing state similarity in deterministic Markov decision processes\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.34,pp\. 10069–10076\.Cited by:[Table 23](https://arxiv.org/html/2609.25757#A6.T23.10.2.1.1.1),[§6](https://arxiv.org/html/2609.25757#S6.p2.1)\.
- Chenet al\.\(2026\)T\. Chen, Y\. Wang, M\. Li, Y\. Qin, H\. Shi, Z\. Li, Y\. Hu, Y\. Zhang, K\. Wang, Y\. Chen, H\. Wang, J\. Wang, T\. Yang, R\. Xu, R\. Wu, Y\. Mu, Y\. Yang, H\. Dong, and P\. LuoRMBench: memory\-dependent robotic manipulation benchmark with insights into policy design\.External Links:2603\.01229,[Link](https://arxiv.org/abs/2603.01229)Cited by:[§6](https://arxiv.org/html/2609.25757#S6.p3.1)\.
- Cherepanovet al\.\(2026\)E\. Cherepanov, N\. Kachaev, A\. Kovalev, and A\. PanovMemory, benchmark & robots: a benchmark for solving complex tasks with reinforcement learning\.InThe Fourteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=9cLPurIZMj)Cited by:[§E\.4](https://arxiv.org/html/2609.25757#A5.SS4.p1.1),[§E\.5](https://arxiv.org/html/2609.25757#A5.SS5.p1.1),[§4\.2](https://arxiv.org/html/2609.25757#S4.SS2.p4.1)\.
- Daiet al\.\(2026\)Y\. Dai, H\. Fu, J\. Lee, Y\. Liu, H\. Zhang, J\. Yang, C\. Finn, N\. Fazeli, and J\. ChaiRoboMME: benchmarking and understanding memory for robotic generalist policies\.InForty\-third International Conference on Machine Learning,External Links:[Link](https://openreview.net/forum?id=8m30ogkPk2)Cited by:[§6](https://arxiv.org/html/2609.25757#S6.p3.1)\.
- Dannet al\.\(2016\)C\. Dann, K\. Hofmann, and S\. NowozinMemory lens: how much memory does an agent use?\.arXiv preprint arXiv:1611\.06928\.Cited by:[§1](https://arxiv.org/html/2609.25757#S1.p5.1),[§6](https://arxiv.org/html/2609.25757#S6.p3.1)\.
- de Haanet al\.\(2019\)P\. de Haan, D\. Jayaraman, and S\. LevineCausal confusion in imitation learning\.InAdvances in Neural Information Processing Systems,H\. Wallach, H\. Larochelle, A\. Beygelzimer, F\. d'Alché\-Buc, E\. Fox, and R\. Garnett \(Eds\.\),Vol\.32,pp\.\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2019/file/947018640bf36a2bb609d3557a285329-Paper.pdf)Cited by:[§6](https://arxiv.org/html/2609.25757#S6.p3.1)\.
- Fernset al\.\(2004\)N\. Ferns, P\. Panangaden, and D\. PrecupMetrics for finite Markov decision processes\.InProceedings of the 20th Conference on Uncertainty in Artificial Intelligence,UAI ’04,Arlington, Virginia, USA,pp\. 162–169\.External Links:ISBN 0974903906Cited by:[§6](https://arxiv.org/html/2609.25757#S6.p2.1)\.
- Fox and Tishby \(2012\)R\. Fox and N\. TishbyBounded planning in passive POMDPs\.InProceedings of the 29th International Conference on Machine Learning,Cited by:[§6](https://arxiv.org/html/2609.25757#S6.p3.1)\.
- Fox and Tishby \(2016\)R\. Fox and N\. TishbyMinimum\-information LQG control part II: retentive controllers\.In2016 IEEE 55th Conference on Decision and Control \(CDC\),pp\. 5603–5609\.Cited by:[§6](https://arxiv.org/html/2609.25757#S6.p3.1)\.
- Gangwaniet al\.\(2020\)T\. Gangwani, J\. Lehman, Q\. Liu, and J\. PengLearning belief representations for imitation learning in pomdps\.InProceedings of The 35th Uncertainty in Artificial Intelligence Conference,R\. P\. Adams and V\. Gogate \(Eds\.\),Proceedings of Machine Learning Research, Vol\.115,pp\. 1061–1071\.External Links:[Link](https://proceedings.mlr.press/v115/gangwani20a.html)Cited by:[§6](https://arxiv.org/html/2609.25757#S6.p1.1)\.
- Gaoet al\.\(2026\)Y\. Gao, J\. J\. Liu, S\. Li, and S\. SongGated memory policy: in\-context memorization and adaptation\.External Links:2604\.18933,[Link](https://arxiv.org/abs/2604.18933)Cited by:[§6](https://arxiv.org/html/2609.25757#S6.p3.1)\.
- Givanet al\.\(2003\)R\. Givan, T\. Dean, and M\. GreigEquivalence notions and model minimization in Markov decision processes\.Artificial intelligence147\(1\-2\),pp\. 163–223\.Cited by:[§6](https://arxiv.org/html/2609.25757#S6.p2.1)\.
- Gray \(1973\)R\. M\. GrayA new class of lower bounds to information rates of stationary sources via conditional rate\-distortion functions\.IEEE Transactions on Information Theory19\(4\),pp\. 480–489\.Cited by:[§A\.1](https://arxiv.org/html/2609.25757#A1.SS1.p2.1)\.
- Koulgiet al\.\(2003\)P\. Koulgi, E\. Tuncel, S\.L\. Regunathan, and K\. RoseOn zero\-error source coding with decoder side information\.IEEE Transactions on Information Theory49\(1\),pp\. 99–111\.External Links:[Document](https://dx.doi.org/10.1109/TIT.2002.806154)Cited by:[§2\.3](https://arxiv.org/html/2609.25757#S2.SS3.p2.2)\.
- Kumaret al\.\(2021\)A\. Kumar, Z\. Fu, D\. Pathak, and J\. MalikRMA: rapid motor adaptation for legged robots\.InProceedings of Robotics: Science and Systems,Virtual\.External Links:[Document](https://dx.doi.org/10.15607/RSS.2021.XVII.011)Cited by:[Table 23](https://arxiv.org/html/2609.25757#A6.T23.10.10.1.1.1),[§1](https://arxiv.org/html/2609.25757#S1.p4.1),[§6](https://arxiv.org/html/2609.25757#S6.p2.1)\.
- Lambet al\.\(2023\)A\. Lamb, R\. Islam, Y\. Efroni, A\. R\. Didolkar, D\. Misra, D\. J\. Foster, L\. P\. Molu, R\. Chari, A\. Krishnamurthy, and J\. LangfordGuaranteed discovery of control\-endogenous latent states with multi\-step inverse models\.Transactions on Machine Learning Research\.Note:External Links:ISSN 2835\-8856,[Link](https://openreview.net/forum?id=TNocbXm5MZ)Cited by:[§C\.3](https://arxiv.org/html/2609.25757#A3.SS3.p2.1),[Appendix F](https://arxiv.org/html/2609.25757#A6.SS0.SSS0.Px1.p1.1),[Table 23](https://arxiv.org/html/2609.25757#A6.T23.10.5.1.1.1),[§6](https://arxiv.org/html/2609.25757#S6.p2.1)\.
- Leeet al\.\(2024\)S\. Lee, Y\. Wang, H\. Etukuru, H\. J\. Kim, N\. M\. M\. Shafiullah, and L\. PintoBehavior generation with latent actions\.InForty\-first International Conference on Machine Learning,External Links:[Link](https://openreview.net/forum?id=hoVwecMqV5)Cited by:[§1](https://arxiv.org/html/2609.25757#S1.p6.1),[§3\.1](https://arxiv.org/html/2609.25757#S3.SS1.p1.2)\.
- Liet al\.\(2006\)L\. Li, T\. J\. Walsh, and M\. L\. LittmanTowards a unified theory of state abstraction for MDPs\.InInternational Symposium on Artificial Intelligence and Mathematics \(ISAIM\),Cited by:[Table 23](https://arxiv.org/html/2609.25757#A6.T23.10.3.1.1.1)\.
- Lianget al\.\(2024\)Y\. Liang, K\. Ellis, and J\. HenriquesRapid motor adaptation for robotic manipulator arms\.In2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),Vol\.,pp\. 16404–16413\.External Links:[Document](https://dx.doi.org/10.1109/CVPR52733.2024.01552)Cited by:[Table 23](https://arxiv.org/html/2609.25757#A6.T23.10.10.1.1.1),[§6](https://arxiv.org/html/2609.25757#S6.p2.1)\.
- Littmanet al\.\(2001\)M\. L\. Littman, R\. S\. Sutton, and S\. SinghPredictive representations of state\.InAdvances in Neural Information Processing Systems,Vol\.14\.Cited by:[Table 23](https://arxiv.org/html/2609.25757#A6.T23.10.4.1.1.1),[§1](https://arxiv.org/html/2609.25757#S1.p4.1),[§6](https://arxiv.org/html/2609.25757#S6.p1.1)\.
- Mhammediet al\.\(2023\)Z\. Mhammedi, D\. J\. Foster, and A\. RakhlinRepresentation learning with multi\-step inverse kinematics: an efficient and optimal approach to rich\-observation RL\.InProceedings of the 40th International Conference on Machine Learning,A\. Krause, E\. Brunskill, K\. Cho, B\. Engelhardt, S\. Sabato, and J\. Scarlett \(Eds\.\),Proceedings of Machine Learning Research, Vol\.202,pp\. 24659–24700\.External Links:[Link](https://proceedings.mlr.press/v202/mhammedi23a.html)Cited by:[§C\.3](https://arxiv.org/html/2609.25757#A3.SS3.p2.1),[Appendix F](https://arxiv.org/html/2609.25757#A6.SS0.SSS0.Px1.p1.1),[Table 23](https://arxiv.org/html/2609.25757#A6.T23.10.5.1.1.1),[§6](https://arxiv.org/html/2609.25757#S6.p2.1)\.
- Moradet al\.\(2023\)S\. Morad, R\. Kortvelesy, M\. Bettini, S\. Liwicki, and A\. ProrokPOPGym: benchmarking partially observable reinforcement learning\.InThe Eleventh International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=chDrutUTs0K)Cited by:[§6](https://arxiv.org/html/2609.25757#S6.p3.1)\.
- Niet al\.\(2024\)T\. Ni, B\. Eysenbach, E\. Seyedsalehi, M\. Ma, C\. Gehring, A\. Mahajan, and P\. BaconBridging state and history representations: understanding self\-predictive rl\.InInternational Conference on Learning Representations,Cited by:[§6](https://arxiv.org/html/2609.25757#S6.p1.1)\.
- Niet al\.\(2023\)T\. Ni, M\. Ma, B\. Eysenbach, and P\. BaconWhen do transformers shine in rl? decoupling memory from credit assignment\.Advances in Neural Information Processing Systems36,pp\. 50429–50452\.Cited by:[§E\.2](https://arxiv.org/html/2609.25757#A5.SS2.p1.1),[§E\.4](https://arxiv.org/html/2609.25757#A5.SS4.p1.1),[§1](https://arxiv.org/html/2609.25757#S1.p6.1),[§4\.2](https://arxiv.org/html/2609.25757#S4.SS2.p4.1),[§6](https://arxiv.org/html/2609.25757#S6.p3.1)\.
- Nixon \(2026\)A\. T\. NixonThe myhill\-nerode theorem for bounded interaction: canonical abstractions via agent\-bounded indistinguishability\.arXiv preprint arXiv:2603\.21399\.Cited by:[Table 23](https://arxiv.org/html/2609.25757#A6.T23.10.7.1.1.1)\.
- Parisotto and Salakhutdinov \(2021\)E\. Parisotto and R\. SalakhutdinovEfficient transformers in reinforcement learning using actor\-learner distillation\.InInternational Conference on Learning Representations,Cited by:[§6](https://arxiv.org/html/2609.25757#S6.p3.1)\.
- Paull and Unger \(1959\)M\. C\. Paull and S\. H\. UngerMinimizing the number of states in incompletely specified sequential switching functions\.IRE Transactions on Electronic ComputersEC\-8\(3\),pp\. 356–367\.Cited by:[§A\.3](https://arxiv.org/html/2609.25757#A1.SS3.p1.1),[Table 23](https://arxiv.org/html/2609.25757#A6.T23.10.8.1.1.1),[§2\.3](https://arxiv.org/html/2609.25757#S2.SS3.p2.2)\.
- Penget al\.\(2019\)X\. B\. Peng, A\. Kanazawa, S\. Toyer, P\. Abbeel, and S\. LevineVariational discriminator bottleneck: improving imitation learning, inverse RL, and GANs by constraining information flow\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=HyxPx3R9tm)Cited by:[§3\.1](https://arxiv.org/html/2609.25757#S3.SS1.p1.2)\.
- Pfleeger \(1973\)C\. P\. PfleegerState reduction in incompletely specified finite\-state machines\.IEEE Transactions on ComputersC\-22\(12\),pp\. 1099–1102\.Cited by:[§A\.3](https://arxiv.org/html/2609.25757#A1.SS3.SSS0.Px3.p1.1),[Table 23](https://arxiv.org/html/2609.25757#A6.T23.10.8.1.1.1)\.
- Pleineset al\.\(2025\)M\. Pleines, M\. Pallasch, F\. Zimmer, and M\. PreussMemory gym: towards endless tasks to benchmark memory capabilities of agents\.Journal of Machine Learning Research26\(6\),pp\. 1–40\.External Links:[Link](http://jmlr.org/papers/v26/24-0043.html)Cited by:[§6](https://arxiv.org/html/2609.25757#S6.p3.1)\.
- Shahet al\.\(2026\)R\. Shah, Y\. Li, F\. Bello, Y\. Zhu, and R\. Martín\-MartínMemory retrieval in visuomotor policies for long\-horizon robot control\.arXiv preprint arXiv:2606\.25136\.Cited by:[§6](https://arxiv.org/html/2609.25757#S6.p3.1)\.
- Shalizi and Crutchfield \(2001\)C\. R\. Shalizi and J\. P\. CrutchfieldComputational mechanics: pattern and prediction, structure and simplicity\.Journal of statistical physics104\(3\),pp\. 817–879\.Cited by:[Table 23](https://arxiv.org/html/2609.25757#A6.T23.10.4.1.1.1),[§1](https://arxiv.org/html/2609.25757#S1.p4.1),[§6](https://arxiv.org/html/2609.25757#S6.p1.1)\.
- Stillet al\.\(2010\)S\. Still, J\. P\. Crutchfield, and C\. J\. EllisonOptimal causal inference: estimating stored information and approximating causal architecture\.Chaos: An Interdisciplinary Journal of Nonlinear Science20\(3\)\.Cited by:[§6](https://arxiv.org/html/2609.25757#S6.p1.1)\.
- Subramanianet al\.\(2022\)J\. Subramanian, A\. Sinha, R\. Seraj, and A\. MahajanApproximate information state for approximate planning and reinforcement learning in partially observed systems\.Journal of Machine Learning Research23\(12\),pp\. 1–83\.Cited by:[Table 23](https://arxiv.org/html/2609.25757#A6.T23.10.6.1.1.1),[§1](https://arxiv.org/html/2609.25757#S1.p4.1),[§6](https://arxiv.org/html/2609.25757#S6.p2.1)\.
- Swamyet al\.\(2022\)G\. Swamy, S\. Choudhury, J\. Bagnell, and S\. WuSequence model imitation learning with unobserved contexts\.Advances in Neural Information Processing Systems35,pp\. 17665–17676\.Cited by:[§6](https://arxiv.org/html/2609.25757#S6.p3.1)\.
- Taoet al\.\(2025\)R\. Y\. Tao, K\. Guo, C\. Allen, and G\. KonidarisBenchmarking partial observability in reinforcement learning with a suite of memory\-improvable domains\.InReinforcement Learning Conference,External Links:[Link](https://openreview.net/forum?id=HUTCbYOW5E)Cited by:[§B\.1](https://arxiv.org/html/2609.25757#A2.SS1.p3.1)\.
- Tishbyet al\.\(2000\)N\. Tishby, F\. C\. Pereira, and W\. BialekThe information bottleneck method\.ArXivphysics/0004057\.External Links:[Link](https://arxiv.org/abs/physics/0004057)Cited by:[§1](https://arxiv.org/html/2609.25757#S1.p6.1),[§3\.1](https://arxiv.org/html/2609.25757#S3.SS1.p1.2)\.
- Villasevilet al\.\(2025\)M\. T\. Villasevil, A\. Tang, Y\. Liu, and C\. FinnLearning long\-context diffusion policies via past\-token prediction\.In9th Annual Conference on Robot Learning,External Links:[Link](https://openreview.net/forum?id=o0LBjJxUeS)Cited by:[§6](https://arxiv.org/html/2609.25757#S6.p3.1)\.
- Walsh \(2026\)M\. WalshSupport sufficiency as action\-sufficient compression: a single\-cycle rate\-regret formulation\.External Links:2606\.09858,[Link](https://arxiv.org/abs/2606.09858)Cited by:[Table 23](https://arxiv.org/html/2609.25757#A6.T23.10.3.1.1.1),[§6](https://arxiv.org/html/2609.25757#S6.p3.1)\.
- Wanget al\.\(2026\)K\. Wang, S\. Yeom, J\. Cao, Y\. Zhi, N\. Shinde, and M\. YipRemember what you did?: learning behavioral memories for partially observable object manipulation\.arXiv preprint arXiv:2606\.21188\.Cited by:[§6](https://arxiv.org/html/2609.25757#S6.p3.1)\.
- Weinzaepfelet al\.\(2026\)P\. Weinzaepfel, C\. Wolf, B\. M\. Sariyildiz, G\. Bono, and G\. MonaciCompressing observation history into agent memory: distilling transformers into recurrent transformers\.arXiv preprint arXiv:2606\.21562\.Cited by:[§6](https://arxiv.org/html/2609.25757#S6.p3.1)\.
- Wenet al\.\(2020\)C\. Wen, J\. Lin, T\. Darrell, D\. Jayaraman, and Y\. GaoFighting copycat agents in behavioral cloning from observation histories\.InAdvances in Neural Information Processing Systems,H\. Larochelle, M\. Ranzato, R\. Hadsell, M\.F\. Balcan, and H\. Lin \(Eds\.\),Vol\.33,pp\. 2564–2575\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2020/file/1b113258af3968aaf3969ca67e744ff8-Paper.pdf)Cited by:[§6](https://arxiv.org/html/2609.25757#S6.p3.1)\.
- Witsenhausen \(1976\)H\. WitsenhausenThe zero\-error side information problem and chromatic numbers \(corresp\.\)\.IEEE Transactions on Information Theory22\(5\),pp\. 592–593\.Cited by:[§A\.3](https://arxiv.org/html/2609.25757#A1.SS3.SSS0.Px1.p1.1),[§2\.3](https://arxiv.org/html/2609.25757#S2.SS3.p2.2)\.
- Wuet al\.\(2024\)L\. Wu, B\. Evans, R\. Islam, R\. Seraj, Y\. Efroni, and A\. LambGeneralizing multi\-step inverse models for representation learning to finite\-memory pomdps\.External Links:2404\.14552,[Link](https://arxiv.org/abs/2404.14552)Cited by:[§6](https://arxiv.org/html/2609.25757#S6.p2.1)\.
- Yaminet al\.\(2026\)K\. Yamin, N\. Deka, M\. Swaroop, A\. Ting, J\. Schneider, and B\. WilderWhat must generalist agents remember?\.arXiv preprint arXiv:2606\.18746\.Cited by:[Appendix F](https://arxiv.org/html/2609.25757#A6.SS0.SSS0.Px1.p1.1),[Table 23](https://arxiv.org/html/2609.25757#A6.T23.10.9.1.1.1),[§6](https://arxiv.org/html/2609.25757#S6.p3.1)\.
- Yu \(2026\)Y\. YuOn the capability separation between world\-model policy learning and imitated world\-action models\.arXiv preprint arXiv:2608\.22197\.Cited by:[§C\.1](https://arxiv.org/html/2609.25757#A3.SS1.SSS0.Px1.p7.1),[§2\.1](https://arxiv.org/html/2609.25757#S2.SS1.p2.1)\.
- Yueet al\.\(2025\)W\. Yue, B\. Liu, and P\. StoneLearning memory mechanisms for decision making through demonstration\.InNew Frontiers in Associative Memories,External Links:[Link](https://openreview.net/forum?id=gmVqnpseHr)Cited by:[§6](https://arxiv.org/html/2609.25757#S6.p3.1)\.
- Zhanget al\.\(2021\)A\. Zhang, R\. T\. McAllister, R\. Calandra, Y\. Gal, and S\. LevineLearning invariant representations for reinforcement learning without reconstruction\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=-2FCwDKRREu)Cited by:[Table 23](https://arxiv.org/html/2609.25757#A6.T23.10.2.1.1.1)\.
- Zhanget al\.\(2026\)Z\. Zhang, Y\. Chen, M\. Imani, and T\. LanMinimal Markovization via stable quotients in holonomy\-cover decision processes\.arXiv preprint arXiv:2607\.27132\.Cited by:[Table 23](https://arxiv.org/html/2609.25757#A6.T23.10.7.1.1.1),[§6](https://arxiv.org/html/2609.25757#S6.p2.1)\.
- Zouet al\.\(2026\)M\. Zou, Z\. Guo, L\. Liang, Z\. Wang, Q\. Wang, Q\. Wen, I\. King, L\. Qu, and Z\. XuRemember the decision, not the description: a rate\-distortion framework for agent memory\.External Links:2605\.10870,[Link](https://arxiv.org/abs/2605.10870)Cited by:[Appendix F](https://arxiv.org/html/2609.25757#A6.SS0.SSS0.Px1.p1.1),[Table 23](https://arxiv.org/html/2609.25757#A6.T23.10.9.1.1.1),[§6](https://arxiv.org/html/2609.25757#S6.p3.1)\.

Guide\.Appendix[A](https://arxiv.org/html/2609.25757#A1)proves the theorems and gives the finite\-instance certificates, including the matching lower and upper bounds for the non\-transitive corridor \([A\.3](https://arxiv.org/html/2609.25757#A1.SS3)\)\. Appendix[B](https://arxiv.org/html/2609.25757#A2)specifies the benchmarks, the observation convention and the estimator controls;[B\.5](https://arxiv.org/html/2609.25757#A2.SS5)compares the full\-trajectory gate with the relaxed one, and[B\.6](https://arxiv.org/html/2609.25757#A2.SS6)reports the injected\-leak audit of §[4\.1](https://arxiv.org/html/2609.25757#S4.SS1)with both detection rules and their clean\-data baseline\. Appendix[C](https://arxiv.org/html/2609.25757#A3)validates the target on toys \(a stochastic expert in[C\.2](https://arxiv.org/html/2609.25757#A3.SS2)\), Task A, the readout tasks and the corridor\. Appendix[D](https://arxiv.org/html/2609.25757#A4)covers learning: the event\-agnostic objective, its held\-out pixel evaluation, its predictive surplus and the jitter intervention in[D\.1](https://arxiv.org/html/2609.25757#A4.SS1), and the codebook sweep of §[5](https://arxiv.org/html/2609.25757#S5)in[D\.5](https://arxiv.org/html/2609.25757#A4.SS5)\. Appendix[E](https://arxiv.org/html/2609.25757#A5)reports additional domains: pixels, the physical weighing task with its hybrid rollout \([E\.3](https://arxiv.org/html/2609.25757#A5.SS3)\), the external Passive T\-maze with small\-state GRUs \([E\.4](https://arxiv.org/html/2609.25757#A5.SS4)\), and the certified community benchmarks of §[4\.2](https://arxiv.org/html/2609.25757#S4.SS2)\([E\.5](https://arxiv.org/html/2609.25757#A5.SS5)\)\. Appendix[F](https://arxiv.org/html/2609.25757#A6)compares representation targets\.

## Appendix AProofs and exact certificates

#### Conventions\.

All statements are made under the expert occupancy at a fixed stepttand hold almost surely\. For concise notation, we state them for*reachable*histories, namely histories with positive probability under the expert\. We therefore assume that the reachable history space is countable at each step, as in the enumerated POMDP used by the solver; in the general case, “for all reachable histories” is replaced by “almost surely\.” Under \(A2\),GE,t=γt​\(Ht\)G\_\{E,t\}=\\gamma\_\{t\}\(H\_\{t\}\)for a measurableγt\\gamma\_\{t\}\. By Definition[1](https://arxiv.org/html/2609.25757#Thmdefinition1), the expert’s action distribution is then a function of\(Ot,GE,t\)\(O\_\{t\},G\_\{E,t\}\), which we write asπE\(⋅∣o,g\)\\pi\_\{E\}\(\\cdot\\mid o,g\)\. A*context*is a random variableCCgenerated fromHtH\_\{t\}by an encoderp⁡\(c∣h\)p\(c\\mid h\), and a decoder is a mapq\(⋅∣o,c\)q\(\\cdot\\mid o,c\)into action distributions\. Zero distortion means𝔼\[δ\(πE\(⋅∣St\),q\(⋅∣Ot,C\)\)\]=0\\mathbb\{E\}\[\\delta\(\\pi\_\{E\}\(\\cdot\\mid S\_\{t\}\),q\(\\cdot\\mid O\_\{t\},C\)\)\]=0\. By \(A3\) andδ≥0\\delta\\geq 0, this condition is equivalent toq\(⋅∣Ot,C\)=πE\(⋅∣St\)q\(\\cdot\\mid O\_\{t\},C\)=\\pi\_\{E\}\(\\cdot\\mid S\_\{t\}\)almost surely\. Assumption \(A1\) is used only to ensure thatH⁡\(GE,t∣Ot\)<∞H\(G\_\{E,t\}\\mid O\_\{t\}\)<\\infty\.

### A\.1Instantaneous minimality

###### Theorem 3\(Behavioral sufficiency\)\.

Under \(A1\)–\(A3\), zero distortion impliesH⁡\(GE,t∣Ot,C\)=0H\(G\_\{E,t\}\\mid O\_\{t\},C\)=0; under \(A2\),GE,tG\_\{E,t\}is the coarsest sufficient context\.

###### Theorem 4\(Rate–distortion reduction\)\.

Under \(A1\)\(A2\), for allDD,RE​\(D\)=RG\|O​\(D\)R\_\{E\}\(D\)=R\_\{G\\mid O\}\(D\), the conditional rate–distortion function of the sourceGE,tG\_\{E,t\}givenOtO\_\{t\}\. In particularRE​\(0\)=H⁡\(GE,t∣Ot\)R\_\{E\}\(0\)=H\(G\_\{E,t\}\\mid O\_\{t\}\)\.

###### Proof of Theorem[3](https://arxiv.org/html/2609.25757#Thmtheorem3)\.

Fix\(o,c\)\(o,c\)such thatP⁡\(Ot=o,C=c\)\>0P\(O\_\{t\}=o,C=c\)\>0, and lets,s′s,s^\{\\prime\}be states in the fiber𝒮o\\mathcal\{S\}\_\{o\}that each occur with positive probability jointly withC=cC=c\. Zero distortion impliesπE\(⋅∣s\)=q\(⋅∣o,c\)=πE\(⋅∣s′\)\\pi\_\{E\}\(\\cdot\\mid s\)=q\(\\cdot\\mid o,c\)=\\pi\_\{E\}\(\\cdot\\mid s^\{\\prime\}\)\. Hences∼Es′s\\sim\_\{E\}s^\{\\prime\}andgE​\(s\)=gE​\(s′\)g\_\{E\}\(s\)=g\_\{E\}\(s^\{\\prime\}\), soGE,tG\_\{E,t\}is almost surely a function of\(Ot,C\)\(O\_\{t\},C\); equivalently,H⁡\(GE,t∣Ot,C\)=0H\(G\_\{E,t\}\\mid O\_\{t\},C\)=0\. This argument does not use \(A2\)\. For the second claim, \(A2\) makes the contextC=GE,t=γt​\(Ht\)C=G\_\{E,t\}=\\gamma\_\{t\}\(H\_\{t\}\)admissible, and the decoderq\(⋅∣o,g\):=πE\(⋅∣o,g\)q\(\\cdot\\mid o,g\):=\\pi\_\{E\}\(\\cdot\\mid o,g\)attains zero distortion\. Thus,GE,tG\_\{E,t\}is sufficient\. By the first claim, every other zero\-distortion contextC′C^\{\\prime\}satisfiesH⁡\(GE,t∣Ot,C′\)=0H\(G\_\{E,t\}\\mid O\_\{t\},C^\{\\prime\}\)=0\. Therefore,GE,tG\_\{E,t\}is a function of\(Ot,C′\)\(O\_\{t\},C^\{\\prime\}\)for every sufficient context, which establishes that it is the coarsest such context\. ∎

LetRG\|O​\(D\)R\_\{G\\mid O\}\(D\)be the conditional rate–distortion function\([Gray, 1973](https://arxiv.org/html/2609.25757#bib.bib11)\)of the sourceGE,tG\_\{E,t\}with side informationOtO\_\{t\}at encoder and decoder, reproduction alphabet the action distributions, and distortiond\(\(o,g\),q\)=δ\(πE\(⋅∣o,g\),q\)d\\big\(\(o,g\),q\\big\)=\\delta\\big\(\\pi\_\{E\}\(\\cdot\\mid o,g\),q\\big\):

RG\|O\(D\)=inf\{I\(C;GE,t∣Ot\):p\(c∣g,o\),q\(⋅∣o,c\),𝔼\[d\(\(Ot,GE,t\),q\(⋅∣Ot,C\)\)\]≤D\}\.R\_\{G\\mid O\}\(D\)=\\inf\\big\\\{I\(C;G\_\{E,t\}\\mid O\_\{t\}\):\\ p\(c\\mid g,o\),\\ q\(\\cdot\\mid o,c\),\\ \\mathbb\{E\}\\big\[d\\big\(\(O\_\{t\},G\_\{E,t\}\),q\(\\cdot\\mid O\_\{t\},C\)\\big\)\\big\]\\leq D\\big\\\}\.BecauseπE\(⋅∣St\)=πE\(⋅∣Ot,GE,t\)\\pi\_\{E\}\(\\cdot\\mid S\_\{t\}\)=\\pi\_\{E\}\(\\cdot\\mid O\_\{t\},G\_\{E,t\}\), the distortion of any encoder–decoder pair in either problem depends on the joint law of\(Ot,GE,t,C\)\(O\_\{t\},G\_\{E,t\},C\)only\.

###### Proof of Theorem[4](https://arxiv.org/html/2609.25757#Thmtheorem4)\.

\(≥\\geq\) Let\(p⁡\(c∣h\),q\)\(p\(c\\mid h\),q\)be feasible forRE​\(D\)R\_\{E\}\(D\)\. Under \(A2\),GE,t=γt​\(Ht\)G\_\{E,t\}=\\gamma\_\{t\}\(H\_\{t\}\), so the chain rule givesI\(C;Ht∣Ot\)=I\(C;GE,t,Ht∣Ot\)=I\(C;GE,t∣Ot\)\+I\(C;Ht∣GE,t,Ot\)≥I\(C;GE,t∣Ot\)I\(C;H\_\{t\}\\mid O\_\{t\}\)=I\(C;G\_\{E,t\},H\_\{t\}\\mid O\_\{t\}\)=I\(C;G\_\{E,t\}\\mid O\_\{t\}\)\+I\(C;H\_\{t\}\\mid G\_\{E,t\},O\_\{t\}\)\\geq I\(C;G\_\{E,t\}\\mid O\_\{t\}\)\. Define the induced encoderp′​\(c∣g,o\):=P⁡\(C=c∣GE,t=g,Ot=o\)p^\{\\prime\}\(c\\mid g,o\):=P\(C=c\\mid G\_\{E,t\}=g,O\_\{t\}=o\)\. The pair\(p′,q\)\(p^\{\\prime\},q\)produces the same joint law of\(Ot,GE,t,C\)\(O\_\{t\},G\_\{E,t\},C\), hence the same distortion, and is feasible forRG\|O​\(D\)R\_\{G\\mid O\}\(D\)with rateI⁡\(C;GE,t∣Ot\)≤I⁡\(C;Ht∣Ot\)I\(C;G\_\{E,t\}\\mid O\_\{t\}\)\\leq I\(C;H\_\{t\}\\mid O\_\{t\}\)\. Taking infima,RG\|O​\(D\)≤RE​\(D\)R\_\{G\\mid O\}\(D\)\\leq R\_\{E\}\(D\)\. \(≤\\leq\) Let\(p′​\(c∣g,o\),q\)\(p^\{\\prime\}\(c\\mid g,o\),q\)be feasible forRG\|O​\(D\)R\_\{G\\mid O\}\(D\)and setp⁡\(c∣h\):=p′​\(c∣γt​\(h\),o⁡\(h\)\)p\(c\\mid h\):=p^\{\\prime\}\\big\(c\\mid\\gamma\_\{t\}\(h\),o\(h\)\\big\)\. The joint law of\(Ot,GE,t,C\)\(O\_\{t\},G\_\{E,t\},C\)and the distortion are unchanged, andCCdepends onHtH\_\{t\}only through\(GE,t,Ot\)\(G\_\{E,t\},O\_\{t\}\), soI\(C;Ht∣GE,t,Ot\)=0I\(C;H\_\{t\}\\mid G\_\{E,t\},O\_\{t\}\)=0andI⁡\(C;Ht∣Ot\)=I⁡\(C;GE,t∣Ot\)I\(C;H\_\{t\}\\mid O\_\{t\}\)=I\(C;G\_\{E,t\}\\mid O\_\{t\}\)\. HenceRE​\(D\)≤RG\|O​\(D\)R\_\{E\}\(D\)\\leq R\_\{G\\mid O\}\(D\)\. AtD=0D=0, every feasible pair hasH⁡\(GE,t∣Ot,C\)=0H\(G\_\{E,t\}\\mid O\_\{t\},C\)=0\(Theorem[3](https://arxiv.org/html/2609.25757#Thmtheorem3)\), soI⁡\(C;GE,t∣Ot\)=H⁡\(GE,t∣Ot\)−H⁡\(GE,t∣Ot,C\)=H⁡\(GE,t∣Ot\)I\(C;G\_\{E,t\}\\mid O\_\{t\}\)=H\(G\_\{E,t\}\\mid O\_\{t\}\)\-H\(G\_\{E,t\}\\mid O\_\{t\},C\)=H\(G\_\{E,t\}\\mid O\_\{t\}\), attained byC=GE,tC=G\_\{E,t\}; thusRE​\(0\)=RG\|O​\(0\)=H⁡\(GE,t∣Ot\)R\_\{E\}\(0\)=R\_\{G\\mid O\}\(0\)=H\(G\_\{E,t\}\\mid O\_\{t\}\), finite by \(A1\)\. ∎

###### Corollary 1\(Harmless ambiguity and selective forgetting\)\.

\(a\) LetUUbe a hidden variable such thatGE,tG\_\{E,t\}is a function of\(Ot,U\)\(O\_\{t\},U\)\(e\.g\. the hidden part of the state\)\. ThenH⁡\(U∣Ot\)=H⁡\(GE,t∣Ot\)\+H⁡\(U∣GE,t,Ot\)H\(U\\mid O\_\{t\}\)=H\(G\_\{E,t\}\\mid O\_\{t\}\)\+H\(U\\mid G\_\{E,t\},O\_\{t\}\); withMMequiprobable hidden modes partitioned intoRRclasses ofM/RM/Rmodes each,H⁡\(U∣GE,t,Ot\)=log2⁡\(M/R\)H\(U\\mid G\_\{E,t\},O\_\{t\}\)=\\log\_\{2\}\(M/R\)\. \(b\) IfH⁡\(GE,t∣Ot,C\)=0H\(G\_\{E,t\}\\mid O\_\{t\},C\)=0andH⁡\(C∣Ot\)=H⁡\(GE,t∣Ot\)H\(C\\mid O\_\{t\}\)=H\(G\_\{E,t\}\\mid O\_\{t\}\), thenH⁡\(C∣GE,t,Ot\)=0H\(C\\mid G\_\{E,t\},O\_\{t\}\)=0andI\(C;V∣GE,t,Ot\)=0I\(C;V\\mid G\_\{E,t\},O\_\{t\}\)=0for every random variableVV, in particular for any nuisanceUnuiU\_\{\\rm nui\}\. \(c\) For everyDDandϵ\>0\\epsilon\>0there is an encoder with distortion at mostDDand rate at mostRE​\(D\)\+ϵR\_\{E\}\(D\)\+\\epsilonof the formp⁡\(c∣h\)=p′​\(c∣γt​\(h\),o⁡\(h\)\)p\(c\\mid h\)=p^\{\\prime\}\(c\\mid\\gamma\_\{t\}\(h\),o\(h\)\), and for itI\(C;V∣GE,t,Ot\)=0I\(C;V\\mid G\_\{E,t\},O\_\{t\}\)=0for every source\-side variableVVwhose joint distribution is fixed before the encoder draws its independent randomness\.

###### Proof\.

\(a\)H⁡\(U∣Ot\)=H⁡\(U,GE,t∣Ot\)=H⁡\(GE,t∣Ot\)\+H⁡\(U∣GE,t,Ot\)H\(U\\mid O\_\{t\}\)=H\(U,G\_\{E,t\}\\mid O\_\{t\}\)=H\(G\_\{E,t\}\\mid O\_\{t\}\)\+H\(U\\mid G\_\{E,t\},O\_\{t\}\)becauseH⁡\(GE,t∣U,Ot\)=0H\(G\_\{E,t\}\\mid U,O\_\{t\}\)=0; the second statement is the entropy of a uniform variable onM/RM/Rvalues\. \(b\)H⁡\(C∣GE,t,Ot\)=H⁡\(C∣Ot\)−I⁡\(C;GE,t∣Ot\)=H⁡\(C∣Ot\)−\(H⁡\(GE,t∣Ot\)−H⁡\(GE,t∣Ot,C\)\)=0H\(C\\mid G\_\{E,t\},O\_\{t\}\)=H\(C\\mid O\_\{t\}\)\-I\(C;G\_\{E,t\}\\mid O\_\{t\}\)=H\(C\\mid O\_\{t\}\)\-\\big\(H\(G\_\{E,t\}\\mid O\_\{t\}\)\-H\(G\_\{E,t\}\\mid O\_\{t\},C\)\\big\)=0, and0≤I\(C;V∣GE,t,Ot\)≤H\(C∣GE,t,Ot\)=00\\leq I\(C;V\\mid G\_\{E,t\},O\_\{t\}\)\\leq H\(C\\mid G\_\{E,t\},O\_\{t\}\)=0\. \(c\) The \(≤\\leq\) direction of the proof of Theorem[4](https://arxiv.org/html/2609.25757#Thmtheorem4)turns any encoder forRG\|O​\(D\)R\_\{G\\mid O\}\(D\)withinϵ\\epsilonof the infimum into an encoder of the stated form with the same rate and distortion; for itP⁡\(C=c∣Ht=h,V=v\)=p′​\(c∣γt​\(h\),o⁡\(h\)\)P\(C=c\\mid H\_\{t\}=h,V=v\)=p^\{\\prime\}\(c\\mid\\gamma\_\{t\}\(h\),o\(h\)\)for every such source\-sideVV, soCCis conditionally independent ofVVgiven\(GE,t,Ot\)\(G\_\{E,t\},O\_\{t\}\)\. ∎

Part \(c\) is an existence statement: nuisance information is never*required*to reach the frontier, but not every point on the frontier must be nuisance\-free\.

###### Corollary 2\(Closed form\)\.

SupposeGE,tG\_\{E,t\}is uniform onR≥2R\\geq 2classes and independent ofOtO\_\{t\}, the expert is deterministic with distinct actions across classes, andδ\\deltais total variation\. ThenRE​\(D\)=log2⁡R−h2​\(D\)−D​log2⁡\(R−1\)R\_\{E\}\(D\)=\\log\_\{2\}R\-h\_\{2\}\(D\)\-D\\log\_\{2\}\(R\-1\)for0≤D≤1−1/R0\\leq D\\leq 1\-1/RandRE​\(D\)=0R\_\{E\}\(D\)=0forD≥1−1/RD\\geq 1\-1/R, whereh2h\_\{2\}is the binary entropy\.

###### Proof\.

By Theorem[4](https://arxiv.org/html/2609.25757#Thmtheorem4)it suffices to computeRG\|O​\(D\)R\_\{G\\mid O\}\(D\)\. SinceGE,t⟂OtG\_\{E,t\}\\perp O\_\{t\}, a scheme conditioned onOtO\_\{t\}is a family of unconditional schemes indexed byoo, with rate𝔼o​\[I⁡\(C;GE,t∣Ot=o\)\]≥𝔼o​\[RG​\(Do\)\]≥RG​\(𝔼o​Do\)≥RG​\(D\)\\mathbb\{E\}\_\{o\}\[I\(C;G\_\{E,t\}\\mid O\_\{t\}=o\)\]\\geq\\mathbb\{E\}\_\{o\}\[R\_\{G\}\(D\_\{o\}\)\]\\geq R\_\{G\}\(\\mathbb\{E\}\_\{o\}D\_\{o\}\)\\geq R\_\{G\}\(D\)by convexity of the unconditional rate–distortion functionRGR\_\{G\}, and a scheme that ignoresooattainsRG​\(D\)R\_\{G\}\(D\); soRG\|O=RGR\_\{G\\mid O\}=R\_\{G\}\. With a deterministic expert taking actionaga\_\{g\}in classgg,δ\(πE\(⋅∣g\),q\)=1−q\(ag\)\\delta\\big\(\\pi\_\{E\}\(\\cdot\\mid g\),q\\big\)=1\-q\(a\_\{g\}\), which is linear inqq; for a fixed encoder the decoder minimizing𝔼⁡\[1−q⁡\(aG\)∣c\]\\mathbb\{E\}\[1\-q\(a\_\{G\}\)\\mid c\]is a point mass on the most probable class, so restricting reproductions to class estimatesG^\\hat\{G\}with Hamming distortion𝟏\[G^≠GE,t\]\\mathbf\{1\}\[\\hat\{G\}\\neq G\_\{E,t\}\]loses nothing\. Converse: ifP⁡\(G^≠GE,t\)=Pe≤D≤1−1/RP\(\\hat\{G\}\\neq G\_\{E,t\}\)=P\_\{e\}\\leq D\\leq 1\-1/R, Fano’s inequality givesH⁡\(GE,t∣G^\)≤h2​\(Pe\)\+Pe​log2⁡\(R−1\)≤h2​\(D\)\+D​log2⁡\(R−1\)H\(G\_\{E,t\}\\mid\\hat\{G\}\)\\leq h\_\{2\}\(P\_\{e\}\)\+P\_\{e\}\\log\_\{2\}\(R\-1\)\\leq h\_\{2\}\(D\)\+D\\log\_\{2\}\(R\-1\), the right\-hand side being non\-decreasing on\[0,1−1/R\]\[0,1\-1/R\], soI⁡\(GE,t,G^\)≥log2⁡R−h2​\(D\)−D​log2⁡\(R−1\)I\(G\_\{E,t\};\\hat\{G\}\)\\geq\\log\_\{2\}R\-h\_\{2\}\(D\)\-D\\log\_\{2\}\(R\-1\)\. Achievability: takeG^\\hat\{G\}uniform andGE,t=G^G\_\{E,t\}=\\hat\{G\}with probability1−D1\-D, otherwise uniform on the otherR−1R\-1classes; the marginal ofGE,tG\_\{E,t\}is uniform, the distortion isDD, andI⁡\(GE,t,G^\)=log2⁡R−h2​\(D\)−D​log2⁡\(R−1\)I\(G\_\{E,t\};\\hat\{G\}\)=\\log\_\{2\}R\-h\_\{2\}\(D\)\-D\\log\_\{2\}\(R\-1\)\. ForD≥1−1/RD\\geq 1\-1/Ra constantG^\\hat\{G\}has rate00and distortion1−1/R1\-1/R\. ∎

###### Proposition 1\(Deterministic frontier\)\.

LetREdet​\(D\)=infH⁡\(C∣Ot\)R^\{\\det\}\_\{E\}\(D\)=\\inf H\(C\\mid O\_\{t\}\)over deterministic encodersC=f⁡\(Ht\)C=f\(H\_\{t\}\)and decoders with distortion at mostDD, and letRG\|Odet​\(D\)R^\{\\det\}\_\{G\\mid O\}\(D\)be the same infimum overC=f⁡\(GE,t,Ot\)C=f\(G\_\{E,t\},O\_\{t\}\)\. Under \(A1\)\(A2\),RE​\(D\)≤REdet​\(D\)≤RG\|Odet​\(D\)R\_\{E\}\(D\)\\leq R^\{\\det\}\_\{E\}\(D\)\\leq R^\{\\det\}\_\{G\\mid O\}\(D\)for allDD, with equality throughout atD=0D=0\.

###### Proof\.

For a deterministic encoderI⁡\(C;Ht∣Ot\)=H⁡\(C∣Ot\)−H⁡\(C∣Ht,Ot\)=H⁡\(C∣Ot\)I\(C;H\_\{t\}\\mid O\_\{t\}\)=H\(C\\mid O\_\{t\}\)\-H\(C\\mid H\_\{t\},O\_\{t\}\)=H\(C\\mid O\_\{t\}\), so every deterministic feasible pair is feasible forRE​\(D\)R\_\{E\}\(D\)with the same rate:RE≤REdetR\_\{E\}\\leq R^\{\\det\}\_\{E\}\. Under \(A2\) everyf⁡\(GE,t,Ot\)f\(G\_\{E,t\},O\_\{t\}\)equals the deterministic functionf⁡\(γt​\(Ht\),o⁡\(Ht\)\)f\(\\gamma\_\{t\}\(H\_\{t\}\),o\(H\_\{t\}\)\)ofHtH\_\{t\}, with the same joint law of\(Ot,GE,t,C\)\(O\_\{t\},G\_\{E,t\},C\), hence the same distortion and the sameH⁡\(C∣Ot\)H\(C\\mid O\_\{t\}\):REdet≤RG\|OdetR^\{\\det\}\_\{E\}\\leq R^\{\\det\}\_\{G\\mid O\}\. AtD=0D=0,C=GE,tC=G\_\{E,t\}is feasible forRG\|OdetR^\{\\det\}\_\{G\\mid O\}with rateH⁡\(GE,t∣Ot\)=RE​\(0\)H\(G\_\{E,t\}\\mid O\_\{t\}\)=R\_\{E\}\(0\)\(Theorem[4](https://arxiv.org/html/2609.25757#Thmtheorem4)\), which closes the chain\. ∎

### A\.2Recurrent minimality and support conditions

###### Assumption 1\(Continuation\-support homogeneity, A4\)\.

For alltt, equality of\(Ot,GE,t\)\(O\_\{t\},G\_\{E,t\}\)for two histories implies equality of their continuation supportsUt​\(h\)U\_\{t\}\(h\)\.

###### Lemma 1\.

Under \(A4\),∼t\\sim\_\{t\}is an equivalence for everyttand∼t=∼st\\sim\_\{t\}=\\sim^\{s\}\_\{t\};Γt:=\[Ht\]∼t\\Gamma\_\{t\}:=\[H\_\{t\}\]\_\{\\sim\_\{t\}\}is a right congruence,Γt\+1=Φt​\(Γt,Ot\+1,At\)\\Gamma\_\{t\+1\}=\\Phi\_\{t\}\(\\Gamma\_\{t\},O\_\{t\+1\},A\_\{t\}\), andGE,tG\_\{E,t\}is a function of\(Ot,Γt\)\(O\_\{t\},\\Gamma\_\{t\}\)\.

###### Theorem 5\(Exact regime\)\.

Under \(A1\)–\(A4\),H⁡\(Ct∣Ot\)≥H⁡\(Γt∣Ot\)≥H⁡\(GE,t∣Ot\)H\(C\_\{t\}\\mid O\_\{t\}\)\\geq H\(\\Gamma\_\{t\}\\mid O\_\{t\}\)\\geq H\(G\_\{E,t\}\\mid O\_\{t\}\)for every zero\-distortion realization, andCt=ΓtC\_\{t\}=\\Gamma\_\{t\},F=ΦF=\\Phiattains it:Rtmem​\(0\)=H⁡\(Γt∣Ot\)R^\{\\rm mem\}\_\{t\}\(0\)=H\(\\Gamma\_\{t\}\\mid O\_\{t\}\)\.

Recall from §[2](https://arxiv.org/html/2609.25757#S2)the continuation supportUt​\(h\)=supp⁡PE​\(At,Ot\+1∣Ht=h\)U\_\{t\}\(h\)=\\operatorname\{supp\}P\_\{E\}\(A\_\{t\},O\_\{t\+1\}\\mid H\_\{t\}=h\), the extended historyh​u=\(h,at,ot\+1\)hu=\(h,a\_\{t\},o\_\{t\+1\}\)foru=\(at,ot\+1\)u=\(a\_\{t\},o\_\{t\+1\}\), the relation∼t\\sim\_\{t\}of Definition[3](https://arxiv.org/html/2609.25757#Thmdefinition3), and the strong relation∼st\\sim^\{s\}\_\{t\}:h∼sTh′h\\sim^\{s\}\_\{T\}h^\{\\prime\}iffh∼Th′h\\sim\_\{T\}h^\{\\prime\}, and fort<Tt<T,h∼sth′h\\sim^\{s\}\_\{t\}h^\{\\prime\}iffOt​\(h\)=Ot​\(h′\)O\_\{t\}\(h\)=O\_\{t\}\(h^\{\\prime\}\),GE,t​\(h\)=GE,t​\(h′\)G\_\{E,t\}\(h\)=G\_\{E,t\}\(h^\{\\prime\}\),Ut​\(h\)=Ut​\(h′\)U\_\{t\}\(h\)=U\_\{t\}\(h^\{\\prime\}\)andhu∼st\+1h′uhu\\sim^\{s\}\_\{t\+1\}h^\{\\prime\}ufor everyu∈Ut​\(h\)u\\in U\_\{t\}\(h\)\. A recurrent realization hasCt=Ft​\(Ct−1,Ot,At−1\)C\_\{t\}=F\_\{t\}\(C\_\{t\-1\},O\_\{t\},A\_\{t\-1\}\)with a fixedC0C\_\{0\}and deterministicFtF\_\{t\}, soCt=ct​\(Ht\)C\_\{t\}=c\_\{t\}\(H\_\{t\}\)is a deterministic function of the history andH⁡\(Ct∣Ot\)=I⁡\(Ct;Ht∣Ot\)H\(C\_\{t\}\\mid O\_\{t\}\)=I\(C\_\{t\};H\_\{t\}\\mid O\_\{t\}\); zero distortion meansπ^\(⋅∣Ot,Ct\)=πE\(⋅∣St\)\\hat\{\\pi\}\(\\cdot\\mid O\_\{t\},C\_\{t\}\)=\\pi\_\{E\}\(\\cdot\\mid S\_\{t\}\)almost surely for everyttunder the expert occupancy\. A set of histories is*∼t\\sim\_\{t\}\-compatible*if its members are pairwise∼t\\sim\_\{t\}\-related\. Two facts are used repeatedly: \(F1\) ifh∼th′h\\sim\_\{t\}h^\{\\prime\}orh∼sth′h\\sim^\{s\}\_\{t\}h^\{\\prime\}thenOt​\(h\)=Ot​\(h′\)O\_\{t\}\(h\)=O\_\{t\}\(h^\{\\prime\}\)andGE,t​\(h\)=GE,t​\(h′\)G\_\{E,t\}\(h\)=G\_\{E,t\}\(h^\{\\prime\}\); \(F2\)u∈Ut​\(h\)u\\in U\_\{t\}\(h\)iffh​uhuis a reachable history att\+1t\+1, and thenOt\+1​\(h​u\)=ot\+1O\_\{t\+1\}\(hu\)=o\_\{t\+1\}\.

###### Lemma 2\(Strong congruence\)\.

For everytt,∼st\\sim^\{s\}\_\{t\}is an equivalence relation and a right congruence: ifh∼sth′h\\sim^\{s\}\_\{t\}h^\{\\prime\}andu∈Ut​\(h\)=Ut​\(h′\)u\\in U\_\{t\}\(h\)=U\_\{t\}\(h^\{\\prime\}\)thenhu∼st\+1h′uhu\\sim^\{s\}\_\{t\+1\}h^\{\\prime\}u\. ConsequentlyΓts:=\[Ht\]∼st\\Gamma^\{s\}\_\{t\}:=\[H\_\{t\}\]\_\{\\sim^\{s\}\_\{t\}\}satisfiesΓt\+1s=Φts​\(Γts,Ot\+1,At\)\\Gamma^\{s\}\_\{t\+1\}=\\Phi^\{s\}\_\{t\}\(\\Gamma^\{s\}\_\{t\},O\_\{t\+1\},A\_\{t\}\)for a deterministicΦts\\Phi^\{s\}\_\{t\}, and\(Ot,GE,t\)\(O\_\{t\},G\_\{E,t\}\)is a function ofΓts\\Gamma^\{s\}\_\{t\}\.

###### Proof\.

Reflexivity and symmetry are immediate\. Transitivity by backward induction ontt:∼sT\\sim^\{s\}\_\{T\}is equality of the functionh↦\(OT​\(h\),GE,T​\(h\)\)h\\mapsto\(O\_\{T\}\(h\),G\_\{E,T\}\(h\)\)\. Fort<Tt<T, leth1∼sth2h\_\{1\}\\sim^\{s\}\_\{t\}h\_\{2\}andh2∼sth3h\_\{2\}\\sim^\{s\}\_\{t\}h\_\{3\}; thenUt​\(h1\)=Ut​\(h2\)=Ut​\(h3\)=:UU\_\{t\}\(h\_\{1\}\)=U\_\{t\}\(h\_\{2\}\)=U\_\{t\}\(h\_\{3\}\)=:Uand for everyu∈Uu\\in U,h1u∼st\+1h2uh\_\{1\}u\\sim^\{s\}\_\{t\+1\}h\_\{2\}uandh2u∼st\+1h3uh\_\{2\}u\\sim^\{s\}\_\{t\+1\}h\_\{3\}u, soh1u∼st\+1h3uh\_\{1\}u\\sim^\{s\}\_\{t\+1\}h\_\{3\}uby the induction hypothesis; with \(F1\) this givesh1∼sth3h\_\{1\}\\sim^\{s\}\_\{t\}h\_\{3\}\. The right\-congruence property is the last conjunct of the definition; it says that the class ofh​uhuis determined by the class ofhhand byuu, which definesΦts\\Phi^\{s\}\_\{t\}on reachable pairs\. \(F1\) gives the last claim\. ∎

###### Proof of Lemma[1](https://arxiv.org/html/2609.25757#Thmlemma1)\.

We show by backward induction onttthat under \(A4\)∼t=∼st\\sim\_\{t\}=\\sim^\{s\}\_\{t\}; the remaining claims then follow from Lemma[2](https://arxiv.org/html/2609.25757#Thmlemma2)\. Att=Tt=Tthe relations coincide by definition\. Fort<Tt<T,h∼sth′h\\sim^\{s\}\_\{t\}h^\{\\prime\}impliesh∼th′h\\sim\_\{t\}h^\{\\prime\}, because the strong relation requires agreement on every continuation of the common support and∼t\+1=∼st\+1\\sim\_\{t\+1\}=\\sim^\{s\}\_\{t\+1\}by induction\. Conversely leth∼th′h\\sim\_\{t\}h^\{\\prime\}\. By \(F1\) the two histories share\(Ot,GE,t\)\(O\_\{t\},G\_\{E,t\}\), so \(A4\) givesUt​\(h\)=Ut​\(h′\)=:UU\_\{t\}\(h\)=U\_\{t\}\(h^\{\\prime\}\)=:UandUt​\(h\)∩Ut​\(h′\)=UU\_\{t\}\(h\)\\cap U\_\{t\}\(h^\{\\prime\}\)=U; the last conjunct of Definition[3](https://arxiv.org/html/2609.25757#Thmdefinition3)then sayshu∼t\+1h′uhu\\sim\_\{t\+1\}h^\{\\prime\}u, i\.e\.hu∼st\+1h′uhu\\sim^\{s\}\_\{t\+1\}h^\{\\prime\}uby induction, for everyu∈Uu\\in U, which is exactlyh∼sth′h\\sim^\{s\}\_\{t\}h^\{\\prime\}\. Hence∼t\\sim\_\{t\}is an equivalence,Γt=Γts\\Gamma\_\{t\}=\\Gamma^\{s\}\_\{t\}is well defined,Γt\+1=Φt​\(Γt,Ot\+1,At\)\\Gamma\_\{t\+1\}=\\Phi\_\{t\}\(\\Gamma\_\{t\},O\_\{t\+1\},A\_\{t\}\)withΦt=Φts\\Phi\_\{t\}=\\Phi^\{s\}\_\{t\}, andGE,tG\_\{E,t\}is a function ofΓt\\Gamma\_\{t\}, a fortiori of\(Ot,Γt\)\(O\_\{t\},\\Gamma\_\{t\}\)\. ∎

The same induction directly establishes transitivity of∼t\\sim\_\{t\}under \(A4\)\. Ifh1∼th2∼th3h\_\{1\}\\sim\_\{t\}h\_\{2\}\\sim\_\{t\}h\_\{3\}, all three histories share\(Ot,GE,t\)\(O\_\{t\},G\_\{E,t\}\); \(A4\) makes their supports equal, and transitivity att\+1t\+1then transfers tott\. Assumption \(A4\) is sufficient but not necessary: the re\-reveal example in the remark below violates \(A4\) at one step, although∼t\\sim\_\{t\}is transitive at every step\.

###### Proof of Theorem[2](https://arxiv.org/html/2609.25757#Thmtheorem2)\.

*Compatibility of cells\.*Fix a zero\-distortion realization and callℋt\(o,c\)=\{hreachable:Ot\(h\)=o,ct\(h\)=c\}\\mathcal\{H\}\_\{t\}\(o,c\)=\\\{h\\ \\text\{reachable\}:O\_\{t\}\(h\)=o,\\ c\_\{t\}\(h\)=c\\\}its cells\. We show by backward induction onttthat every cell is∼t\\sim\_\{t\}\-compatible\. Leth,h′∈ℋt​\(o,c\)h,h^\{\\prime\}\\in\\mathcal\{H\}\_\{t\}\(o,c\)\. Zero distortion givesπE\(⋅∣s\)=π^\(⋅∣o,c\)=πE\(⋅∣s′\)\\pi\_\{E\}\(\\cdot\\mid s\)=\\hat\{\\pi\}\(\\cdot\\mid o,c\)=\\pi\_\{E\}\(\\cdot\\mid s^\{\\prime\}\)for \(almost\) all statesssconsistent withhhands′s^\{\\prime\}consistent withh′h^\{\\prime\}, soGE,t​\(h\)=GE,t​\(h′\)G\_\{E,t\}\(h\)=G\_\{E,t\}\(h^\{\\prime\}\)by Definition[1](https://arxiv.org/html/2609.25757#Thmdefinition1)and \(A2\)\. Ift=Tt=Tthis ish∼Th′h\\sim\_\{T\}h^\{\\prime\}\. Ift<Tt<T, takeu=\(a,o′\)∈Ut​\(h\)∩Ut​\(h′\)u=\(a,o^\{\\prime\}\)\\in U\_\{t\}\(h\)\\cap U\_\{t\}\(h^\{\\prime\}\)\. By \(F2\) bothh​uhuandh′​uh^\{\\prime\}uare reachable and shareOt\+1=o′O\_\{t\+1\}=o^\{\\prime\};Ft\+1F\_\{t\+1\}being deterministic, they also shareCt\+1=Ft\+1​\(c,o′,a\)C\_\{t\+1\}=F\_\{t\+1\}\(c,o^\{\\prime\},a\), so they lie in one cell att\+1t\+1and are∼t\+1\\sim\_\{t\+1\}\-related by the induction hypothesis\. Henceh∼th′h\\sim\_\{t\}h^\{\\prime\}\. Transitivity of∼t\\sim\_\{t\}is never used\.*Lower bound\.*By the previous paragraphGE,tG\_\{E,t\}is a function of\(Ot,Ct\)\(O\_\{t\},C\_\{t\}\)\(equivalently, Theorem[3](https://arxiv.org/html/2609.25757#Thmtheorem3)applied to the contextCtC\_\{t\}\), soH⁡\(Ct∣Ot\)≥I⁡\(Ct;GE,t∣Ot\)=H⁡\(GE,t∣Ot\)H\(C\_\{t\}\\mid O\_\{t\}\)\\geq I\(C\_\{t\};G\_\{E,t\}\\mid O\_\{t\}\)=H\(G\_\{E,t\}\\mid O\_\{t\}\)for every zero\-distortion realization, andRtmem​\(0\)≥H⁡\(GE,t∣Ot\)R^\{\\rm mem\}\_\{t\}\(0\)\\geq H\(G\_\{E,t\}\\mid O\_\{t\}\)\.*Upper bound\.*Ct:=ΓtsC\_\{t\}:=\\Gamma^\{s\}\_\{t\}withFt:=ΦtsF\_\{t\}:=\\Phi^\{s\}\_\{t\}is a recurrent realization \(Lemma[2](https://arxiv.org/html/2609.25757#Thmlemma2);C0C\_\{0\}is the class of the empty history\)\. The policyπ^\(⋅∣o,γ\):=πE\(⋅∣o,g\)\\hat\{\\pi\}\(\\cdot\\mid o,\\gamma\):=\\pi\_\{E\}\(\\cdot\\mid o,g\), withggthe common value ofGE,tG\_\{E,t\}on the classγ\\gamma, is well defined by \(F1\) and has zero distortion\. HenceRtmem​\(0\)≤H⁡\(Γts∣Ot\)R^\{\\rm mem\}\_\{t\}\(0\)\\leq H\(\\Gamma^\{s\}\_\{t\}\\mid O\_\{t\}\)\. \(This realization has as many states as∼st\\sim^\{s\}\_\{t\}has classes; in the countable setting adopted here the infimum definingRtmem​\(0\)R^\{\\rm mem\}\_\{t\}\(0\)ranges over countable state spaces, and on the enumerated POMDPs of the experiments all class counts are finite\.\) Strictness of both sides is shown in the Remark below\. ∎

###### Proof of Theorem[5](https://arxiv.org/html/2609.25757#Thmtheorem5)\.

Under \(A4\),∼t\\sim\_\{t\}is an equivalence \(Lemma[1](https://arxiv.org/html/2609.25757#Thmlemma1)\), so a∼t\\sim\_\{t\}\-compatible set is contained in a single class:Γt\\Gamma\_\{t\}is a function of\(Ot,Ct\)\(O\_\{t\},C\_\{t\}\)for every zero\-distortion realization, andH⁡\(Ct∣Ot\)≥I⁡\(Ct;Γt∣Ot\)=H⁡\(Γt∣Ot\)≥H⁡\(GE,t∣Ot\)H\(C\_\{t\}\\mid O\_\{t\}\)\\geq I\(C\_\{t\};\\Gamma\_\{t\}\\mid O\_\{t\}\)=H\(\\Gamma\_\{t\}\\mid O\_\{t\}\)\\geq H\(G\_\{E,t\}\\mid O\_\{t\}\), the last step becauseGE,tG\_\{E,t\}is a function ofΓt\\Gamma\_\{t\}\. The realizationCt=Γt=ΓtsC\_\{t\}=\\Gamma\_\{t\}=\\Gamma^\{s\}\_\{t\},F=ΦF=\\Phiof the previous paragraph attainsH⁡\(Γt∣Ot\)H\(\\Gamma\_\{t\}\\mid O\_\{t\}\), soRtmem​\(0\)=H⁡\(Γt∣Ot\)R^\{\\rm mem\}\_\{t\}\(0\)=H\(\\Gamma\_\{t\}\\mid O\_\{t\}\)\. ∎

#### Remark \(both bounds can be strict\)\.

\(i\) Let three equiprobable reachable historiesh1,h2,h3h\_\{1\},h\_\{2\},h\_\{3\}at stept=T−1t=T\-1share\(Ot,GE,t\)\(O\_\{t\},G\_\{E,t\}\), withUt​\(h1\)=\{u,v1\}U\_\{t\}\(h\_\{1\}\)=\\\{u,v\_\{1\}\\\},Ut​\(h3\)=\{u,v3\}U\_\{t\}\(h\_\{3\}\)=\\\{u,v\_\{3\}\\\},Ut​\(h2\)=\{w\}U\_\{t\}\(h\_\{2\}\)=\\\{w\\\}for distinctu,v1,v3,wu,v\_\{1\},v\_\{3\},w, and let the expert require different actions afterh1​uh\_\{1\}uandh3​uh\_\{3\}u\(distinct classes atTT\)\. Thenh1∼th2h\_\{1\}\\sim\_\{t\}h\_\{2\}andh2∼th3h\_\{2\}\\sim\_\{t\}h\_\{3\}hold vacuously whileh1≁th3h\_\{1\}\\not\\sim\_\{t\}h\_\{3\}:∼t\\sim\_\{t\}is not transitive and \(A4\) fails\. HereH⁡\(GE,t∣Ot\)=0H\(G\_\{E,t\}\\mid O\_\{t\}\)=0; the three supports are distinct, soΓts\\Gamma^\{s\}\_\{t\}has three classes andH⁡\(Γts∣Ot\)=log2⁡3H\(\\Gamma^\{s\}\_\{t\}\\mid O\_\{t\}\)=\\log\_\{2\}3; a zero\-distortion realization may mergeh2h\_\{2\}with either neighbour but neverh1h\_\{1\}withh3h\_\{3\}, soRtmem​\(0\)=h2​\(1/3\)R^\{\\rm mem\}\_\{t\}\(0\)=h\_\{2\}\(1/3\), strictly inside\[0,log2⁡3\]\[0,\\log\_\{2\}3\]\. \(ii\) If a behavioral class was revealed beforett, is absent fromOtO\_\{t\}, is re\-revealed att\+1t\+1, and is first used aftert\+1t\+1, then histories that differ only in that class have disjoint continuation supports attt\(their next observations differ\), are vacuously∼t\\sim\_\{t\}\-related and are merged byΓt\\Gamma\_\{t\}, whereasΓts\\Gamma^\{s\}\_\{t\}separates them because their supports differ:Γs\\Gamma^\{s\}charges a bit that the future will re\-provide\. In this instance∼t\\sim\_\{t\}is transitive at every step although \(A4\) fails attt\. Without transitivity, the zero\-distortion memories are exactly the closed compatible state assignments of Appendix[A\.3](https://arxiv.org/html/2609.25757#A1.SS3), whose minimum entropy we compute by dynamic programming\.

###### Corollary 3\(Codebook\)\.

Under \(A4\), any sufficient realization hasK≥maxt,o⁡\|Γt\|o≥RK\\geq\\max\_\{t,o\}\|\\Gamma\_\{t\}\|\_\{o\}\\geq R; the left side can exceedRR\(A′: the joint class\(β1,β2\)\(\\beta\_\{1\},\\beta\_\{2\}\)has four values whileR1=R2=2R\_\{1\}=R\_\{2\}=2during gap1\)\.

###### Proof\.

Here\|Γt\|o\|\\Gamma\_\{t\}\|\_\{o\}is the number of classes ofΓt\\Gamma\_\{t\}withOt=oO\_\{t\}=oandR=maxt,o⁡\|GE,t\|oR=\\max\_\{t,o\}\|G\_\{E,t\}\|\_\{o\}the largest number of behavioral classes in a fiber\. Under \(A4\) the exact\-regime paragraph of the proof of Theorem[5](https://arxiv.org/html/2609.25757#Thmtheorem5)shows that, for every sufficient realization and every\(t,o\)\(t,o\),Γt\\Gamma\_\{t\}is a function of\(Ot,Ct\)\(O\_\{t\},C\_\{t\}\); the mapc↦Γtc\\mapsto\\Gamma\_\{t\}on the codes that occur together withOt=oO\_\{t\}=ois therefore onto the\|Γt\|o\|\\Gamma\_\{t\}\|\_\{o\}classes withOt=oO\_\{t\}=o, andK≥\|Γt\|oK\\geq\|\\Gamma\_\{t\}\|\_\{o\}\. SinceGE,tG\_\{E,t\}is a function ofΓt\\Gamma\_\{t\},\|Γt\|o≥\|GE,t\|o\|\\Gamma\_\{t\}\|\_\{o\}\\geq\|G\_\{E,t\}\|\_\{o\}, whose maximum isRR\. In A′during gap1, bothβ1\\beta\_\{1\}\(used at the grasp\) andβ2\\beta\_\{2\}\(used at the place\) must be carried although each pending decision has only two behavioral classes, so\|Γt\|o=\|β1∨β2\|=4\|\\Gamma\_\{t\}\|\_\{o\}=\|\\beta\_\{1\}\\vee\\beta\_\{2\}\|=4\. ∎

###### Proof of Theorem[1](https://arxiv.org/html/2609.25757#Thmtheorem1)\.

*Well\-definedness\.*If∼t\\sim\_\{t\}is transitive it is an equivalence \(reflexive and symmetric by definition\), soΓt\\Gamma\_\{t\}is the partition into classes\. For a classγ\\gammaatttandu=\(a,o′\)u=\(a,o^\{\\prime\}\)withu∈Ut​\(h\)u\\in U\_\{t\}\(h\)for someh∈γh\\in\\gamma, setΦt\(γ,o′,a\):=\[hu\]∼t\+1\\Phi\_\{t\}\(\\gamma,o^\{\\prime\},a\):=\[hu\]\_\{\\sim\_\{t\+1\}\}\. This does not depend on the representative: ifh,h′′∈γh,h^\{\\prime\\prime\}\\in\\gammaboth haveu∈Ut​\(h\)∩Ut​\(h′′\)u\\in U\_\{t\}\(h\)\\cap U\_\{t\}\(h^\{\\prime\\prime\}\), the last conjunct of Definition[3](https://arxiv.org/html/2609.25757#Thmdefinition3)giveshu∼t\+1h′′uhu\\sim\_\{t\+1\}h^\{\\prime\\prime\}u, so both extensions lie in one class\. HenceCt:=ΓtC\_\{t\}:=\\Gamma\_\{t\},Ft:=ΦtF\_\{t\}:=\\Phi\_\{t\}is a recurrent realization on reachable pairs; by \(F1\)GE,tG\_\{E,t\}is constant on each class, soπ^\(⋅∣o,γ\):=πE\(⋅∣o,g\)\\hat\{\\pi\}\(\\cdot\\mid o,\\gamma\):=\\pi\_\{E\}\(\\cdot\\mid o,g\)has zero distortion andRtmem​\(0\)≤H⁡\(Γt∣Ot\)R^\{\\rm mem\}\_\{t\}\(0\)\\leq H\(\\Gamma\_\{t\}\\mid O\_\{t\}\)\.*Lower bound\.*By the compatibility paragraph of the proof of Theorem[2](https://arxiv.org/html/2609.25757#Thmtheorem2), which does not use transitivity, every\(Ot,Ct\)\(O\_\{t\},C\_\{t\}\)\-cell of a zero\-distortion realization is∼t\\sim\_\{t\}\-compatible; a compatible set is contained in a single class of an equivalence, soΓt\\Gamma\_\{t\}is a function of\(Ot,Ct\)\(O\_\{t\},C\_\{t\}\)andH⁡\(Ct∣Ot\)≥I⁡\(Ct;Γt∣Ot\)=H⁡\(Γt∣Ot\)H\(C\_\{t\}\\mid O\_\{t\}\)\\geq I\(C\_\{t\};\\Gamma\_\{t\}\\mid O\_\{t\}\)=H\(\\Gamma\_\{t\}\\mid O\_\{t\}\)\. Under \(A4\),∼t=∼st\\sim\_\{t\}=\\sim^\{s\}\_\{t\}\(Lemma[1](https://arxiv.org/html/2609.25757#Thmlemma1)\), which recovers Theorem[5](https://arxiv.org/html/2609.25757#Thmtheorem5)\. ∎

### A\.3The general case: closed compatible state assignments

A deterministic recurrent realization assigns every reachable history to one state\. For rate accounting it is without loss of generality to refine that state by the current observation,Ct′=\(Ot,Ct\)C^\{\\prime\}\_\{t\}=\(O\_\{t\},C\_\{t\}\): this leavesH⁡\(Ct′∣Ot\)=H⁡\(Ct∣Ot\)H\(C^\{\\prime\}\_\{t\}\\mid O\_\{t\}\)=H\(C\_\{t\}\\mid O\_\{t\}\)unchanged and yields a partition into joint\(Ot,Ct\)\(O\_\{t\},C\_\{t\}\)\-cells\. We optimize over these partitions rather than overlapping covers; an overlapping cover, as in incompletely specified machines\([Paull and Unger, 1959](https://arxiv.org/html/2609.25757#bib.bib7)\), additionally requires a selection map whose induced partition determines the entropy\.

###### Definition 4\(Closed compatible state assignment\)\.

A sequence of partitions\(𝒫t\)t≤T\(\\mathcal\{P\}\_\{t\}\)\_\{t\\leq T\}of the reachable histories is a closed compatible state assignment if \(i\) every cell of𝒫t\\mathcal\{P\}\_\{t\}is∼t\\sim\_\{t\}\-compatible, and \(ii\) it is closed: ifh,h′h,h^\{\\prime\}lie in one cell of𝒫t−1\\mathcal\{P\}\_\{t\-1\}andu∈Ut−1​\(h\)∩Ut−1​\(h′\)u\\in U\_\{t\-1\}\(h\)\\cap U\_\{t\-1\}\(h^\{\\prime\}\), thenh​uhuandh′​uh^\{\\prime\}ulie in one cell of𝒫t\\mathcal\{P\}\_\{t\}\. Its rate atttisH⁡\(Pt∣Ot\)H\(P\_\{t\}\\mid O\_\{t\}\)under the expert occupancy, wherePtP\_\{t\}is the cell index\. Because compatibility requires equal observations, every cell lies within one observation fiber\.

###### Proposition 2\.

Under \(A1\)–\(A3\), the joint\(Ot,Ct\)\(O\_\{t\},C\_\{t\}\)\-cell partitions of zero\-distortion recurrent realizations are exactly the closed compatible state assignments, up to the rate\-preserving refinement above, and for everytt,Rtmem​\(0\)=infH⁡\(Pt∣Ot\)R^\{\\rm mem\}\_\{t\}\(0\)=\\inf H\(P\_\{t\}\\mid O\_\{t\}\)over closed compatible state assignments\. On a finite POMDP the infimum is a minimum\.

###### Proof\.

\(⇒\\Rightarrow\) Refine a realization toCt′=\(Ot,Ct\)C^\{\\prime\}\_\{t\}=\(O\_\{t\},C\_\{t\}\)\. Its cells are compatible by the compatibility paragraph of the proof of Theorem[2](https://arxiv.org/html/2609.25757#Thmtheorem2); closure follows from the deterministic update because histories in one cell share bothOt−1O\_\{t\-1\}andCt−1C\_\{t\-1\}\. \(⇐\\Leftarrow\) Given a closed compatible assignment, letCtC\_\{t\}be its cell index\. Closure makesFt​\(cell⁡\(h\),o′,a\):=cell⁡\(h​u\)F\_\{t\}\(\\mathrm\{cell\}\(h\),o^\{\\prime\},a\):=\\mathrm\{cell\}\(hu\)well defined on reachable pairs, and compatibility gives, by \(F1\), a single value ofGE,tG\_\{E,t\}per cell, soπ^\(⋅∣o,c\):=πE\(⋅∣o,g\)\\hat\{\\pi\}\(\\cdot\\mid o,c\):=\\pi\_\{E\}\(\\cdot\\mid o,g\)has zero distortion\. Finally,H⁡\(Ct′∣Ot\)=H⁡\(Ct∣Ot\)H\(C^\{\\prime\}\_\{t\}\\mid O\_\{t\}\)=H\(C\_\{t\}\\mid O\_\{t\}\)for the forward construction, while the reverse construction hasCt=PtC\_\{t\}=P\_\{t\}, establishing the rate identity\. ∎

#### Exact computation\.

On an enumerated finite POMDP, a backward dynamic program enumerates partition sequences satisfying compatibility and closure; concentrating its objective on stepttgivesRtmem​\(0\)R^\{\\rm mem\}\_\{t\}\(0\)\. Two exact reductions keep small instances tractable: histories with isomorphic futures are interchangeable, and future\-independent components factorize\. The solver returnsh2​\(1/3\)h\_\{2\}\(1/3\)on the three\-history instance of Remark \(i\), strictly inside\[0,log2⁡3\]\[0,\\log\_\{2\}3\], andH⁡\(Γt∣O¯t\)H\(\\Gamma\_\{t\}\\mid\\bar\{O\}\_\{t\}\)at every step of Task A atM=4M=4\(at most 109 memoized partition types per level, 0\.9 s\) and of A′\. The latter instances are also certified directly by Theorem[1](https://arxiv.org/html/2609.25757#Thmtheorem1)\. The unpruned search grows rapidly with the number of mutually compatible histories with non\-isomorphic futures \(more than10710^\{7\}partition types per level at Task AM=16M=16\), so it is a certificate for small instances rather than a general algorithm\. The drifting corridor of Appendix[C\.4](https://arxiv.org/html/2609.25757#A3.SS4), our non\-transitive family, already has up to 384 histories per level atW=3W=3with maximal compatible sets of up to 96 histories and hundreds of thousands of non\-transitive triples; exact search is impractical \(no termination within 25 minutes\); there the minimum is instead pinned by the matching bounds below\. The general problem is a recurrent analogue of zero\-error source coding with decoder side information\([Witsenhausen, 1976](https://arxiv.org/html/2609.25757#bib.bib9);[Alon and Orlitsky, 1996](https://arxiv.org/html/2609.25757#bib.bib10)\); we do not characterize its computational complexity\.

#### A certified upper bound beyond the reach of the exact DP\.

On the drifting corridor \(W≥3W\\geq 3\), any closed compatible assignment is realisable, so a constructed one certifies an upper bound\. A greedy construction \(cells merged only when pairwise compatible, every merge propagated to the successors that share a continuation, chains rolled back when an implied merge is incompatible; verified from scratch for compatibility and successor consistency\) attains exactly theW=1W=1ladder,22bits in the first hall and11bit in the second, forW=3W=3,55and77\(384384to896896histories per level\), and forW=9,11,13W=9,11,13with largest\-cells\-first ordering\. The hall\-wise brackets therefore tighten from\[0,2\+log2⁡W\]\[0,\\,2\+\\log\_\{2\}W\]to\[0,2\]\[0,2\]and\[0,1\]\[0,1\]: the drift, re\-revealed by every hall observation, need not be stored\. Sufficient learned codes match the22\-bit construction in the first hall and retain a modest0\.20\.2–0\.40\.4\-bit surplus in the second\.

###### Proposition 3\(Incompatibility entropy lower bound\)\.

Within each observation fiberooat steptt, connect two reachable histories when they are incompatible, and weight vertices byP⁡\(Ht=h∣Ot=o\)P\(H\_\{t\}=h\\mid O\_\{t\}=o\)\. LetHχ​\(t,o\)H\_\{\\chi\}\(t,o\)be the minimum entropy of a proper coloring of this weighted graph\. Under \(A1\)–\(A3\),

Rtmem​\(0\)≥∑oP⁡\(Ot=o\)​Hχ​\(t,o\)\.R\_\{t\}^\{\\rm mem\}\(0\)\\geq\\sum\_\{o\}P\(O\_\{t\}=o\)H\_\{\\chi\}\(t,o\)\.Equality is certified whenever a closed compatible state assignment attains this lower bound\.

###### Proof\.

Every joint\(Ot,Ct\)\(O\_\{t\},C\_\{t\}\)cell of a zero\-distortion realization is compatible \(Theorem[2](https://arxiv.org/html/2609.25757#Thmtheorem2)\), so its state labels form a proper coloring within each observation fiber\. Minimizing over all proper colorings relaxes successor consistency and hence lower\-bounds the conditional entropy of every realization\. A closed compatible assignment supplies the matching realizable upper bound \(Proposition[2](https://arxiv.org/html/2609.25757#Thmproposition2)\)\. ∎

#### Finite\-instance equality on the corridor\.

Dropping closure leaves a necessary condition: histories with the sameOtO\_\{t\}that are pairwise*incompatible*must occupy different memory states, so within each observation cell a zero\-distortion memory is a proper coloring of the incompatibility graph, and the minimum entropy over proper colorings, weighted by the cell probabilities, lower\-boundsRtmem​\(0\)R^\{\\rm mem\}\_\{t\}\(0\)\. Histories with identical incompatibility neighbourhoods can be merged without loss \(moving a vertex from the smaller color class to the larger preserves proper coloring and majorizes the mass vector\), after which every cell has at most1212types and the minimum\-entropy coloring is computed exactly by a subset dynamic program, cross\-checked by enumerating all proper colorings\. ForW=3,5,7W=3,5,7the bound equals the certified upper bound at every step \(00,11,22and11bits along the episode\), so the minima of these finite instances are pinned exactly at theW=1W=1ladder\. A one\-line certificate suffices: no pairwise\-compatible set of histories has conditional mass aboveα=1/4\\alpha=1/4in the first hall or1/21/2in the second \(the type graph is a disjoint union ofK4K\_\{4\}’s, respectivelyK2K\_\{2\}’s\), henceH≥log2⁡\(1/α\)H\\geq\\log\_\{2\}\(1/\\alpha\); the plain clique bound is loose \(at most1\.871\.87of the22bits\)\. The same agreement holds at every step forW=9W=9,1111and1313\(up to16641664histories per level; exact coloring on every cell\), although these instances belong to one structural family \(identical type graphs\) and only the largest\-cells\-first construction attains the bound\. The lower bound is solved exactly; the upper bound is an independently verified constructive certificate\. On the three\-history instance of Remark \(i\) the coloring bound equals the exact dynamic program,h2​\(1/3\)=0\.918h\_\{2\}\(1/3\)=0\.918bit, whereas the largest\-compatible\-set certificate gives onlylog2⁡\(3/2\)=0\.585\\log\_\{2\}\(3/2\)=0\.585, so the coloring bound is the one to use in general\. These are computations on finite instances, not a result for the general non\-transitive case, whose complexity we do not characterize\. State\-count minimization of incompletely specified machines is NP\-hard\([Pfleeger, 1973](https://arxiv.org/html/2609.25757#bib.bib8)\); this related result is not a complexity proof for our entropy objective\.

### A\.4Nonzero distortion

LetRtmem​\(D\)R^\{\\rm mem\}\_\{t\}\(D\)be the infimum ofH⁡\(Ct∣Ot\)H\(C\_\{t\}\\mid O\_\{t\}\)over deterministic recurrent realizations whose expected distortion is at mostDDat every step\. The instantaneous functionRE​\(D\)R\_\{E\}\(D\)is defined in §[2](https://arxiv.org/html/2609.25757#S2)for the source at steptt\.

###### Proposition 4\.

Under \(A1\)–\(A3\),Rtmem​\(D\)R^\{\\rm mem\}\_\{t\}\(D\)is non\-increasing inDDand

Rtmem​\(D\)≥RE​\(D\)=RG\|O​\(D\)for every​D≥0\.R^\{\\rm mem\}\_\{t\}\(D\)\\geq R\_\{E\}\(D\)=R\_\{G\\mid O\}\(D\)\\qquad\\text\{for every \}D\\geq 0\.The inequality can be strict, including atD=0D=0\.

###### Proof\.

The feasible sets are nested inDD, which gives monotonicity\. Any deterministic recurrent realization induces at stepttan admissible instantaneous encoderCt=ct​\(Ht\)C\_\{t\}=c\_\{t\}\(H\_\{t\}\)and decoderπ^\(⋅∣Ot,Ct\)\\hat\{\\pi\}\(\\cdot\\mid O\_\{t\},C\_\{t\}\)\. BecauseCtC\_\{t\}is a function ofHtH\_\{t\},H⁡\(Ct∣Ot\)=I⁡\(Ct;Ht∣Ot\)H\(C\_\{t\}\\mid O\_\{t\}\)=I\(C\_\{t\};H\_\{t\}\\mid O\_\{t\}\); taking the infimum over the more restricted recurrent class therefore givesRtmem​\(D\)≥RE​\(D\)R^\{\\rm mem\}\_\{t\}\(D\)\\geq R\_\{E\}\(D\)\. Theorem[4](https://arxiv.org/html/2609.25757#Thmtheorem4)gives the equality on the right\. At zero distortion the lower bound isH⁡\(GE,t∣Ot\)H\(G\_\{E,t\}\\mid O\_\{t\}\), whereas Theorem[1](https://arxiv.org/html/2609.25757#Thmtheorem1)givesH⁡\(Γt∣Ot\)H\(\\Gamma\_\{t\}\\mid O\_\{t\}\)in the transitive regime; A′is strict in both gaps\. ∎

This proposition supplies only a lower bound forD\>0D\>0\. The recurrent constraint couples encoders across steps, and deterministic finite realizations need not admit time\-sharing, so we claim neither convexity nor equality with the instantaneous frontier\.

### A\.5Behavioral future sufficiency

ForJ≥0J\\geq 0define∼\(J\)t\\sim^\{\(J\)\}\_\{t\}by unrolling Definition[3](https://arxiv.org/html/2609.25757#Thmdefinition3)JJsteps:h∼\(0\)th′h\\sim^\{\(0\)\}\_\{t\}h^\{\\prime\}iff\(Ot,GE,t\)​\(h\)=\(Ot,GE,t\)​\(h′\)\(O\_\{t\},G\_\{E,t\}\)\(h\)=\(O\_\{t\},G\_\{E,t\}\)\(h^\{\\prime\}\); forJ≥1J\\geq 1andt<Tt<T,h∼\(J\)th′h\\sim^\{\(J\)\}\_\{t\}h^\{\\prime\}iffh∼\(0\)th′h\\sim^\{\(0\)\}\_\{t\}h^\{\\prime\}andhu∼\(J−1\)t\+1h′uhu\\sim^\{\(J\-1\)\}\_\{t\+1\}h^\{\\prime\}ufor allu∈Ut​\(h\)∩Ut​\(h′\)u\\in U\_\{t\}\(h\)\\cap U\_\{t\}\(h^\{\\prime\}\); and∼\(J\)T:=∼\(0\)T\\sim^\{\(J\)\}\_\{T\}:=\\sim^\{\(0\)\}\_\{T\}\. Then∼\(T−t\)t=∼t\\sim^\{\(T\-t\)\}\_\{t\}=\\sim\_\{t\}and∼\(J\+1\)t⊆∼\(J\)t\\sim^\{\(J\+1\)\}\_\{t\}\\subseteq\\sim^\{\(J\)\}\_\{t\}, since each unrolling adds conjuncts\. Writeξt:j=\(Ot\+1:t\+j,At:t\+j−1\)\\xi\_\{t:j\}=\(O\_\{t\+1:t\+j\},A\_\{t:t\+j\-1\}\)for a continuation of lengthjj, so that\(Ht,ξt:j\)=Ht\+j\(H\_\{t\},\\xi\_\{t:j\}\)=H\_\{t\+j\}\. A continuation is reachable fromhhiff each of its steps lies in the support of the history built so far, so the continuations reachable from bothhhandh′h^\{\\prime\}are exactly those built step by step from common supports, and unrolling the definition gives

h∼\(J\)th′⇔Ot\(h\)=Ot\(h′\)andPE\(At\+j∣h,ξ\)=PE\(At\+j∣h′,ξ\)for all0≤j≤Jand allξ=ξt:jreachable from both,h\\sim^\{\(J\)\}\_\{t\}h^\{\\prime\}\\iff O\_\{t\}\(h\)=O\_\{t\}\(h^\{\\prime\}\)\\ \\text\{and\}\\ P\_\{E\}\(A\_\{t\+j\}\\mid h,\\xi\)=P\_\{E\}\(A\_\{t\+j\}\\mid h^\{\\prime\},\\xi\)\\\\ \\text\{for all \}0\\leq j\\leq J\\text\{ and all \}\\xi=\\xi\_\{t:j\}\\text\{ reachable from both\},\(5\)because under \(A2\)PE\(At\+j∣Ht\+j\)=πE\(⋅∣Ot\+j,GE,t\+j\)P\_\{E\}\(A\_\{t\+j\}\\mid H\_\{t\+j\}\)=\\pi\_\{E\}\(\\cdot\\mid O\_\{t\+j\},G\_\{E,t\+j\}\)determinesGE,t\+jG\_\{E,t\+j\}within the fiber ofOt\+jO\_\{t\+j\}, andOt\+jO\_\{t\+j\}is part ofξ\\xiforj≥1j\\geq 1\.

###### Lemma 3\.

Under \(A4\) every∼\(J\)t\\sim^\{\(J\)\}\_\{t\}is an equivalence;Γt\(J\):=\[Ht\]∼\(J\)t\\Gamma^\{\(J\)\}\_\{t\}:=\[H\_\{t\}\]\_\{\\sim^\{\(J\)\}\_\{t\}\}satisfiesΓt\(0\)⪯Γt\(1\)⪯⋯⪯Γt\(T−t\)=Γt\\Gamma^\{\(0\)\}\_\{t\}\\preceq\\Gamma^\{\(1\)\}\_\{t\}\\preceq\\dots\\preceq\\Gamma^\{\(T\-t\)\}\_\{t\}=\\Gamma\_\{t\}, each partition refining the previous one, andH⁡\(Γt\(J\)∣Ot\)H\(\\Gamma^\{\(J\)\}\_\{t\}\\mid O\_\{t\}\)is non\-decreasing inJJ\.

###### Proof\.

Induction onJJ, for allttsimultaneously:∼\(0\)t\\sim^\{\(0\)\}\_\{t\}is equality of a function\. If∼\(J−1\)t\+1\\sim^\{\(J\-1\)\}\_\{t\+1\}is an equivalence andh1∼\(J\)th2∼\(J\)th3h\_\{1\}\\sim^\{\(J\)\}\_\{t\}h\_\{2\}\\sim^\{\(J\)\}\_\{t\}h\_\{3\}, the three share\(Ot,GE,t\)\(O\_\{t\},G\_\{E,t\}\), \(A4\) equalizes their supports, and transitivity at\(J−1,t\+1\)\(J\-1,t\+1\)givesh1u∼\(J−1\)t\+1h3uh\_\{1\}u\\sim^\{\(J\-1\)\}\_\{t\+1\}h\_\{3\}ufor all commonuu, i\.e\.h1∼\(J\)th3h\_\{1\}\\sim^\{\(J\)\}\_\{t\}h\_\{3\}\. Refinement is∼\(J\+1\)t⊆∼\(J\)t\\sim^\{\(J\+1\)\}\_\{t\}\\subseteq\\sim^\{\(J\)\}\_\{t\}; since the coarser partition is then a function of the finer one,H⁡\(Γt\(J\)∣Ot\)≤H⁡\(Γt\(J\+1\)∣Ot\)H\(\\Gamma^\{\(J\)\}\_\{t\}\\mid O\_\{t\}\)\\leq H\(\\Gamma^\{\(J\+1\)\}\_\{t\}\\mid O\_\{t\}\)\. ∎

The population BFS distortion at horizonJJof a codeCt=ct​\(Ht\)C\_\{t\}=c\_\{t\}\(H\_\{t\}\)with decodersqj\(⋅∣c,o,ξ\)q\_\{j\}\(\\cdot\\mid c,o,\\xi\),j=0,…,Jj=0,\\dots,J, is

DBFS=∑j=0Jwj𝔼\[δ\(PE\(At\+j∣Ht,ξt:j\),qj\(⋅∣Ct,Ot,ξt:j\)\)\],wj\>0,D\_\{\\rm BFS\}=\\sum\_\{j=0\}^\{J\}w\_\{j\}\\,\\mathbb\{E\}\\Big\[\\delta\\big\(P\_\{E\}\(A\_\{t\+j\}\\mid H\_\{t\},\\xi\_\{t:j\}\),\\ q\_\{j\}\(\\cdot\\mid C\_\{t\},O\_\{t\},\\xi\_\{t:j\}\)\\big\)\\Big\],\\qquad w\_\{j\}\>0,with the expectation over\(Ht,ξt:j\)\(H\_\{t\},\\xi\_\{t:j\}\)under the expert occupancy, so that every continuation reachable from a history receives positive weight \(common reachable coverage, \(ii\)\);δ\\deltais distributional,δ=0\\delta=0iff the distributions coincide \(\(iii\); for a stochastic expert this is zero KL or zero excess log\-loss, not zero log\-loss\); all prefixesj=0,…,Jj=0,\\dots,Jare predicted \(\(iv\)\); the full horizon isJ=T−tJ=T\-t\(\(i\)\)\. Thej=0j=0term is the imitation loss; the targetAt\+jA\_\{t\+j\}is never an input\.

###### Proposition 5\(BFS minimizers\)\.

Under \(A1\)–\(A3\), full\-horizon prediction over all common reachable continuations with a distributional loss,DBFS=0D\_\{\\rm BFS\}=0iff every\(Ot,Ct\)\(O\_\{t\},C\_\{t\}\)\-cell is∼\(J\)t\\sim^\{\(J\)\}\_\{t\}\-compatible; under \(A4\), iff\(Ot,Ct\)\(O\_\{t\},C\_\{t\}\)refinesΓt\(J\)\\Gamma^\{\(J\)\}\_\{t\}, and withJ=T−tJ=T\-tthe rate\-minimal zero\-BFS code hasH⁡\(Ct∣Ot\)=H⁡\(Γt∣Ot\)H\(C\_\{t\}\\mid O\_\{t\}\)=H\(\\Gamma\_\{t\}\\mid O\_\{t\}\)\.

###### Proof\.

\(⇒\\Rightarrow\) LetDBFS=0D\_\{\\rm BFS\}=0and leth,h′h,h^\{\\prime\}lie in the same cell\(o,c\)\(o,c\)\. For any0≤j≤J0\\leq j\\leq Jand anyξ=ξt:j\\xi=\\xi\_\{t:j\}reachable from both, the pairs\(h,ξ\)\(h,\\xi\)and\(h′,ξ\)\(h^\{\\prime\},\\xi\)have positive probability, so by \(iii\)qj\(⋅∣c,o,ξ\)=PE\(At\+j∣h,ξ\)q\_\{j\}\(\\cdot\\mid c,o,\\xi\)=P\_\{E\}\(A\_\{t\+j\}\\mid h,\\xi\)andqj\(⋅∣c,o,ξ\)=PE\(At\+j∣h′,ξ\)q\_\{j\}\(\\cdot\\mid c,o,\\xi\)=P\_\{E\}\(A\_\{t\+j\}\\mid h^\{\\prime\},\\xi\); the two expert conditionals agree, and equation[5](https://arxiv.org/html/2609.25757#A1.E5)givesh∼\(J\)th′h\\sim^\{\(J\)\}\_\{t\}h^\{\\prime\}\. \(⇐\\Leftarrow\) If every cell is∼\(J\)t\\sim^\{\(J\)\}\_\{t\}\-compatible, setqj\(⋅∣c,o,ξ\):=PE\(At\+j∣h,ξ\)q\_\{j\}\(\\cdot\\mid c,o,\\xi\):=P\_\{E\}\(A\_\{t\+j\}\\mid h,\\xi\)for anyhhin the cell\(o,c\)\(o,c\)from whichξ\\xiis reachable; by equation[5](https://arxiv.org/html/2609.25757#A1.E5)any two suchhhgive the same value \(pairwise compatibility suffices; no transitivity is used\), and these decoders haveDBFS=0D\_\{\\rm BFS\}=0\. Under \(A4\),∼\(J\)t\\sim^\{\(J\)\}\_\{t\}is an equivalence \(Lemma[3](https://arxiv.org/html/2609.25757#Thmlemma3)\), so a cell is compatible iff it is contained in a class, i\.e\. iff\(Ot,Ct\)\(O\_\{t\},C\_\{t\}\)refinesΓt\(J\)\\Gamma^\{\(J\)\}\_\{t\}\. ForJ=T−tJ=T\-t,Γt\(J\)=Γt\\Gamma^\{\(J\)\}\_\{t\}=\\Gamma\_\{t\}; a code refiningΓt\\Gamma\_\{t\}hasΓt\\Gamma\_\{t\}as a function of\(Ot,Ct\)\(O\_\{t\},C\_\{t\}\)and thereforeH⁡\(Ct∣Ot\)≥H⁡\(Γt∣Ot\)H\(C\_\{t\}\\mid O\_\{t\}\)\\geq H\(\\Gamma\_\{t\}\\mid O\_\{t\}\), with equality forCt=ΓtC\_\{t\}=\\Gamma\_\{t\}\(for stochastic codes the same holds withI⁡\(Ct;Ht∣Ot\)I\(C\_\{t\};H\_\{t\}\\mid O\_\{t\}\)in place ofH⁡\(Ct∣Ot\)H\(C\_\{t\}\\mid O\_\{t\}\)\)\. ∎

#### Why every prefix is predicted \(first divergence\)\.

Suppose that the objective included only thej=Jj=Jterm\. Leth,h′h,h^\{\\prime\}share a cell but be∼\(J\)t\\sim^\{\(J\)\}\_\{t\}\-incompatible, and letj⋆≤Jj^\{\\star\}\\leq Jbe the smallestjjfor which some common continuationξt:j⋆\\xi\_\{t:j^\{\\star\}\}yields different expert conditionals\. By the minimality ofj⋆j^\{\\star\}, the action prefixAt:t\+j⋆−1A\_\{t:t\+j^\{\\star\}\-1\}supplied alongξt:j⋆\\xi\_\{t:j^\{\\star\}\}is common tohhandh′h^\{\\prime\}, so thej⋆j^\{\\star\}term is positive\. Every longer continuation, however, containsAt\+j⋆A\_\{t\+j^\{\\star\}\}, whose supports may already differ betweenhhandh′h^\{\\prime\}, as they can for a deterministic expert\. In that case, noξt:J\\xi\_\{t:J\}is common, and theJJterm alone can vanish on an incompatible cell\. Predicting every prefix, as required by \(iv\), assigns a cost at the first step at which the supplied actions have not already revealed the divergence\. Supplying intermediate expert actions under teacher forcing is therefore not label leakage: these actions instantiate the same continuation consumed by the recurrent transition, and the prediction target is never provided as input\. With finite demonstrations, finiteJJ, and sampled continuations, the trained objective is an empirical surrogate forDBFSD\_\{\\rm BFS\}; whenJJis finite, its target isΓt\(J\)\\Gamma^\{\(J\)\}\_\{t\}rather thanΓt\\Gamma\_\{t\}\.

### A\.6Closed\-form verification of the theoretical quantities

Before running any learning experiment, we validated every partition\-refinement quantity in Appendix[B\.1](https://arxiv.org/html/2609.25757#A2.SS1)against closed\-form results on the finite toy problems \(the T0 gate\)\. First, Blahut–Arimoto on the conditional source\(GE,t,Ot\)\(G\_\{E,t\},O\_\{t\}\)agrees withH⁡\(G∣O\)H\(G\\mid O\)atD=0D=0to×10−167\.8\\\!\\times\\\!10^\{\-16\}\. Second, forM∈\{4,8,16\}M\\in\\\{4,8,16\\\}, Blahut–Arimoto on the full history source\(Ht,Ot\)\(H\_\{t\},O\_\{t\}\)agrees with the reduced source to×10−151\.3\\\!\\times\\\!10^\{\-15\}, as predicted by Theorem[4](https://arxiv.org/html/2609.25757#Thmtheorem4)\. Third, the solver recovers the deterministic frontierRG\|Odet​\(D\)R^\{\\det\}\_\{G\\mid O\}\(D\)shown in Figure[4](https://arxiv.org/html/2609.25757#A3.F4)\. Fourth, at everytt,H⁡\(Γt∣Ot\)H\(\\Gamma\_\{t\}\\mid O\_\{t\}\)matches the hand\-derived toy staircases:\[0,0,1,2\]\[0,0,1,2\]for the reveal toy andlog2⁡\(R1​R2\)→log2⁡R2→0\\log\_\{2\}\(R\_\{1\}R\_\{2\}\)\\to\\log\_\{2\}R\_\{2\}\\to 0for the gap toy\. Fifth, partition refinement returns the strict bracket on a re\-reveal instance \(Γs\\Gamma^\{s\}requires 2 bits whereΓ\\Gammarequires 1\), refuses to constructΓt\\Gamma\_\{t\}on a purpose\-built non\-transitive instance, and produces aΓt\(J\)\\Gamma^\{\(J\)\}\_\{t\}table monotone inJJ\(Lemma[3](https://arxiv.org/html/2609.25757#Thmlemma3)\)\. This refinement pass takes below one second for every dataset, with at most 768 histories per level in the corridor; the general exact dynamic program has the more limited scope described in Appendix[A\.3](https://arxiv.org/html/2609.25757#A1.SS3)\.

## Appendix BExperimental design and measurement

These sections fix the information structure before comparing learned rates\. They distinguish symbolic behavioral requirements, code entropy, codebook capacity, and physical storage, and document the empirical sufficiency gate\.

### B\.1Benchmarks and the symbolic observation convention

Table 3:Benchmark families, regimes and frozen primary configurations \(NN= training episodes;K=16K=16codes throughout; the Franka runs use batch 512\)\.Observation convention and frozen binner\.For behaviorally decodable classes, defineO¯t:=ft​\(Ot\)=\(phase,probe symbol,ft1​\(Ot\),ft2​\(Ot\)\)\\bar\{O\}\_\{t\}:=f\_\{t\}\(O\_\{t\}\)=\(\\text\{phase\},\\ \\text\{probe symbol\},\\ f^\{1\}\_\{t\}\(O\_\{t\}\),\\ f^\{2\}\_\{t\}\(O\_\{t\}\)\)\. Here,ftif^\{i\}\_\{t\}is a frozen nearest\-class\-mean binner of the raw end\-effector position at steptt, with class means fitted once on half of the recorded episodes\. The binner is applied only when the classes are separable, defined as class\-conditional means more than55cm apart, or more than 25 standard deviations of the sensor noise; otherwise, it returns a null symbol\. On every recorded dataset used in the paper, including A′at gaps 6, 10, and 20, the 2048\-episode set, and Task A, the frozen binner reproduces the class on100%100\\%of held\-out episodes at each separable step\. Specifically, there are00disagreements among2,8162\{,\}816–7,0407\{,\}040held\-out \(episode, step\) pairs per dataset, both with and without conditioning on the other class\. Except for the Task A sag convention disclosed below, the solver labels are therefore exactly those produced byft​\(Ot\)f\_\{t\}\(O\_\{t\}\), and this binned observation is a coarsening ofOtO\_\{t\}\. Under this convention, the data\-processing inequality givesH⁡\(Ct∣O¯t\)≥H⁡\(Ct∣Ot\)H\(C\_\{t\}\\mid\\bar\{O\}\_\{t\}\)\\geq H\(C\_\{t\}\\mid O\_\{t\}\), so the reported rates upper\-bound the raw\-observation rate\. Table[5](https://arxiv.org/html/2609.25757#A2.T5)reports a complementary held\-out probe: a classifier on raw\-observation windows recovers a class only when the convention marks it as separable \(0\.990\.99–1\.001\.00\) and otherwise performs at chance\.

Table 4:Design choices, premises and evidence \(§[4](https://arxiv.org/html/2609.25757#S4)\)\.The sag symbol\.On Task A the symbolic observation carries a mass\-half symbol during transport and placing\. The online binner of Appendix[B\.4](https://arxiv.org/html/2609.25757#A2.SS4)reproduces it on99\.5%99\.5\\%of transport steps but only80%80\\%of placing steps, so the placing convention credits the observation with a partly decodable symbol\. This symbol is behaviorally irrelevant and no reported quantity at the place step involves mass; we retain the convention and state the discrepancy explicitly\.

Benchmark hygiene\([Tao et al\., 2025](https://arxiv.org/html/2609.25757#bib.bib48);[Agarwal et al\., 2026](https://arxiv.org/html/2609.25757#bib.bib35), cf\.\)\.Without sensor noise, sub\-millimeter contact artifacts made the mass decodable at7979–93%93\\%\(2 mm noise and a non\-contact attach removed it\); a capture offset in the first pixel collection produced apparently memory\-free solutions that the sufficiency metrics flagged \(data recollected\); and a deterministic expert makes the inverse\-dynamics objective degenerate in the nuisance, which motivated the corridor’s behaviorally equivalent action randomness\.

Coverage of the empirical POMDP\.The solver is exact on the finite POMDP induced by the recorded symbolic histories\. On 13 of the 14 datasets every latent configuration of the generating program is recorded \(Good–Turing missing mass00, no unseen key\), and the reachable histories per level number at most 8 \(A′\), 32 \(tier 16\), 64 \(P\-I\) and 252 \(Task A\)\. The exception is Task A atM=32M=32, where 2 of 128 mass–slot pairs are unrecorded \(Good–Turing missing mass0\.020\.02\)\. Their symbolic sequences are reconstructed from the generating program’s binary mass code and mass\-half sag rule; the same reconstruction reproduces every observed key when applied from a different donor\. Under this reconstructionH⁡\(Γs∣O¯\)=1\.00H\(\\Gamma^\{s\}\\mid\\bar\{O\}\)=1\.00bit, as at every otherMM, and no recorded label is modified\.

Estimator bias\.All entropies are plug\-in estimates over the recorded episodes with exact occupancy weights\. We compute the Miller–Madow correction toH⁡\(C∣O¯\)H\(C\\mid\\bar\{O\}\)from the number of*occupied*\(c,o¯\)\(c,\\bar\{o\}\)cells at each step\. The correction is at most0\.0090\.009bit for the 2048\-episode set across 16 runs and all steps, at most0\.0350\.035bit for the 512\-episode sets, and at most0\.0170\.017bit for the 128 own\-occupancy episodes of the reference policy\. The worst\-case correction of4⋅15/\(2⋅512​ln⁡2\)=0\.0854\\cdot 15/\(2\\cdot 512\\ln 2\)=0\.085bit is never approached because each symbol occupies at most 2–4 codes\. These corrections assess finite\-sample bias in the reported rates; matching2\.002\.00bits is interpreted jointly with sufficiency and the replication checks\. We do*not*remove self\-controlled observation components, such as the end\-effector pose, fromOtO\_\{t\}; instead, the body\-memory probe detects when they carry memory, as discussed in §[3](https://arxiv.org/html/2609.25757#S3)\.

Table 5:Leak probe on A′\(tier 4, gap 6\): held\-out accuracy \(n=128n=128episodes\) of a classifier on raw observation windows, per phase and hidden variable; chance in parentheses\. Bold entries are the steps at which the visibility convention marks the class visible\.The solver applies backward partition refinement from Definition[3](https://arxiv.org/html/2609.25757#Thmdefinition3)to an enumerated finite POMDP constructed from the recorded data\. We verify that the mapping from latent state to symbolic\-observation and action sequences is deterministic\. The solver computes∼t\\sim\_\{t\}and∼st\\sim^\{s\}\_\{t\}, performs per\-step checks of \(A2\), \(A4\), and transitivity, refuses to constructΓt\\Gamma\_\{t\}when the relation is non\-transitive, and returns exact conditional entropies\. The symbolic observation is\(phase,probe bit,visible classes\)\(\\text\{phase\},\\text\{probe bit\},\\text\{visible classes\}\)\. A class is defined as visible atttwhen it can be decoded from the raw observation, using class\-conditional end\-effector separation\>5\>5cm after conditioning on the other factor\. Without this convention, information already present in the raw observation would be incorrectly counted as memory\. For Task A, the symbolic observation also contains a mass\-half symbol during the sag phases\.

### B\.2Realization and training variants

The residual proposal ise~t=et−1\+f⁡\(et−1,go​\(Ot\),ga​\(At−1\)\)\\tilde\{e\}\_\{t\}=e\_\{t\-1\}\+f\(e\_\{t\-1\},g\_\{o\}\(O\_\{t\}\),g\_\{a\}\(A\_\{t\-1\}\)\), followed by deterministic nearest\-code quantization\. Quantization is hard at training and evaluation; gradients use straight\-through and soft\-to\-hard rate estimators\. Unless stated otherwise,K=16K=16, the code vector has dimension3232, and the hidden width is128128\. The code is the only learned recurrent state transmitted across steps\. In the matched\-side\-information controls, the transition and prior instead receive the frozen per\-step binnerO¯t=ft​\(Ot\)\\bar\{O\}\_\{t\}=f\_\{t\}\(O\_\{t\}\)\.

The tables use compact variant labels:−\-R is the plain imitation\-and\-rate objective;−\-RF adds behavioral future sufficiency \(BFS\); scaffold uses an auxiliary continuous recurrent path during training, annealed to zero before evaluation\. A bypass retains that continuous path and is therefore excluded from sole\-carrier minimality claims\. DIACRITIC names this implementation, not a separate architectural contribution\.

BFS predicts every prefix of a future continuation from a frozen code and the intervening observations and actions\. Its ideal full\-horizon minimizers have the compatibility property of Proposition[5](https://arxiv.org/html/2609.25757#Thmproposition5); the finite sampled objective is only a surrogate\. Event\-agnostic future\-behavior supervision instead predicts a random\-offset action without future observations \(§[5](https://arxiv.org/html/2609.25757#S5)\)\. These objectives differ precisely in decoder side information\. Conditioning the prior matches the conditional\-rate objective, although the conditional and unconditional priors are empirically indistinguishable in the reported families\.

#### Additional attribution evidence\.

With a persistent GRU bypass, readout\-code rates are0\.90\.9–1\.91\.9bits on the toy,1\.061\.06on A′, and1\.521\.52in the corridor despite a two\-bit behavioral requirement\. One closed\-loop policy instead externalizes the placement class in its pose: end\-effector separation grows from0\.040\.04to2\.072\.07cm during transport \(p=×10−14p=5\\\!\\times\\\!10^\{\-14\}\) at roughly22cm imitation error\. Expert\- and policy\-occupancy sufficiency agree on31/3231/32reference policies, but the additional domains demonstrate that this agreement is not universal\. These controls motivate the attribution requirement in §[3](https://arxiv.org/html/2609.25757#S3)\.

### B\.3Rate accounting

The quantity∑tH⁡\(Ct∣Ct−1,Ot\)\\sum\_\{t\}H\(C\_\{t\}\\mid C\_\{t\-1\},O\_\{t\}\)assigns a cost of 0 both to a code that stores only behavioral information and to one that additionally stores nuisance information, because past observations can pass throughCt−1C\_\{t\-1\}without cost\. Conversely,∑tH⁡\(Ct∣Ct−1\)\\sum\_\{t\}H\(C\_\{t\}\\mid C\_\{t\-1\}\)over\-penalizes information that remains permanently visible\. At zero distortion, the per\-step quantityH⁡\(Ct∣Ot\)H\(C\_\{t\}\\mid O\_\{t\}\)is characterized by Theorems[1](https://arxiv.org/html/2609.25757#Thmtheorem1)and[2](https://arxiv.org/html/2609.25757#Thmtheorem2)and assigns a cost when information must persist beyond its visibility\. Throughout the paper, “pay” refers to representational burden rather than communication cost\.

### B\.4Side\-information and estimator controls

#### Resolution sweep \(N1\)\.

The reported rates condition onO¯t=ft​\(Ot\)\\bar\{O\}\_\{t\}=f\_\{t\}\(O\_\{t\}\)\. To test whether the code carries raw\-observation detail thatO¯t\\bar\{O\}\_\{t\}omits, we refine the conditioning by the raw end\-effector position on a grid of decreasing cell size and recomputeH^​\(Ct∣fr​\(Ot\)\)\\hat\{H\}\(C\_\{t\}\\mid f\_\{r\}\(O\_\{t\}\)\)for the 16 plain−\-R runs of Figure[7](https://arxiv.org/html/2609.25757#A3.F7)\(11 sufficient\)\. In the first gap the sufficient seeds give2\.0102\.010bits atO¯t\\bar\{O\}\_\{t\}and2\.0042\.004,2\.0042\.004,1\.9831\.983,1\.9061\.906,1\.5501\.550bits with cells of1010,55,22,11,0\.50\.5cm\. The decrease is fragmentation: permuting the codes within eachO¯t\\bar\{O\}\_\{t\}cell, which preservesH⁡\(C∣O¯\)H\(C\\mid\\bar\{O\}\)exactly and destroys any dependence on position, gives the same decrease to within0\.0080\.008bit at every resolution \(at0\.50\.5cm there are 237 occupied cells with 5\.4 episodes each\)\. The Miller–Madow values range from2\.0122\.012to1\.7621\.762, and a two\-fold cross\-validated logistic model of the code from the raw 27\-dimensional observation andO¯t\\bar\{O\}\_\{t\}gives2\.0672\.067bits, an upper bound onH⁡\(Ct∣Ot\)H\(C\_\{t\}\\mid O\_\{t\}\)that does not fragment\. Thus neither diagnostic provides evidence that raw observation predicts the first\-gap code beyondO¯t\\bar\{O\}\_\{t\}\. In the second gap the sufficient seeds carry1\.521\.52bits, of which0\.040\.04–0\.080\.08bit exceeds the permutation null and is explained by the end\-effector position of four low\-β\\betaseeds; the cross\-validated bound is1\.481\.48bits, which we quote wherever a leak\-resistant post\-use rate is needed\.

#### Estimator bias \(N2\)\.

RecomputingSΓS\_\{\\Gamma\}with Miller–Madow corrections applied consistently to the joint and marginal entropies changes it by at most0\.0140\.014across 160 runs, and the number of sufficient seeds is identical at thresholds0\.80\.8,0\.90\.9and0\.950\.95in every cell: the scores are bimodal, with every sufficient run atSΓ=1\.000S\_\{\\Gamma\}=1\.000on every second\-gap step and the best insufficient run at0\.370\.37\.

#### Matched side information \(E1\)\.

The pre\-specified control feeds the recurrent transition and the conditional prior a one\-hot ofO¯t=ft​\(Ot\)\\bar\{O\}\_\{t\}=f\_\{t\}\(O\_\{t\}\)computed online by a frozen per\-step binner \(phase from the observation, probe symbol from the probe channels, visible classes by nearest per\-step class mean of the end\-effector position, mass half by a per\-step height threshold on Task A\); the binner agrees with the symbolic labels on every recorded step of A′and on every Task A step except the mass\-half symbol, which it reproduces on99\.5%99\.5\\%of transport steps but only80%80\\%of placing steps, where the end\-effector height no longer separates the masses \(the symbol is behaviorally irrelevant, and no reported quantity at the place step involves the mass\); no unseen symbol occurred in closed loop\. Results \(8 seeds per cell; 128 closed\-loop episodes per seed\): on A′\(2048\-episode set,N=1152N=1152\)6/86/8seeds are sufficient atβ=0\\beta=0and5/85/8at10−310^\{\-3\}, with first\-gap rate2\.002\.00on the sufficient seeds and mean closed\-loop success0\.840\.84and0\.920\.92\(raw input:6/86/8,5/85/8,0\.960\.96,0\.970\.97\); on Task A all four values ofMMgive8/88/8, grasp rate0\.000\.00–0\.100\.10bit, and success0\.970\.97–1\.001\.00\.

Two stronger variants also restrict the behavioral head: a hierarchical policy whose symbolic\-action head reads\(Ct,O¯t\)\(C\_\{t\},\\bar\{O\}\_\{t\}\)and whose memory\-free controller executes fromOtO\_\{t\}, and an intent policy whose head regresses the action from\(Ct,O¯t\)\(C\_\{t\},\\bar\{O\}\_\{t\}\)and whose controller refines it fromOtO\_\{t\}\. The hierarchical variant is sufficient on8/88/8Task A seeds atM=4,8,16,32M=4,8,16,32\(behavioral error0\.0000\.000; closed\-loop success0\.730\.73–0\.790\.79\), and the intent variant on8/88/8at the testedM=4,32M=4,32\(success0\.930\.93–0\.980\.98\)\.

Both give0/80/8on A′: the hierarchical head fails at cross\-entropy weights11,1010, and3030, and the intent head at both values ofβ\\beta; they retainβ1\\beta\_\{1\}\(used three steps after the reveal\) but loseβ2\\beta\_\{2\}\(nineteen steps\) before the place step, whose error is at chance\. A zero\-distortion recurrent realization using only\(Ct,O¯t\)\(C\_\{t\},\\bar\{O\}\_\{t\}\)nevertheless exists under Theorem[1](https://arxiv.org/html/2609.25757#Thmtheorem1), so these failures concern optimization under the stricter read\-side architecture rather than the theoretical rate target\.

#### Strict read\-side architectures with forecast supervision \(E1×\\timesE4\)\.

Table[6](https://arxiv.org/html/2609.25757#A2.T6)reports the two read\-side\-restricted architectures of the previous paragraph trained with the behavioral\-forecast target of Appendix[D\.2](https://arxiv.org/html/2609.25757#A4.SS2); the training\-time head also reads only\(O¯t,Ct\)\(\\bar\{O\}\_\{t\},C\_\{t\}\)\. Across all 64 runs,SΓ\>0\.9S\_\{\\Gamma\}\>0\.9at every first\- and second\-gap step andSG\>0\.9S\_\{G\}\>0\.9at every grasp and place step \(the minimum of either score over all checked runs and steps is1\.0001\.000\), and all eight cells pass the pre\-specified numerical gate\. The hierarchical head directly exposes the learned behavioral decision\. For the intent architecture, nearest\-class decoding of the intent vector is only a diagnostic proxy and is not the action executed by the controller; we therefore omit it from the behavioral\-error column\. Sufficiency and rate are evaluated from the code in both architectures\.

Table 6:Strict read\-side architectures on A′with behavioral\-forecast supervision \(8 seeds per cell; theory2\.00→1\.002\.00\\to 1\.00; forecaster removed at evaluation\)\. Without the forecast target the same architectures are sufficient on0/80/8seeds in every cell\.

### B\.5Sufficiency\-gate sensitivity

The main text uses one gate:SΓ\>0\.9S\_\{\\Gamma\}\>0\.9at every step of both gaps with a positive exact requirement andSG\>0\.9S\_\{G\}\>0\.9at the first grasp and first place step\. A relaxed gate that inspects only the second gap \(meanSΓ\>0\.9S\_\{\\Gamma\}\>0\.9\) and placement \(meanSG\>0\.9S\_\{G\}\>0\.9\) tests what survives the grasp but does not check the first\-gap join\. Of22112211A′\-family runs,928928pass the relaxed gate and879879the full gate; no run passes the full gate only, except on the weighing task, where the relaxed gate is undefined\. The task\-informed forecast cells are unchanged \(8/88/8at gaps 6, 10, 20 andM=16M=16\) exceptM=32M=32\(6→56\\to 5of88\), and the readout tasks are unchanged\. Counts that change \(relaxed→\\tofull\): plain−\-R on A′atβ=×10−3\\beta=2\\\!\\times\\\!10^\{\-3\}and×10−33\\\!\\times\\\!10^\{\-3\},4→24\\to 2and5→35\\to 3;−\-RF at×10−33\\\!\\times\\\!10^\{\-3\},2→12\\to 1; GRU bypass,6→16\\to 1; privileged join head at gap 20,8→58\\to 5, and atK=128K=128on the 4\-bit tier,2→12\\to 1; forecast supervision on the 4\-bit tier,3→03\\to 0\(K=16K=16\) and5→25\\to 2\(K=64K=64\); P\-I with forecast supervision atM=128M=128and512512,5→45\\to 4and2→02\\to 0; label stride 4,8→78\\to 7\. The bypass row is the instructive one: a continuous state can restore the placement class late, so a gate on the second gap alone mistakes late recovery for a correct recurrent state throughout\. Every sufficient seed under the full gate carries2\.002\.00–2\.152\.15bits in the first gap\. Under the full gate,479479of the615615passes among the16881688runs with stored per\-step requirements have every score above0\.9990\.999, and thresholds of0\.80\.8and0\.950\.95give650650and576576passes\.

### B\.6Observation side information: an injected\-leak audit

Memory diagnostics and direct observation probes exposed three benchmark flaws \(the pixel frame captured one step late, Appendix[E\.1](https://arxiv.org/html/2609.25757#A5.SS1); the resting pose that remembered the mass, Appendix[E\.3](https://arxiv.org/html/2609.25757#A5.SS3); sub\-millimetre contact artefacts, Appendix[B\.1](https://arxiv.org/html/2609.25757#A2.SS1)\)\. These cases motivate a controlled audit\. On A′\(gap 6,N=1152N=1152\) we*inject*a side channel of controlled magnitude: during the second gap the observed height of the object in the hand is shifted by±δ\\pm\\deltaaccording to the pending place class, on top of the22mm sensor noise\. The injection leaves the symbolic observation unchanged, so the requirement under that convention remains one bit in the second gap\.

For eachδ\\deltawe record a dataset, train the plain sole\-carrier policy \(β=10−3\\beta=10^\{\-3\}\) on1616seeds, and evaluate every policy in closed loop with the same leak\. The low\-rate prediction and flag A were fixed before the runs; flags B and B′were formulated afterwards\.

Table 7:Injected\-leak audit\. Top:1616seeds per magnitudeδ\\delta\.*Correct*means held\-out class error below0\.050\.05at the first place step\. Flag A, fixed before the runs, pairs correctness with second\-gap rate below0\.90\.9bit\. The exploratory flags pair correctness with failure of the full gate \(B\) or second\-gap memory alone \(B′\)\. Bottom: false alarms on separate clean recordings, conditioned on correctness\.
Atδ=8\\delta=8mm, success and slot accuracy are higher and action error is lower than in the clean condition\. Conventional performance metrics can therefore improve on a contaminated benchmark\. Flag A fails as a sensitive detector: only one of the9696runs is flagged\. Codes in behaviorally correct leaked policies can retain other content, with rates of1\.01\.0–2\.72\.7bits in the22mm condition, so a code rate below the requirement is not a necessary signature of leakage\.

#### Exploratory behavioral–representation disagreement\.

After observing these data, we defined flag B: a correct place decision paired with failure of the full sufficiency gate\. It flags11/1511/15,8/148/14, and13/1613/16correct policies atδ=2,4,8\\delta=2,4,8mm\. Applied unchanged to14321432separate clean runs recorded before the rule existed, it flags9/1669/166correct unsupervised runs and51/62751/627correct runs overall \(Table[7](https://arxiv.org/html/2609.25757#A2.T7)\)\. These are empirical alarm rates, not a leakage certificate\.

Two mismatches prevent a theorem\-based inference\. Behavioral sufficiency \(Theorem[3](https://arxiv.org/html/2609.25757#Thmtheorem3), Appendix[A\.1](https://arxiv.org/html/2609.25757#A1.SS1)\) assumes zero distortion, whereas “correct” here permits5%5\\%classification error\. A uniform binary target with symmetric4%4\\%error can haveH⁡\(G∣C\)=h2​\(0\.04\)≈0\.242H\(G\\mid C\)=h\_\{2\}\(0\.04\)\\approx 0\.242bit andSG≈0\.758S\_\{G\}\\approx 0\.758, satisfying our correctness criterion while failing the gate without any leak\. Moreover, Theorem[1](https://arxiv.org/html/2609.25757#Thmtheorem1)concerns reproduction over the full horizon; flag B pairs one place decision with a gate that also tests the first gap and grasp\. Most clean firings involve a correctly remembered place class but a lost grasp class\. They are genuine insufficiencies and do not imply a side channel\.

Flag B′instead pairs the place decision only with second\-gap memory\. It flags0/1660/166correct clean unsupervised runs and18/62718/627overall\. Its injected\-leak counts are11/1511/15,8/148/14, and4/164/16at2,4,82,4,8mm\. At88mm the code can copy the visible leak during the second gap, restoringSΓS\_\{\\Gamma\}there; thus even this matched signal is not monotone in leak magnitude\. Both rules are exploratory diagnostics requiring follow\-up\.

#### Direct observation probes\.

The raw\-observation probes in Appendix[B\.1](https://arxiv.org/html/2609.25757#A2.SS1)directly test whetherOtO\_\{t\}reveals a behavioral class not credited toO¯t\\bar\{O\}\_\{t\}, using held\-out prediction and the corresponding symbolic\-observation baseline\. Such probes address the side\-information discrepancy itself\. Behavioral–representation disagreement can motivate that check, but neither it nor the absence of a low code rate establishes whether a leak is present\.

### B\.7Independent\-recording replication

To check that the exact agreement between learned and theoretical rates is not a property of the recordings on which the solver was built, we recorded fresh expert datasets with new seeds \(A′gap 6: 1024 episodes; Task AM=4M=4and3232: 512 each,M=512M=512: 1024, all stratified over the masses\) and re\-evaluated 96 saved models under teacher forcing, rebuilding the solver from the new recordings\. On every dataset, the rebuilt solver returns the sameH⁡\(Γt∣O¯t\)H\(\\Gamma\_\{t\}\\mid\\bar\{O\}\_\{t\}\),H⁡\(GE,t∣O¯t\)H\(G\_\{E,t\}\\mid\\bar\{O\}\_\{t\}\)andH⁡\(Γts∣O¯t\)H\(\\Gamma^\{s\}\_\{t\}\\mid\\bar\{O\}\_\{t\}\)at every step \(maximum difference0\.0000\.000\), together with the same \(A4\)\-failure steps and visibility convention\. Every A′configuration preserves its seedwise sufficiency verdict; the first\-gap rate of sufficient seeds is2\.002\.00–2\.022\.02bits and post\-use rates move by at most0\.030\.03bit\. DIACRITIC also remains sufficient on every Task A seed \(8/88/8,8/88/8, and16/1616/16atM=4,32,512M=4,32,512\), with grasp rate0\.020\.02–0\.100\.10bit and at most0\.050\.05bit about mass\. System identification remains at1/81/8and0/80/8forM=4M=4and3232\. AtM=512M=512, its placement\-sufficiency count falls from7/87/8on the training recordings to4/84/8on fresh recordings; averaged over all seeds, its fresh\-data grasp rate is8\.708\.70bits,I⁡\(C;mass∣O¯\)=1\.26I\(C;\\mathrm\{mass\}\\mid\\bar\{O\}\)=1\.26bits, and mean placementSGS\_\{G\}changes from0\.9230\.923to0\.9020\.902\. This threshold sensitivity does not alter the rate separation, but it precludes a claim of seedwise replication for that baseline \(released code\)\.

We also re\-evaluated 72 models from the large\-MMP\-I and coverage\-controlled Task A experiments on independently recorded, stratified datasets\. For forecast\-distilled P\-I, the training\-to\-fresh all\-seed rates are2\.36→2\.482\.36\\to 2\.48atM=128M=128,2\.63→2\.722\.63\\to 2\.72atM=512,N=448M=512,N=448, and2\.31→2\.382\.31\\to 2\.38atM=512,N=1000M=512,N=1000; the corresponding late\-gate counts are5→35\\to 3,3→33\\to 3, and2→02\\to 0of 8\. The capacity\-relieved P\-I system\-identification rates replicate as6\.45→6\.496\.45\\to 6\.49and8\.72→8\.698\.72\\to 8\.69bits\. In the fully covered Task A cell, DIACRITIC and system identification both remain place\-sufficient on8/88/8seeds; their training\-to\-fresh rates are0\.01→0\.010\.01\\to 0\.01and8\.17→8\.198\.17\\to 8\.19bits, and their mass information is0\.01→0\.010\.01\\to 0\.01and1\.43→1\.361\.43\\to 1\.36bits\. The same policies obtain closed\-loop success0\.990\.99and0\.230\.23, respectively\. Thus full mode coverage does not remove the rate or closed\-loop separation on Task A, whereas large\-MMP\-I sufficiency is not stable across recordings\.

### B\.8Inference latency and storage footprint

#### Measurement\.

All three carriers use the paper’s widths \(d=32d=32, hidden width128128; the Transformer has two layers and four heads\)\. One policy step at history lengthTTis timed as the median of200200steps after2020warm\-up steps, on one CPU core at batch11and on a GPU at batch6464\(Table[8](https://arxiv.org/html/2609.25757#A2.T8)\)\. The Transformer re\-encodes its full token history at every step, so its per\-step cost grows withTTand its carried state is the token history itself \(64​\(T−1\)64\(T\-1\)floats:8\.28\.2kB atT=33T=33,262262kB atT=1024T=1024\)\. A key–value cache would remove the re\-encoding but still stores2​L​T​h2LThfloats per episode, about22kB per step at these widths, so the footprint still grows with the episode\. The recurrent carriers keep a fixed state of128128bytes \(GRU,3232floats\) or129129bytes \(DIACRITIC, the code vector plus one code index\) at a constant per\-step cost \(0\.550\.55ms on the CPU,1\.11\.1–1\.21\.2ms on the GPU\)\. The comparison isolates the cost of committing information in advance; it is not a recommendation of one deployment architecture\.

Table 8:Per\-step inference latency \(median of 200 steps\) as a function of the history lengthTT, with the widths of the paper’s models\. The Transformer re\-encodes its token history at every step \(no key–value cache\), so these timings describe that implementation; its carried history is64​\(T−1\)64\(T\-1\)floats \(8\.2 kB atT=33T=33, 262 kB atT=1024T=1024\), compared with 128 bytes for the GRU and 129 bytes for DIACRITIC \(state plus an 8\-bit code index\)\.

### B\.9Training\-time cost of forecast supervision

Table 9:Cost bookkeeping on A′gap 20 \(one B200 GPU; medians over runs\)\. The forecaster and the auxiliary head exist only during training\.Fitting the default forecaster adds about1\.5%1\.5\\%to the student’s training time; the measured 100\-step random\-offset forecaster fit takes under0\.2%0\.2\\%\. Separate*task\-informed*controls at A′gap 20 yield8/88/8sufficient seeds with a 100\-step forecaster and7/87/8with recorded\-action targets \(Appendix[D\.1](https://arxiv.org/html/2609.25757#A4.SS1)\)\. These learning results do not establish either performance claim for the frozen event\-agnostic configuration\. Deployment costs remain those of Appendix[B\.8](https://arxiv.org/html/2609.25757#A2.SS8)\.

### B\.10Compute and data collection

We collected demonstrations on one RTX 4090 using 64 environments, requiring≈\\approx8 s per 64 episodes\. Offline training used one GPU per run on an NVIDIA B200, with an RTX PRO 6000 as fallback; a 10k\-step run requires≈\\approx12–25 min\. Closed\-loop evaluation was performed in Isaac Sim on an RTX PRO 6000, requiring≈\\approx65 s per 128 episodes\. Metered scheduler usage over all cluster allocations for this work, including development, failed, and unreported runs, comprises3,0333\{,\}033training tasks: 403\.1 GPU\-hours on B200 across1,9251\{,\}925tasks and 178\.0 GPU\-hours on RTX PRO 6000 across1,1081\{,\}108tasks\. The2,9932\{,\}993closed\-loop evaluation tasks used 66\.9 GPU\-hours and the 52 independent re\-recording tasks 2\.5 GPU\-hours on RTX PRO 6000, for a total of≈\\approx651 GPU\-hours\. Corridor training, the external benchmarks, and teacher\-forced re\-evaluation on independent recordings ran in 464 CPU\-only tasks\. Local development, data collection, and the pixel evaluation used one RTX 4090 and were not metered\. The two training GPU types were used interchangeably for identical jobs, and we do not report a timing comparison between them\.

## Appendix CValidation of the behavioral memory target

### C\.1Finite toy validation

Figure 4:Toy validation\.\(a\)RE​\(0\)=log2⁡RR\_\{E\}\(0\)=\\log\_\{2\}Ris independent ofMM: the world uncertaintyH⁡\(Ubeh\)H\(U\_\{\\rm beh\}\)is fixed atlog2⁡16\\log\_\{2\}16while the behavioral rate islog2⁡R\\log\_\{2\}R\. \(b\) Blahut–Arimoto on the full history source agrees with the closed form for the reduced source \(Theorem[4](https://arxiv.org/html/2609.25757#Thmtheorem4)\) forM=4,8,16M=4,8,16; the deterministic frontier lies above the stochastic curve\. \(c\)–\(e\) The gap toy \(unified configuration, 8 seeds\): reliability as a function of the gap2length for−\-R and−\-RF, the learned post\-use rate of successful runs relative toH⁡\(Γt∣Ot\)=1H\(\\Gamma\_\{t\}\\mid O\_\{t\}\)=1, and the per\-step rate at gap2=6\{\}\_\{2\}=6against the three\-level staircase\. Panels \(c\)–\(e\) use near\-zero distortion at the second decision as their diagnostic\.#### Toy results\.

All toy results can be reproduced from the released result files; these experiments use CPUs and the stated numbers of seeds\.

*Gate \(T0\)\.*On the reveal toy, the solver returnsH⁡\(Γt\|O¯t\)=\[0,0,1,2\]H\(\\Gamma\_\{t\}\|\\bar\{O\}\_\{t\}\)=\[0,0,1,2\]bits\. On the gap toy, it returns\[0,0,1,2,2,2,2,1,1,1,1\]\[0,0,1,2,2,2,2,1,1,1,1\], corresponding to 2 bits at grasp, 1 bit in the gap, and 0 after use\. It reports bounds without constructingΓt\\Gamma\_\{t\}on a hand\-constructed non\-transitive instance and separately flags a violation of \(A2\)\. On a re\-reveal instance, it returns aΓ\\Gammastaircase of2→1→02\\to 1\\to 0, whileΓs\\Gamma^\{s\}remains at 2, demonstrating that the upper bound in Theorem[2](https://arxiv.org/html/2609.25757#Thmtheorem2)can be strict\.

*Exact rate–distortion \(Figure[4](https://arxiv.org/html/2609.25757#A3.F4)\)\.*Blahut–Arimoto on the reduced source agrees with the closed form in Corollary[2](https://arxiv.org/html/2609.25757#Thmcorollary2)to7\.8×10−167\.8\\times 10^\{\-16\}\. ForM∈\{4,8,16\}M\\in\\\{4,8,16\\\}, the history source gives the same curve, confirming both Theorem[4](https://arxiv.org/html/2609.25757#Thmtheorem4)and independence fromMMto1\.3×10−151\.3\\times 10^\{\-15\}\. The deterministic frontier is\(2\.0,0\),\(1\.5,\.25\),\(0\.81,\.5\),\(0,\.75\)\(2\.0,0\),\(1\.5,\.25\),\(0\.81,\.5\),\(0,\.75\), compared with2\.0/0\.79/0\.21/02\.0/0\.79/0\.21/0for the stochastic frontier \(Proposition[1](https://arxiv.org/html/2609.25757#Thmproposition1)\)\.

*Enumeration regime\.*When the population conditional is available \(β∈\[0,0\.1\]\\beta\\in\[0,0\.1\], 3 seeds\), both \-R and \-RF recoverH^​\(C\|O¯\)=\[0,0,1,2\]\\hat\{H\}\(C\|\\bar\{O\}\)=\[0,0,1,2\]exactly, withSΓ=1S\_\{\\Gamma\}=1,I^​\(C;z\|O¯\)=0\\hat\{I\}\(C;z\|\\bar\{O\}\)=0, andDT​V≤0\.002D\_\{TV\}\\leq 0\.002\. Atβ=0\.3\\beta=0\.3, the code collapses \(0\.67/1\.67 bits,D=0\.042D=0\.042\)\.

*Finite\-sample selective forgetting\.*Withn=48n=48,p=0\.7p=0\.7, and 4 seeds, the nuisance informationI^​\(C;z\|O¯\)\\hat\{I\}\(C;z\|\\bar\{O\}\)decreases as0\.81→0\.55→0\.050\.81\\to 0\.55\\to 0\.05, and the surplusH^​\(C\|Γ,O¯\)\\hat\{H\}\(C\|\\Gamma,\\bar\{O\}\)decreases as1\.54→0\.98→0\.101\.54\\to 0\.98\\to 0\.10, whenβ\\betachanges from0→0\.1→0\.30\\to 0\.1\\to 0\.3\. At the decision step,SΓS\_\{\\Gamma\}changes as0\.54→0\.94→0\.560\.54\\to 0\.94\\to 0\.56, with the final value reflecting collapse\. On the gap toy \(n=512n=512, gap 6\), the gap\-to\-gap rate changes from3\.27→2\.943\.27\\to 2\.94atβ=0\.01\\beta=0\.01to1\.90→1\.401\.90\\to 1\.40atβ=0\.1\\beta=0\.1, whileI^​\(C;z\|O¯\)\\hat\{I\}\(C;z\|\\bar\{O\}\)decreases from0\.36→0\.010\.36\\to 0\.01\. These measurements show reduced surplus under stronger rate pressure, together with a loss of sufficiency when the pressure becomes too large\.

*Factorized ablation\.*At gap 6 with 4 seeds andβ=0\.03\\beta=0\.03, DIACRITIC is sufficient on 3/4 seeds at a gap\-1 rate of 2\.00; the*nonpersistent*variant reaches 0/4, and the*uncond*variant reaches 4/4 and is indistinguishable on this toy\. The*continuous*carrier succeeds only atβ=0\.003\\beta=0\.003, with a 6\.9\-bit KL bound and 0\.32 bit of nuisance information\. This bound does not provide a comparably tight minimality measurement and is not an exact rate\. The*bypass*variant, in which a continuous GRU state persists across time, succeeds on 20/24 runs across 12 seeds×\\times2 learning rates but reportsH^​\(C\|O¯\)=0\.9\\hat\{H\}\(C\|\\bar\{O\}\)=0\.9–1\.9<2\.001\.9<2\.00atD=0D=0\. Because memory bypasses the code, this readout rate underestimates the memory rate, motivating the sole\-carrier requirement\.

*Violation of \(A2\)\.*Under a Bernoulli\(0\.3\) reveal loss, the solver flags \(A2\)\. Replacing the privileged quotient by the history\-conditional action lawPE​\(At∣Ht\)P\_\{E\}\(A\_\{t\}\\mid H\_\{t\}\)defines an observational relationΓobs\\Gamma^\{\\rm obs\}withH⁡\(Γobs∣O¯\)=\[0,0,1\.58,2\.58\]H\(\\Gamma^\{\\rm obs\}\\mid\\bar\{O\}\)=\[0,0,1\.58,2\.58\], where “lost” forms a third class\. Both objectives recover this observational rate, withSΓobs=1S\_\{\\Gamma^\{\\rm obs\}\}=1,DT​V≤0\.007D\_\{TV\}\\leq 0\.007, and zero nuisance information\. This does not restore reproducibility of the demonstrator’s privileged coupling, which lies outside the theorem’s assumptions \(cf\. Remark 4\.7 of[Yu \(2026\)](https://arxiv.org/html/2609.25757#bib.bib17)\)\.

*Discovery aids on the toy\.*Across gaps\{6,10,15\}\\\{6,10,15\\\}and 4 seeds, future\-sufficiency\-triggered code refinement does not change the number of sufficient seeds: \-R gives 3,3,4 with and without refinement, while \-RF gives 3,2,2 versus 3,2,1\. The annealed continuous scaffold is modestly more reliable, with 4,4,3 for scaffold\-R and 3,4,4 for scaffold\-RF\.

### C\.2A stochastic expert: the quotient is defined on action laws

Definition[1](https://arxiv.org/html/2609.25757#Thmdefinition1)groups hidden states by the expert’s action*distribution*, whereas the other experiments use deterministic experts\. In a finite check three equiprobable hidden states are displayed att=1t=1and hidden afterwards; att=3t=3the expert draws one of two actions withP⁡\(L∣s\)=p,p,1−pP\(L\\mid s\)=p,p,1\-p\. States 0 and 1 induce the same non\-degenerate law, and the solver merges them:H⁡\(GE,t∣Ot\)=H⁡\(Γt∣Ot\)=h2​\(1/3\)=0\.918H\(G\_\{E,t\}\\mid O\_\{t\}\)=H\(\\Gamma\_\{t\}\\mid O\_\{t\}\)=h\_\{2\}\(1/3\)=0\.918bit rather thanlog2⁡3=1\.585\\log\_\{2\}3=1\.585\(transitive; \(A4\) holds\)\. For the learner \(cross\-entropy on sampled demonstrations,K=8K=8\) the prediction that the rate would be within0\.10\.1bit of0\.9180\.918on every seed failed: atp=0\.7p=0\.7the quotient is found on4/84/8seeds without rate pressure and on none withβ≥0\.01\\beta\\geq 0\.01\(the code collapses to zero bits, KL0\.20\.2bit; distinguishing the third state is worth only0\.10\.1bit of likelihood at one step\); atp=0\.9p=0\.9it is found on6/86/8seeds atβ=0\\beta=0and on11–44of88with rate pressure\. In all6464runs the two equal\-law states share a code and no run uses three codes: whenever a memory is learned it is the distributional quotient, and the observed failures collapse the code rather than separate the equal\-law states \(cf\. §[5](https://arxiv.org/html/2609.25757#S5)\)\.

### C\.3Task A and readout controls

Figure 5:Task A controls\(8 seeds per point\)\. Left: code rate during grasp and transport as a function oflog2⁡M\\log\_\{2\}M\. The capacity\-relieved system\-identification rate rises approximately linearly withlog2⁡M\\log\_\{2\}M, whereas DIACRITIC withK=16K=16or6464remains at≤0\.13\\leq 0\.13bit\. Middle: retained mass information\. Right: fraction of seeds sufficient at the place step; the multi\-step\-inverse\-only memory never reaches sufficiency\.Codebook capacity for system identification\.AtK=16K=16, a code can represent at most44bits, whereas identifyingM=32M=32masses requires55bits\. The0/80/8sufficiency of theK=16K=16system\-identification baseline at largeMMmay therefore reflect limited capacity\. Table[10](https://arxiv.org/html/2609.25757#A3.T10)repeats this baseline withK=64K=64,128128, and256256, holding the data, training steps, andβ\\betafixed\. Greater capacity is necessary but not sufficient: system identification becomes sufficient atM=4M=4withK=64K=64, atM≤8M\\leq 8withK=128K=128, and on5/85/8seeds atM=16M=16withK=256K=256, but remains insufficient atM=32M=32\. With the largest codebook tested at eachMM, its rate increases from2\.252\.25to5\.225\.22bits; closed\-loop success ranges from0\.150\.15to0\.790\.79\. DIACRITIC withK=64K=64is unchanged at everyM≤32M\\leq 32: it is sufficient on8/88/8seeds, carries0\.050\.05–0\.130\.13bit withI⁡\(C;mass∣O¯\)≤0\.02I\(C;\\mathrm\{mass\}\\mid\\bar\{O\}\)\\leq 0\.02, and reaches success of0\.990\.99–1\.001\.00\. System identification is therefore capable of solving the smaller instances, but its learned rate grows with world complexity and does not deliver comparable control accuracy through a shared carrier\. In settings whereKKis not fixed by a theoretical fiber bound, a practical held\-out selection rule is to increaseKKuntil the sufficiency gate is first met and the measured rate stabilizes across successive values, then freezeKKbefore final evaluation\.

Multi\-step inverse objectives\.Multi\-step inverse kinematics and agent\-centric state discovery\([Mhammedi et al\., 2023](https://arxiv.org/html/2609.25757#bib.bib21);[Lamb et al\., 2023](https://arxiv.org/html/2609.25757#bib.bib16)\)train an encoder to predictAtA\_\{t\}from the representation atttand a future observationOt\+jO\_\{t\+j\}\. We add a recurrent analogue to our model\. This is a MusIK/ACSD\-*style*baseline rather than a direct implementation of either published algorithm, and uses a headq⁡\(At∣Ct,Ot,Ot\+j\)q\(A\_\{t\}\\mid C\_\{t\},O\_\{t\},O\_\{t\+j\}\)withj∼U​\{1,…,8\}j\\sim U\\\{1,\\dots,8\\\}\. We evaluate two variants\. In*inverse only*, the inverse head and VQ losses shape the memory, the policy head receives a detached\(Ot,Ct\)\(O\_\{t\},C\_\{t\}\), and no rate term is used\. In*inverse \+ imitation \+ rate*, the inverse head is added to our objective\.

Table[10](https://arxiv.org/html/2609.25757#A3.T10)shows that inverse\-only memory never reaches sufficiency in these tests: it obtains0/80/8at everyMM, the recorded class\-error metric is0\.300\.30–0\.330\.33, and closed\-loop success of0\.030\.03–0\.060\.06\. Instead, it retains0\.30\.3–0\.60\.6bit about the mass\. A future observation can reveal where the object was placed, allowing the inverse head to predict an earlier action without retaining the slot across the gap\. This explains why the inverse target need not induce the anticipatory representation; the learned nuisance information is consistent with its emphasis on dynamics\. When added to imitation and the rate term, the inverse objective has no measurable effect: the model reaches8/88/8, matches the behavioral code rate andI⁡\(C;mass∣O¯\)I\(C;\\mathrm\{mass\}\\mid\\bar\{O\}\), and obtains closed\-loop success of0\.990\.99–1\.001\.00\. For demonstrations from a deterministic expert, the inverse problem does not depend on the nuisance, allowing the rate term to remove it\. Empirically separating the objectives for a layer\-2 nuisance, an agent\-centric variable ignored by the expert, requires behaviorally equivalent action randomness whose observable effect depends on that nuisance; Appendix[C\.4](https://arxiv.org/html/2609.25757#A3.SS4)provides such a construction\.

Table 10:Task A controls, 8 seeds per cell\. Top block: seeds sufficient at the place step / closed\-loop success \(128 episodes per seed\)\. Bottom block: code rate during grasp \(bits\) /I⁡\(C;mass∣O¯\)I\(C;\\mathrm\{mass\}\\mid\\bar\{O\}\)\(bits\)\.#### World complexity beyond five bits\.

Table[11](https://arxiv.org/html/2609.25757#A3.T11)extends Task A toM=128M=128and512512with datasets stratified over the masses \(each mass recorded 4 and 2 times\)\. The predictions were fixed before the runs: DIACRITIC rate≤0\.2\\leq 0\.2bit during grasp and transport,I⁡\(C;mass∣O¯\)≤0\.05I\(C;\\mathrm\{mass\}\\mid\\bar\{O\}\)\\leq 0\.05bit, closed\-loop success≥0\.9\\geq 0\.9\. System identification atK=16K=16is a capacity\-stress diagnostic \(log2⁡K<log2⁡M\\log\_\{2\}K<\\log\_\{2\}Mby construction\); the capacity\-relieved rows useK=256K=256andK=1024K=1024\. Training usesN=448N=448episodes at everyMM, which covers all 128 masses atM=128M=128but only 352 of 512 atM=512M=512; theN=896N=896row is a data\-scale sensitivity control, and theN=1000N=1000rows, in which every mass appears in training, are the coverage\-controlled comparison\. With full coverage, system identification atK=1024K=1024becomes place\-sufficient on all seeds \(also on fresh recordings, Appendix[B\.7](https://arxiv.org/html/2609.25757#A2.SS7)\), yet still stores8\.28\.2bits and succeeds in closed loop on0\.230\.23of episodes against0\.990\.99for DIACRITIC: the separation is not a coverage artifact\.

Table 11:Task A atM=128M=128and512512\(8 seeds per row; 128 closed\-loop episodes per seed\)\.
#### Non\-zero behavioural memory at growing world complexity \(readout task\)\.

Task A has exact minimum00during grasp because mass does not affect the expert’s action choice\. To test the non\-zero counterpart of the world\-complexity comparison, we retain the Franka grasp–transport–place task, attach the object kinematically so that mass is not re\-revealed, and display the hidden classθ∈\[M\]\\theta\\in\[M\]on a 9\-channel binary readout during three scan steps\. The expert grasps according to the quartile ofθ\\theta\(22bits, used33steps after the readout\) and places according to its half \(*readout\-2*:11bit, used≈20\\approx 20steps later\) or octile \(*readout\-3*:88slots separated by at least2121cm\)\. These nested threshold classes admit the exact gap\-to\-transport profiles2→1→02\\to 1\\to 0and3→3→03\\to 3\\to 0at everyM∈\{32,128,512\}M\\in\\\{32,128,512\\\}; the solver certifies A2, A4 and transitivity, whileH⁡\(Ht∣O¯t\)=log2⁡M\+1H\(H\_\{t\}\\mid\\bar\{O\}\_\{t\}\)=\\log\_\{2\}M\+1bits during transport\. During the first three side\-approach steps, the end\-effector pose reveals the grasp quartile, temporarily making the pending readout\-2 slot conditionally deterministic; its conditional rate returns to11bit once this side information disappears\. This is temporary redundancy with the current observation, not expiration: the recurrent state must preserve the distinction until placement\. The visibility convention accordingly uses unconditional class means for the nested factor\. Every dataset contains10241024stratified episodes \(N=448N=448for training\), and an independent recording provides the replication set\. Table[12](https://arxiv.org/html/2609.25757#A3.T12)reports teacher\-forced re\-evaluation on both recordings and closed\-loop evaluation over128128episodes per policy\.

An analog single\-channel display was tried first; with eight threshold classes even the full\-history forecaster could not decode the slot \(held\-out error0\.280\.28–0\.510\.51\), so we use the binary display\. On readout\-2 the plain rate\-penalised learner \(−\-R,K=16K=16,β=10−3\\beta=10^\{\-3\}\) is sufficient on8/88/8,7/87/8and8/88/8seeds atM=32,128,512M=32,128,512, with first\-gap rate2\.002\.00–2\.032\.03bits \(theory2\.002\.00\) and transport rate1\.061\.06–1\.231\.23\(theory1\.001\.00\)\. Forecast supervision reaches8/88/8at everyMMwith rates2\.022\.02–2\.082\.08and1\.091\.09–1\.271\.27bits, and closed\-loop success0\.950\.95,0\.900\.90and0\.960\.96\(−\-R:0\.940\.94,0\.850\.85,0\.870\.87\)\. The behavioral learners reproduce every verdict on the fresh recordings, with rate changes of at most0\.010\.01bit\. System identification stores3\.33\.3bits atK=16K=16and4\.84\.8,5\.65\.6and7\.57\.5bits with the largest codebook tested \(K=256,256,1024K=256,256,1024;3\.73\.7–7\.07\.0bits aboutθ\\theta\), yet reaches only0\.190\.19–0\.260\.26closed\-loop success with slot accuracy0\.600\.60–0\.730\.73\. Thus the behavioral rate remains fixed at a non\-zero value while the system\-identification rate grows with slope0\.670\.67per bit oflog2⁡M\\log\_\{2\}M\.

Readout\-3 probes the next representational scale\. Forecast supervision reaches sufficiency on6/86/8,5/85/8and4/84/8seeds, identically on the fresh recordings, and the sufficient codes carry3\.033\.03–3\.173\.17bits against the3\.003\.00\-bit prediction\. Closed\-loop evaluation reveals the remaining occupancy gap: across all seeds, success is0\.150\.15–0\.280\.28and slot accuracy0\.470\.47–0\.590\.59; among teacher\-forced sufficient seeds, success is0\.300\.30–0\.480\.48despite0\.960\.96teacher\-forced place accuracy\. AK=8K=8carrier, which meets the cardinality lower boundmaxt,o⁡\|Γt\|o=8\\max\_\{t,o\}\|\\Gamma\_\{t\}\|\_\{o\}=8, trains to sufficiency on0/240/24seeds, whereasK=16K=16realizes the three\-bit code on15/2415/24seeds\. The unsupervised learner reaches that code on6/246/24seeds; atM=128M=128and512512, those sufficient seeds attain0\.840\.84closed\-loop success\. These results establish three\-bit offline attainability while locating reliable closed\-loop learning at two bits\.

Table 12:Readout task, 8 seeds per cell,K=16K=16unless noted \(sys\-ID largestKK: 256, 256, 1024\); each cell listsM=32M=32/128128/512512\. Rates are all\-seed means except for theK=16K=16forecast row on readout\-3, which averages sufficient seeds; closed\-loop success is always averaged over all seeds\. Behavioral\-learner verdicts replicate on the fresh recordings; the largest\-KKreadout\-2 sys\-ID count atM=512M=512changes from1/81/8to0/80/8\. Theory is in brackets\.
#### Separating the information and control costs of system identification\.

The sys\-ID baseline decodesθ\\thetafrom\(Ot,Ct\)\(O\_\{t\},C\_\{t\}\), forcing world and behavioral information through the same rate\-penalized codebook\. We separate these roles on Task A atM=32M=32\(all masses observed in training\) and at the full\-coverageM=512M=512cell \(N=1000N=1000\)\. A*dual carrier*adds a recurrent hard\-VQ codeCtidC^\{\\rm id\}\_\{t\}\(K=256K=256/10241024\) read only by theθ\\thetahead, while the policy reads only\(Ot,Ct\)\(O\_\{t\},C\_\{t\}\)withK=16K=16; both codes are rate\-penalized and share the observation and action encoders\. We also test a shared carrier with theθ\\thetaloss down\-weighted to0\.10\.1, and a dual carrier that stops the identification gradient at the shared encoders\. The dual models reach0\.900\.90–0\.970\.97closed\-loop success, compared with0\.070\.07–0\.520\.52across the displayed shared\-carrier controls\. Their behavioral codes carry0\.280\.28–0\.490\.49bit with at most0\.030\.03bit about mass, while the identification codes carry4\.14\.1–6\.06\.0bits, including2\.52\.5–5\.25\.2bits aboutθ\\theta\. Down\-weighting alone leaves4\.54\.5/6\.36\.3bits in the shared carrier and reaches0\.380\.38/0\.520\.52success; stop\-gradient changes the dual result little\. Independent recordings reproduce every behavioral sufficiency verdict and rate to within0\.040\.04bit\. Thus system identification retains a distinct, additive rate burden, whereas its control penalty is not a consequence of that rate alone: in these controls it appears when both objectives share one regularized carrier\.

Table 13:Task A, sys\-ID controls \(8 seeds\): transport rates of the behavioural codeCCand of the identification codeCidC^\{\\rm id\}\(bits\), mass information carried byCC,θ\\thetainformation carried byCidC^\{\\rm id\}, and closed\-loop success \(mean, min\)\. “shared” = theθ\\thetahead readsCCitself\.AcrossM=4M=4to512512the capacity\-relieved system\-identification rate is2\.25,3\.66,4\.53,5\.22,6\.84,8\.712\.25,3\.66,4\.53,5\.22,6\.84,8\.71bits, with a fitted slope of0\.880\.88inlog2⁡M\\log\_\{2\}Mand intercept0\.800\.80\. AtM=128M=128and512512, mean closed\-loop success is respectively0\.170\.17and0\.190\.19, despite sufficiency on3/83/8and7/87/8seeds\. TheM=512M=512teacher\-forced result is not robust to the evaluation sample: on fresh expert recordings the count falls to4/84/8, while mean placementSGS\_\{G\}changes from0\.9230\.923to0\.9020\.902\(Appendix[B\.7](https://arxiv.org/html/2609.25757#A2.SS7)\)\. More generally, the gate tests whether the code contains the discrete placement class under expert occupancy; it does not certify continuous\-action accuracy or stability under policy\-induced observations, and the present evaluations do not separate these two possible sources of the low closed\-loop success\.

### C\.4Signpost corridor: exact requirements and learned rates

The A′manipulation task satisfies \(A4\), whereas Task A violates \(A4\) but remains transitive\. To test whether the empirical phenomena depend on this design, we introduce a partially observed grid corridor that shares only the solver and learner with A′\. ItsW=1W=1setting is in the exact regime; the drift variantsW\>1W\>1are non\-transitive, and matching lower and upper bounds certify the rates of the recordedW=3,5,7W=3,5,7instances \(Proposition[3](https://arxiv.org/html/2609.25757#Thmproposition3)\)\. The agent moves from left to right throughlog2⁡M\\log\_\{2\}Msignpost cells, each of which reveals one bit of a hidden goal codeθ∈\[M\]\\theta\\in\[M\]\. It then traverses a wide hall of gap1cells, reaches a junction where the expert turns according toβ1​\(θ\)\\beta\_\{1\}\(\\theta\), traverses a second hall of gap2cells, and reaches a second junction governed byβ2​\(θ\)\\beta\_\{2\}\(\\theta\)\. Observations are symbolic and indicate the start texture, sign bit, hall identity and lateral cell, or junction identity; actions are forward, sidestep, and turn\. In the first cell of each hall, the expert sidesteps left or right uniformly at random, providing behaviorally equivalent action randomness\.

A hidden lateral*drift*w∈\[W\]w\\in\[W\], displayed on the start sign, shifts the agent sideways at every hall step\. The drift changes the transitions experienced by the agent and is revealed again by its observed lateral position, but never affects the expert’s action\. It therefore serves as an agent\-centric, behaviorally irrelevant nuisance, analogous to a current rather than a mass\. AtW=1W=1, the solver certifies \(A2\), \(A4\), and transitivity at every step and returns the staircaseH⁡\(Γt∣O¯t\)=2→2→1H\(\\Gamma\_\{t\}\\mid\\bar\{O\}\_\{t\}\)=2\\to 2\\to 1across hall 1, junction 1, and hall 2, compared withH⁡\(Ht∣O¯t\)=3H\(H\_\{t\}\\mid\\bar\{O\}\_\{t\}\)=3–77bits\. AtW\>1W\>1, \(A4\) fails in the halls\. The strong\-congruence bracket is\[0,2\+log2⁡W\]\[0,\\,2\+\\log\_\{2\}W\]; the finite\-instance certificates tighten the hall requirements to exactly22and11bits\. We train with the toy trainer using raw VQ,K=16K=16, 6k steps,β=0\.01\\beta=0\.01, and atanh\\tanh\-bounded residual state to prevent codebook collapse\. Each cell contains 8 seeds, and the multi\-step\-inverse and system\-identification heads match those in Appendix[C\.3](https://arxiv.org/html/2609.25757#A3.SS3)\.

Figure 6:Signpost corridor\.\(a\) In the exact regime atM=16M=16, sufficient seeds carry2\.002\.00bits through hall 1 whileH⁡\(GE,t∣O¯t\)=0H\(G\_\{E,t\}\\mid\\bar\{O\}\_\{t\}\)=0, then reduce their rate after junction 1; the history contains 5 bits\. Expiration is partial \(1\.2–1\.8 compared with 1\.0\)\. \(b\) Fraction of sufficient seeds as a function of the reveal→\\touse distance forβ2\\beta\_\{2\}\(hall\-2 length 1–20\)\. The plain realizations decay with distance \(logit slope−0\.11\-0\.11/step,p=0\.01p=0\.01with configuration\-clustered errors\), whereas the annealed scaffold does not\. \(c\) Drift information retained in the halls as a function oflog2⁡W\\log\_\{2\}W\. The multi\-step\-inverse memory carries0\.30\.3–0\.70\.7bit about the drift, with or without the rate term, whereas DIACRITIC and system identification carry≤0\.08\\leq 0\.08bit\. The dashed line denotes the additional rate assigned by the strong congruence\.Staircase and world complexity\.Every sufficient seed carries2\.002\.00–2\.072\.07bits in hall 1 atM=4M=4,1616, and6464, while the history grows from 3 to 7 bits\. The sufficient\-seed counts are7/87/8,4/84/8, and6/86/8for−\-R;8/88/8,8/88/8, and6/86/8for−\-RF; and8/88/8,7/87/8, and6/86/8forβ=0\\beta=0\. System identification is sufficient atM=4M=4, whereθ\\theta*is*the join, but reaches0/80/8atM=16M=16and6464while retaining2\.22\.2–2\.52\.5bits aboutθ\\theta\. The multi\-step\-inverse memory reaches0/80/8at everyMM, with junction error0\.50\.5\. Expiration after junction 1 is partial in this domain:−\-R retains1\.171\.17–1\.361\.36bits relative to the1\.001\.00boundary\. Unlike on A′, increasingβ\\betawithin the admissible range does not reduce this surplus \(β=0\\beta=0:1\.281\.28–1\.301\.30\)\. We therefore treat the discrepancy as a limitation of the realization rather than of the target\.

Learning horizon\.For hall\-2 lengths1,3,6,10,15,201,3,6,10,15,20, corresponding to distances 6–25,−\-R reaches sufficiency on8,8,4,5,5,58,8,4,5,5,5of 8 seeds, and−\-RF reaches8,8,8,5,1,38,8,8,5,1,3\. Increasing the length of hall 1 to 8 and 15 cells reduces−\-R to2/82/8and0/80/8\. We pool the 208 runs atW=1W=1, comprising 416 run×\\timesclass rows and 24 configurations\. For the plain realizations, the distance coefficient is−0\.11\-0\.11/step \(p=0\.01p=0\.01, configuration bootstrap CI\[−0\.20,−0\.05\]\[\-0\.20,\-0\.05\]\), matching the sign and order of the A′coefficient \(−0\.15\-0\.15\)\. The annealed scaffold, which only mitigates the barrier on A′, reaches8/88/8at every corridor distance up to 25, with slope00\. The recurrence of the distance effect outside manipulation supports a temporal optimization interpretation, while the differing benefit of the scaffold shows that its mechanism is domain\-dependent\.

Layer\-2 nuisance\.AtW=3,5,7W=3,5,7drift modes, DIACRITIC retains≤0\.08\\leq 0\.08bit about the drift and2\.002\.00–2\.082\.08bits in hall 1, with no dependence onWW; the strong congruence would instead assign2\.862\.86–3\.823\.82bits\. The multi\-step\-inverse memory retains0\.30/0\.50/0\.470\.30/0\.50/0\.47bit with the rate term and0\.54/0\.71/0\.480\.54/0\.71/0\.48without it, because inferring a sidestep from the observed lateral displacement requires the drift\. This memory never reaches sufficiency\. System identification retains2\.42\.4–2\.62\.6bits aboutθ\\thetaand≤0\.04\\leq 0\.04bit aboutww\. These results separate the targets empirically: a control\-endogenous objective retains an agent\-centric quantity ignored by the expert, whereas the behavioral objective need not\. The objectives become distinguishable only when behaviorally equivalent action randomness makes the inverse problem depend on the nuisance, which is why Task A with a deterministic expert cannot exhibit this separation \(Appendix[C\.3](https://arxiv.org/html/2609.25757#A3.SS3)\)\.

Matched\-distortion controls\.The inverse\-only memory fails in the corridor even withK=64K=64codes, reaching0/80/8at bothW=1W=1andW=5W=5\. AtW=5W=5, it retains≈1\.0\\approx 1\.0bit about the drift and2\.072\.07bits in hall 1, but its junction error remains at chance\. Its failure is therefore not caused by codebook capacity\. When trained*jointly*with imitation and the behavioral code rate term, the same head has no measurable effect: sufficiency is4/84/8atW=1W=1and6/86/8atW=5W=5, matching−\-R in the corresponding cells, and drift information is0\.090\.09–0\.120\.12bit, compared with0\.300\.30–0\.700\.70for the inverse objective alone\. Thus, the inverse objective retains drift only in the absence of pressure to remove it; it does not enforce minimality, and the rate term removes this information at negligible cost to the inverse loss\.

WithK=64K=64, system identification reaches sufficiency on4/84/8seeds, compared with0/80/8atK=16K=16, at a cost of3\.93\.9bits whenθ\\thetacontains 4 bits\. DIACRITIC remains unchanged atK=64K=64, reaching8/88/8at2\.032\.03bits\. Thus capacity contributes to theK=16K=16system\-identification failure but does not fully resolve it; the inverse\-only objective remains at0/80/8even withK=64K=64\(preceding paragraph\)\. The continuous\-carrier ablation, implemented as a GRU bypass, requires a larger learning rate in this domain: it reaches00–2/82/8at×10−43\\\!\\times\\\!10^\{\-4\}and7/87/8at×10−33\\\!\\times\\\!10^\{\-3\}\. When successful, its readout code carries1\.51\.5bits whereΓ\\Gammarequires2\.02\.0, reproducing the undercount observed on the toy and A′\. A code that is not the sole memory carrier therefore does not yield a valid memory rate\.

### C\.5Plain learned code trajectories

Figure 7:Every plain seed, with sufficiency checked first\.A′at gap 6,N=1152N=1152,β=0\\beta=0and10−310^\{\-3\}, eight seeds per cell\. Blue trajectories pass the full gate of §[3](https://arxiv.org/html/2609.25757#S3); red trajectories do not\. The black curve is the certified behavioral memory rate and the orange curve the instantaneous requirement\. Rates use expert occupancy\. A low code rate on a failed seed is not evidence of compression at preserved behavior\.This view complements Figure[3](https://arxiv.org/html/2609.25757#S4.F3): it retains the variation across all plain seeds, while the main figure tests a realization whose behavioral decoder has the same side information as the theoretical model\.

## Appendix DLearning compact behavioral memory

The following experiments concern acquisition of the independently defined target\. Counts use the full gate unless a relaxed or class\-specific diagnostic is explicitly identified\. Rates conditional on passing and all\-seed control outcomes are kept separate\.

### D\.1Event\-agnostic future\-behavior supervision

#### Objective\.

For every training episode and stepttwe draw one offsetj∼U​\{1,…,T−1\}j\\sim U\\\{1,\\dots,T\-1\\\}and a training\-only head regresses, from\(go​\(Ot\),Ct\)\(g\_\{o\}\(O\_\{t\}\),C\_\{t\}\)and an embedding ofjj, the forecast ofAt\+jA\_\{t\+j\}produced by a causal Transformer that was trained from scratch on the same recorded episodes to predict all future actions from the history prefix \(offsets beyond the end of the episode are masked; nothing else is\)\. No reveal step, decision step, divergent action or latent label enters the construction; before the class is revealed the forecaster outputs the conditional mean\. The auxiliary weight is11until60%60\\%of the10410^\{4\}training steps and decreases linearly to00at80%80\\%; the last20%20\\%optimize imitation and rate only\. Everything else is the frozen configuration of Appendix[D\.2](https://arxiv.org/html/2609.25757#A4.SS2)\(K=16K=16,β=10−3\\beta=10^\{\-3\},N=448N=448\)\. The*task\-informed*forecast of Appendix[D\.2](https://arxiv.org/html/2609.25757#A4.SS2)instead predicts only the first grasp and the first place action, supervised from the last reveal step and masked once used\.

Table 14:Forecast supervision on A′\(state observations\), 8 seeds per cell\. Top: sufficient seeds under the full\-trajectory gate \(in parentheses: the same models re\-evaluated on an independent recording\)\. Bottom: first\-gap→\\tosecond\-gap rate of the sufficient seeds \(bits; theory2\.00→1\.002\.00\\to 1\.00\) / closed\-loop success over all seeds \(128 episodes per policy\)\. EA = event\-agnostic random\-offset forecast; TI = task\-informed forecast\.On the independent recordings every seed\-wise verdict of the frozen configuration is reproduced \(3636pass→\\topass,44fail→\\tofail, no switches\); for the task\-informed forecast3030of the3232re\-evaluated verdicts agree, with one switch in each direction, and over all520520models re\-evaluated in this appendix and Appendices[D\.2](https://arxiv.org/html/2609.25757#A4.SS2)and[E\.3](https://arxiv.org/html/2609.25757#A5.SS3)the agreement is96%96\\%\. Overall sufficiency is similar for the two objectives \(36/4036/40and37/4037/40\) with setting\-dependent differences \(M=16M=16:6/86/8against8/88/8;M=32M=32:7/87/8against5/85/8, closed loop0\.930\.93against0\.600\.60\), while the task\-informed target is more rate\-efficient after first use \(1\.061\.06–1\.111\.11against1\.091\.09–1\.261\.26bits\)\. Across these cells, policies that pass the full\-trajectory gate also exhibit high closed\-loop success \(0\.900\.90–1\.001\.00; e\.g\.M=16M=16:0\.790\.79over all seeds,1\.001\.00over sufficient seeds\); we report this as an association, not a causal claim\. A variant that sums the loss over all future offsets is not robust \(3/83/8at gap 20 at weight11,1/81/8on the independent recording;7/87/8only with weight0\.10\.1and annealing\) and is not used\.

#### Pixels\.

Table[15](https://arxiv.org/html/2609.25757#A4.T15)repeats the comparison from pixels; the forecaster reads the same images\. The60%60\\%schedule was the best of three tried on seeds 0–7 at gap 20; seeds 8–15 were run afterwards with nothing changed\. On those held\-out seeds plain training is sufficient on0/80/8\(closed\-loop success0\.080\.08, slot accuracy0\.550\.55\), the event\-agnostic objective on6/86/8\(0\.570\.57,0\.840\.84\) and the task\-informed forecast on7/87/8\(0\.490\.49,0\.820\.82\); the40%40\\%schedule, fixed before any pixel run, gives6/86/8\(0\.650\.65,0\.880\.88\)\. Grasp\-side accuracy, whose class is used two steps after the reveal, is0\.970\.97–1\.001\.00for every learner: the differences arise only at the slot chosen twenty steps later\. Among offline\-sufficient seeds the slot accuracy is0\.970\.97\(event\-agnostic\) and0\.930\.93\(task\-informed\), against0\.530\.53–0\.590\.59for insufficient seeds, so the offline gate predicts the closed\-loop memory decision; task success is further limited by pixel control precision \(action error2\.62\.6–6\.06\.0cm\)\.

Table 15:Pixel A′: sufficient seeds \(full\-trajectory gate\), rates of the sufficient seeds \(theory2\.00→1\.002\.00\\to 1\.00\), and closed loop over all seeds \(128 episodes per policy\)\.
#### What the un\-annealed objective keeps\.

In the second gap onlyβ2\\beta\_\{2\}is required and the symbolic observation is uninformative, soI⁡\(Ct;β1∣β2\)I\(C\_\{t\};\\beta\_\{1\}\\mid\\beta\_\{2\}\)measures the expired class still carried by the code\. It is at most0\.020\.02bit for every learner, annealed or not, from states and from pixels, whileI⁡\(Ct;β2∣β1\)=0\.98I\(C\_\{t\};\\beta\_\{2\}\\mid\\beta\_\{1\}\)=0\.98–1\.001\.00bit\. The surplusH^​\(Ct∣Γt,O¯t\)\\hat\{H\}\(C\_\{t\}\\mid\\Gamma\_\{t\},\\bar\{O\}\_\{t\}\)of the un\-annealed objective is1\.021\.02bits from pixels and1\.161\.16from states, compared with0\.130\.13and0\.170\.17after annealing and0\.000\.00for the task\-informed forecast\.

The scripted expert moves along linear segments whose per\-step displacement is fixed by the starting pose\. During the second gap this is the grasp pose, set by the object’s initial position \(uniform within±2\\pm 2cm\), which is no longer visible and changes the displacement by up to40%40\\%\. Conditioned onβ2\\beta\_\{2\}, the un\-annealed code carries0\.560\.56bit about the initialxx\-position quartile \(0\.010\.01–0\.030\.03after annealing\) and at most0\.030\.03bit about theyyposition, hover nuisance, distractor, or mass\. This information can improve an open\-loop action forecast even though the symbolic behavioral target does not require it\.

At the midpoint of the second gap \(t=28t=28\), a separate five\-fold cross\-validated analysis gives conservative lower bounds on the information explained by these features\. Relative to the1\.471\.47\-bit surplus at that step, initial position explains at least0\.380\.38bit \(26%26\\%\), increasing to0\.510\.51bit \(35%35\\%\) when noisy observation history is included\. These are lower bounds, not a limit on the explainable share; they use a single\-step measurement rather than the phase average reported above\.

#### Intervening on initial\-position variability\.

We re\-recorded A′gap 20 without the±2\\pm 2cm object\-position jitter and retrained the un\-annealed objective \(8 seeds; prediction fixed beforehand: surplus below0\.60\.6bit\)\. All8/88/8seeds are sufficient, with2\.00→1\.232\.00\\to 1\.23bits, second\-gap surplus0\.230\.23bit, and closed\-loop success0\.940\.94\. The original jittered task gives7/87/8sufficient seeds and1\.161\.16bits of surplus\. Removing this source of task variation therefore substantially reduces predictive surplus under retraining, supporting initial\-position variability as an important driver\. The0\.930\.93\-bit difference compares separately trained policies; it is not a decomposition of the position information stored by the original code\. These results help explain excess rate under open\-loop prediction \(Appendix[D\.3](https://arxiv.org/html/2609.25757#A4.SS3)\) while leaving its full content unresolved\.

#### Dependence on the forecaster \(A′gap 20, task\-informed target\)\.

Regressing the recorded future actions directly, without any forecaster, is sufficient on7/87/8seeds \(all\-seed closed\-loop success0\.820\.82\)\. Forecasters trained for300300,100100and3030steps instead of30003000\(held\-out class error of their pending\-action forecast0\.000\.00,0\.080\.08,0\.250\.25\) give7/87/8,8/88/8and4/84/8on the independent recording \(all\-seed closed loop0\.750\.75,0\.780\.78,0\.670\.67\)\. Forecasts that are*consistently wrong*—for55,1010or20%20\\%of the training episodes the target is the forecast of an episode of a different behavioral class—give2/82/8,0/80/8and0/80/8on the clean independent recording, with closed\-loop slot accuracy0\.910\.91,0\.770\.77and0\.620\.62: the student tolerates an imprecise forecaster but distils its systematic errors, and the entropy gate fails once about5%5\\%of the episodes receive a wrong code\. With recorded actions as targets the all\-offset event\-agnostic objective fails \(0/80/8\): before the reveal the recorded future actions are unpredictable, whereas the forecaster supplies their conditional mean\.

### D\.2Task\-informed supervision and teacher\-state controls

This appendix documents the*task\-informed*forecast and the teacher\-state control; the event\-agnostic objective of the main text is documented in Appendix[D\.1](https://arxiv.org/html/2609.25757#A4.SS1)\. The forecaster is a causal Transformer \(two layers, width 128,d=32d=32\) trained from scratch on the recorded training episodes for 3000 steps to output, at every stepttfrom the last reveal step onward, the expert’s action at the first grasp step and at the first place step, each masked once that step is reached, so that its target atttis exactly the pending class\-dependent behavior; it uses only recorded observations and actions\. Its held\-out error is0\.0000\.000\(per\-dimension MSE of standardized actions\) on every dataset\. The student is the plainK=16K=16DIACRITIC \(β=10−3\\beta=10^\{\-3\}\) with one additional training\-time head that regresses the forecaster’s output from\(go​\(Ot\),Ct\)\(g\_\{o\}\(O\_\{t\}\),C\_\{t\}\)under the same mask \(weight 1\); the head and the forecaster are discarded after training, and all rates and rollouts are those of the student alone\. The teacher\-state control replaces the target by the frozen full\-history Transformer baseline’s pre\-quantization state attt\(the full\-history baseline of §[5](https://arxiv.org/html/2609.25757#S5), same seed and data\)\.

Table 16:Distillation controls on A′and P\-I \(8 seeds per cell; full\-trajectory gate of §[4\.1](https://arxiv.org/html/2609.25757#S4.SS1); the event\-agnostic objective is in Table[14](https://arxiv.org/html/2609.25757#A4.T14)\)\. Top block: seeds sufficient\. Bottom block: rates and closed\-loop success of the sufficient seeds under the forecast target\.#### The 4\-bit join\.

On tier 16 the complete anticipatory join requireslog2⁡\(4⋅4\)=4\\log\_\{2\}\(4\\cdot 4\)=4bits under the uniform model; its empirical test\-split entropy is3\.973\.97bits\. Under the gate of §[3](https://arxiv.org/html/2609.25757#S3)task\-informed forecast supervision is sufficient on0/80/8seeds atK=16K=16and2/82/8atK=64K=64; a relaxed gate that inspects only the second gap and the place step is passed by3/83/8and5/85/8, but it tests only the two\-bit remainder after the grasp \(Appendix[B\.5](https://arxiv.org/html/2609.25757#A2.SS5)\)\. The two successfulK=64K=64seeds carry4\.094\.09and4\.214\.21bits in the first gap and attain closed\-loop success0\.450\.45and0\.460\.46\. Forecast supervision thus yields two realizations of the four\-bit representation when the codebook has slack, but not reliable learning at this scale; the privileged\-head control likewise reaches only1/81/8atK=128K=128\(Appendix[C\.3](https://arxiv.org/html/2609.25757#A3.SS3)\)\.

### D\.3Future\-behavior targets without observation side information

The decision\-centric targets of Table[23](https://arxiv.org/html/2609.25757#A6.T23)keep a history distinction when some future decision depends on it\. Their fixed\-demonstrator form is an open\-loop decoder that must predict the future action sequence from\(Ct,Ot\)\(C\_\{t\},O\_\{t\}\)alone, without the future observations and actions that the BFS decoder of §[3](https://arxiv.org/html/2609.25757#S3)receives; a world\-predictive form additionally predicts the future observations\. We train both on the gap toy withβ2\\beta\_\{2\}re\-revealed one step before use \(whereΓ\\Gammaneeds11bit in the first gap and00afterwards while the strong congruence keeps0\.670\.67bit\), on the corridor with driftW∈\{1,5\}W\\in\\\{1,5\\\}, and, for the action\-only form, on Task A and A′\. On the re\-reveal toy the action\-only target carries2\.062\.06bits in the first gap and1\.291\.29after use \(Γ\\Gamma:1\.001\.00and00\): it pays for the re\-revealed class because a future action depends on it, which is precisely the side\-information term of Theorem[1](https://arxiv.org/html/2609.25757#Thmtheorem1); the observation\-and\-action target carries1\.751\.75and0\.850\.85\. On the same toy, the side\-information decoder of §[3](https://arxiv.org/html/2609.25757#S3)\(−\-RF\) lands at1\.121\.12and0\.080\.08bits\. On the corridor withW=5W=5drift modes the two forms separate as the theory predicts: the action\-only target retains0\.010\.01bit about the drift, as DIACRITIC does \(0\.030\.03\), whereas the observation\-predictive target retains0\.850\.85bit, since the drift changes future observations but not future actions\. On the robot the action\-only target is sufficient on8/88/8seeds of every Task A dataset and of A′, but carries1\.611\.61–1\.861\.86bits during grasp and transport on Task A \(DIACRITIC:0\.030\.03\), with onlyI⁡\(C;mass∣O¯\)=0\.03I\(C;\\mathrm\{mass\}\\mid\\bar\{O\}\)=0\.03bit about mass\. It carries2\.162\.16bits after use on A′\(Γ\\Gamma:1\.001\.00\), with closed\-loop success0\.750\.75–0\.870\.87\. Thus, an open\-loop future decoder can charge the code for trajectory detail that future observations would supply\.

#### Label density\.

The forecast target need not be dense\. On A′gap 20 \(reveal→\\toplace distance3333\), applying the distillation loss only every second or fourth step after the last reveal \(1212and66supervised \(step, block\) pairs per episode instead of2323\) gives8/88/8and7/87/8full\-gate sufficiency, respectively, with full\-gate first\-gap rates2\.052\.05and2\.052\.05bits and all\-seed closed\-loop success0\.880\.88and0\.840\.84\(0\.920\.92with every step\)\. Behavior\-aligned long\-range supervision is therefore effective without per\-step labels\.

### D\.4Other training interventions

Table 17:Sufficiency and minimality on A′\(gap 6,N=1152N=1152unless noted\)\. Minimality is the post\-use rate against the1\.001\.00\-bit boundary\. The variational KL bound of the continuous model is not comparable to hard\-code entropy and is therefore not read as a minimality result\. Appendices[D\.4](https://arxiv.org/html/2609.25757#A4.SS4)and[C\.3](https://arxiv.org/html/2609.25757#A3.SS3)report additional variants\.We report training interventions that do not remove the barrier in §[5](https://arxiv.org/html/2609.25757#S5)\. Unless noted otherwise, all experiments use A′with 8 seeds\.*Future\-sufficiency\-driven code refinement*applies the solver’s partition\-refinement principle as a learning rule by splitting the code whose members have the least consistent futures\. It does not change the number of sufficient seeds at any gap on the toy or on A′\(Fisher exactp=1\.0p=1\.0in every cell; pooled coefficient\+0\.35\+0\.35,p=0\.28p=0\.28\)\. A*gap curriculum*warm\-starts gap 10/20 andM=16/32M=16/32from the gap\-6 scaffold model with the same seed\. It yields2/8,1/8,2/8,0/82/8,1/8,2/8,0/8and a pooled coefficient of−1\.6\-1\.6\(p=×10−3p=3\\\!\\times\\\!10^\{\-3\}\), indicating worse performance\.*Doubling the training budget*to 20k steps leaves gap 20 at0/80/8andM=16M=16at0/80/8\. On the toy, a*continuous Gaussian bottleneck*reaches sufficiency only atβ=0\.003\\beta=0\.003, with a KL bound of6\.96\.9bits and0\.320\.32bit of retained nuisance information; this bound is not comparable to hard\-code entropy \(Table[17](https://arxiv.org/html/2609.25757#A4.T17)\)\. A*non\-persistent code*reaches0/80/8\. An*unconditional prior*is indistinguishable from the conditional prior in every family:5/85/8versus5/85/8on A′, and8/8,6/8,4/88/8,6/8,4/8versus7/8,4/8,6/87/8,4/8,6/8in the corridor atM=4,16,64M=4,16,64\. Finally,*BFS*has a modest positive coefficient in the pooled regression \(\+0\.64\+0\.64,p=0\.02p=0\.02\) but does not remove the barrier\.

### D\.5Codebook size between the cardinality bound and learnability

On readout\-3 the exact requirement is three bits and Corollary[3](https://arxiv.org/html/2609.25757#Thmcorollary3)requiresK≥8K\\geq 8\. With task\-informed forecast supervision, sufficient\-seed counts increase monotonically across the tested codebooks:00,11,55,1515, and2323of2424runs \(three values ofMM, eight seeds each\) atK=8K=8,1010,1212,1616, and2424\. Sufficient seeds average2\.992\.99–3\.143\.14bits; all\-seed closed\-loop success is0\.070\.07,0\.100\.10, and0\.600\.60atK=10K=10,1212, and2424\. Some runs succeed with modest slack, andK=24K=24yields the highest observed reliability,23/2423/24\. The sweep distinguishes the representational state\-count bound from reliable acquisition under this optimizer and training configuration \(§[5](https://arxiv.org/html/2609.25757#S5)\); it does not establish that smaller feasible codebooks are unreachable\.

### D\.6Observed rate–fidelity trade\-offs

Figure 8:Empirical rate–distortion in the robot environment\(80 policies, 128 closed\-loop episodes each\)\. Left: post\-use code rate as a function of closed\-loop success among policies that reach sufficiency; many lie exactly atH⁡\(Γ∣O¯\)=1H\(\\Gamma\\mid\\bar\{O\}\)=1\. Right: success of sufficient policies as a function ofβ\\beta\.As the post\-use rate approachesH⁡\(Γ∣O¯\)=1H\(\\Gamma\\mid\\bar\{O\}\)=1, closed\-loop success for sufficient plain−\-R policies decreases from0\.960\.96–1\.001\.00atβ≤×10−4\\beta\\leq 3\\\!\\times\\\!10^\{\-4\}to0\.700\.70–0\.900\.90atβ=×10−3\\beta=2\\\!\\times\\\!10^\{\-3\}\. This trade\-off motivates fixingβ\\betaunder aD≈0D\\approx 0constraint before inspecting any rate \(§[4\.1](https://arxiv.org/html/2609.25757#S4.SS1)\)\. Across 448 models evaluated in closed loop, teacher\-forced and own\-occupancy sufficiency agree on31/3231/32members of the reference set\. Sufficient policies succeed on0\.700\.70–1\.001\.00of episodes, compared with0\.120\.12–0\.450\.45for insufficient policies\. The full pooled regression from §[5](https://arxiv.org/html/2609.25757#S5), including three specifications, run\- and configuration\-clustered errors, and cluster bootstraps, and the corridor regression are reproduced by the released code\.

### D\.7Sufficiency counts and temporal\-distance statistics

Sufficiency counts are reported per 8\- or 16\-seed cell\. Wilson 95% intervals for 8 seeds are\[0,0\.32\]\[0,0\.32\]for0/80/8,\[0\.22,0\.79\]\[0\.22,0\.79\]for4/84/8and\[0\.68,1\]\[0\.68,1\]for8/88/8\. Among 45 pairwise contrasts in the refinement, curriculum, coverage,β\\beta, and BFS sweeps, only two reach Fisherp<0\.05p<0\.05, both involving the scaffold at gap 6 \(5/85/8vs0/80/8\)\. The remaining differences in those sweeps are not established; this comparison set does not include the future\-behavior supervision experiments\.

The controlled distance sweeps of §[5](https://arxiv.org/html/2609.25757#S5)use 16 seeds per cell for gaps6/10/206/10/20\(distances 19/23/33 for the place class\)\. Their outcome is class\-specific sufficiency at use: placement counts are1,3,11,3,1of 16 for−\-R,5,2,05,2,0for−\-RF and10,7,010,7,0for scaffold\-RF, with grasp\-class counts of4242–4545of 48 at distance 3\. These are not the full\-trajectory gate counts used for the frozen supervision comparison\. The gap1sweeps \(distances 11 and 21\) have 8 seeds per cell\. Per\-learner logistic slopes, their cluster\-robust intervals, the interaction test and threshold sensitivity \(0\.80\.8and0\.950\.95leave every count unchanged except one grasp cell\) are in the released analysis code\.

### D\.8Large hidden\-mode counts on P\-I

P\-I on A′\.Figure[9](https://arxiv.org/html/2609.25757#A4.F9)presents the exact\-regime counterpart of the world\-complexity comparison\. The probe channel reveals the fullθ∈\[M\]\\theta\\in\[M\], while each phase contains only two behavioral classes\. ForMMup to 32, every seed retains a learned rate of11–22bits; successful seeds carry exactly2\.002\.00bits, and the slope with respect tolog2⁡M\\log\_\{2\}Mis00\. On the same data, the system\-identification objective increases to33bits, with a slope of\+0\.89\+0\.89per bit oflog2⁡M\\log\_\{2\}M, and loses the behavioral memory\. The unsupervised variants reach0/80/8sufficient seeds atM≥16M\\geq 16, illustrating the optimization barrier in §[5](https://arxiv.org/html/2609.25757#S5)when many bits are revealed but few are behaviorally relevant; forecast supervision is the exception reported there\.

P\-I atM=128M=128and512512with forecast supervision\.We freeze the forecast\-distilled configuration selected atM≤32M\\leq 32\(−\-R,β=10−3\\beta=10^\{\-3\},K=16K=16,N=448N=448; stratified data\)\. On the training recordings, the all\-seed first\-gap means are2\.362\.36and2\.632\.63bits atM=128M=128and512512; among late\-gate sufficient seeds they are2\.172\.17and2\.242\.24bits \(theory2\.002\.00\)\. The corresponding all\-seed means on fresh recordings are2\.482\.48and2\.722\.72\.

Same\-task system identification stores3\.43\.4–3\.53\.5bits atK=16K=16and6\.456\.45and8\.728\.72bits atK=256K=256and10241024\(4\.894\.89and7\.727\.72bits aboutθ\\theta\), is sufficient on0/80/8seeds, and succeeds in closed loop on at most0\.180\.18of episodes\.

Reliability nevertheless degrades withMM: late\-phase sufficiency is8/88/8,6/86/8,5/85/8and3/83/8atM=16M=16,3232,128128and512512\(full\-trajectory4/84/8and3/83/8at the last two\)\. Some verdicts change on fresh recordings, including5→35\\to 3atM=128M=128\. Training on all 512 modes \(N=1000N=1000\) yields only2/82/8late\-gate and0/80/8full\-trajectory seeds on the training recordings, and0/80/8late\-gate seeds on fresh recordings\. Thus, conditional on successful realization, the behavioral rate remains near two bits asMMgrows, but the all\-seed mean rises modestly and reliability deteriorates\. We interpret this as a rate separation from world identification and as an optimization scale boundary, not as reliable flat scaling\.

Figure 9:A′with the fullθ\\thetarevealed\. ForM≤32M\\leq 32, successful seeds remain at 2\.00 bits, the all\-seed average is flat over the plotted range, and the system\-identification rate increases withMM\. Large\-MMresults and their reliability boundary are reported in the text\.

## Appendix EAdditional domains and limits of transfer

These domains test parts of the information structure and the learning surrogate beyond the main comparisons\. Offline sufficiency, memory\-dependent decisions, and complete closed\-loop task success are reported separately\.

### E\.1Pixel observations and observation\-timing checks

#### Setup\.

We recollect A′at gap 6 andN=512N=512using a separate1282128^\{2\}camera for each environment\. The trainer crops the table region, downsamples it to64264^\{2\}, and replacesgog\_\{o\}with a three\-layer CNN\. Its output is concatenated with the low\-dimensional gripper opening, probe tag, and phase\. Because the tag is not visible in the image, the identification channel remains instrumented\. All other components, including the codes, prior, rate, metrics,β=2×10−3\\beta=2\\times 10^\{\-3\}, and 10k training steps, remain unchanged; BFS is not used\.

#### Results\.

\(Forecast supervision from pixels, at gaps 6 and 20, is reported in Appendix[D\.1](https://arxiv.org/html/2609.25757#A4.SS1), Table[15](https://arxiv.org/html/2609.25757#A4.T15); this paragraph concerns plain training\.\) DIACRITIC\-R recovers the exact staircase from pixels on 2 of 8 seeds\. Its gap\-1 rate is1\.991\.99relative toH⁡\(Γ\|O¯\)=1\.99H\(\\Gamma\|\\bar\{O\}\)=1\.99, and its gap\-2 rate is1\.001\.00relative to1\.001\.00\. Both gaps haveSΓ=1\.00S\_\{\\Gamma\}=1\.00; the place phase hasSG=1\.00S\_\{G\}=1\.00;I^​\(C;mass\|O¯\)=0\.01\\hat\{I\}\(C;\\text\{mass\}\|\\bar\{O\}\)=0\.01; and the test MSE is0\.0030\.003\. The remaining seeds fail to recover the place class, with error0\.380\.38–0\.530\.53at the first place\-divergent step\. This outcome matches the discovery\-limited behavior observed with state inputs at the sameNN, where 0–3 of 8 seeds succeed\. The annealed scaffold reaches 0 of 8 seeds from pixels\. Anticipatory memory can therefore emerge from pixel observations at exactly the predicted rate when optimization succeeds; the supervised long\-horizon pixel results are reported in Appendix[D\.1](https://arxiv.org/html/2609.25757#A4.SS1)\.

#### Rate predictions as benchmark diagnostics\.

In the initial pixel collection, each frame was captured*after*the macro step, whereas the state observation was recorded before it\. The diagnostics exposed the offset: every seed had a gap\-2 rate of00andSΓ=0S\_\{\\Gamma\}=0, yet predicted the first memory\-dependent place action with error0\.000\.00\. This mismatch prompted a check of the observation convention used by the theory\. A single\-frame probe, implemented as a CNN that predicts the hidden bits from individual frames, showed that the place class was readable at1\.001\.00from “place” frames and at chance from gap frames, identifying a one\-step capture offset\. After correction, the probe predicts the place class at0\.450\.45, compared with a majority baseline of0\.520\.52, at the first place\-divergent step; it predictsθ\\thetaat0\.030\.03\(1/321/32\-level\) and the mass quartile at0\.290\.29\(0\.330\.33\)\. This case illustrates the importance of direct leakage probes in visual imitation benchmarks with hidden state: the rate mismatch prompted the probe that identified the timing error, despite apparently correct behavior\.

Closed\-loop evaluation from pixels\.At every step, the environment supplies the pixel policy with the corresponding camera frame at1282128^\{2\}, cropped to the table and resized to64264^\{2\}as during training\. We evaluate 16 models, comprising 8 seeds×\\times\{−\-R, scaffold\-R\}, on 128 episodes per model and compute all metrics under the policy’s own occupancy\. The two seeds that are sufficient under teacher forcing remain sufficient under their own occupancy, withSΓ=1\.00S\_\{\\Gamma\}=1\.00and0\.950\.95in gap2\. They carry1\.941\.94bits in gap1, choose the side with accuracy1\.001\.00, choose the slot on0\.920\.92/0\.940\.94of episodes, and achieve success on0\.660\.66/0\.360\.36with mean action error3\.13\.1/5\.75\.7cm\. The fourteen insufficient models choose the slot at chance \(0\.410\.41–0\.660\.66\), achieve success of0\.090\.09–0\.420\.42, and have action errors of77–1212cm\. Symbolic sufficiency therefore predicts closed\-loop memory use from pixel observations as it does from state observations\. The primary cost of pixels is lower action precision: state policies reach0\.50\.5cm, which explains the lower success ceiling of the pixel policies\. One insufficient model exhibits a5\.75\.7cm class\-conditional pose separation in gap2without above\-chance slot selection, indicating body memory that is written but not read\.

Table 18:Closed\-loop evaluation of the 16 pixel policies \(128 episodes each;SΓS\_\{\\Gamma\}under the policy’s own occupancy\)\.

### E\.2Cue–corridor T\-maze

We use the cue–corridor–junction T\-maze from the RL memory literature\([Ni et al\., 2023](https://arxiv.org/html/2609.25757#bib.bib28)\)to verify that exactness does not depend on an instrumented identification phase\. The ordinary first observation displays a cuec∈\[R\]c\\in\[R\]together with an exogenous texturezz\. The agent then traversesLLidentical corridor cells and reaches a junction, where the expert turns according to the cue\. The task contains no instrumented channel\. Assumption \(A4\) holds automatically because the cue affects only the expert’s action, and the solver certifies \(A2\), \(A4\), and transitivity for everyLLup to 80\. Throughout the corridor,H⁡\(Γt∣O¯t\)=log2⁡RH\(\\Gamma\_\{t\}\\mid\\bar\{O\}\_\{t\}\)=\\log\_\{2\}R, whileH⁡\(GE,t∣O¯t\)=0H\(G\_\{E,t\}\\mid\\bar\{O\}\_\{t\}\)=0until the junction\. We use the corridor configuration: raw VQ, atanh\\tanh\-bounded state,β=0\.01\\beta=0\.01,K=16K=16, and 6k steps, with 8 seeds per cell\. A seed is sufficient iffSΓ\>0\.9S\_\{\\Gamma\}\>0\.9at every corridor step and its junction action is exact\.

Table 19:T\-maze: seeds sufficient / corridor rate of sufficient seeds \(bits; theorylog2⁡R\\log\_\{2\}R\)\.Every sufficient seed carries at least the predictedlog2⁡R\\log\_\{2\}Rbits through the corridor and00after the junction\. Most seeds carry exactlylog2⁡R\\log\_\{2\}R, while several retain an additional0\.10\.1–0\.50\.5bit: the range is1\.001\.00–1\.501\.50for a one\-bit cue and exactly2\.002\.00for a two\-bit cue\. Thus, the staircase in §[4\.3](https://arxiv.org/html/2609.25757#S4.SS3)does not require a probe channel\. The learning barrier remains present but is weaker and less regular than on A′\. A single cue bit is retained reliably throughL=20L=20and by33–66of 8 seeds atL=40L=40–8080; the pooled slope is−0\.03\-0\.03/step and is non\-monotone inLL\. With two cue bits, the sufficient\-seed count decreases monotonically as6,4,1,0,06,4,1,0,0of 8 atL=5,10,20,40,80L=5,10,20,40,80, with a slope of−0\.21\-0\.21/step, and every sufficient seed carries exactly2\.002\.00bits\. In this task, the commitment cost therefore increases with both the number of retained bits and the distance, consistent with the ordering observed between A′and the corridor\. We present this as an observation within one task family rather than as an additional distance law, particularly because the one\-bit curve is non\-monotone\. The scaffold again removes most of the one\-bit degradation up toL=40L=40\.

### E\.3Physical weighing and re\-observation

#### Task\.

The A′and readout tasks reveal the hidden class through a probe channel and attach the object kinematically\. Here the hidden variable is the object’s mass \(44classes,0\.10\.1to2\.52\.5kg\) and nothing displays it\. The arm grasps and lifts the object \(a physical grasp throughout\), holds it for four steps, puts it back and releases it; under load the arm sags, and the tool height differs by7\.57\.5–1515mm between adjacent classes against22mm sensor noise\. After a gap of66or2020steps during which the arm hovers, the expert grasps from the side assigned to the mass class \(22bits\) and places the object in the slot of the coarser class; the transport sag re\-reveals the mass\. The symbolic observation shows the sag class at the steps where the class\-conditional mean heights are pairwise separated by more than66mm \(the data\-driven visibility convention of Appendix[B\.1](https://arxiv.org/html/2609.25757#A2.SS1)\)\. Each dataset has10241024stratified episodes \(expert success1024/10241024/1024\), and an independent recording provides the replication set\.

#### Benchmark hygiene\.

In the first recording a probe decoded the mass class from the gap observations with accuracy0\.8850\.885: the loaded arm sets the object down11–1111mm forward and slightly yawed, so the world remembered the class for the policy\. We therefore return the released object to a fixture that re\-seats it \(held at its spawn pose33mm above the table during the last identification step; a plain pose write is dragged back by the simulator’s cached friction anchors\)\. Afterwards the probe is at chance in the gap \(0\.250±0\.0100\.250\\pm 0\.010and0\.244±0\.0090\.244\\pm 0\.009; chance0\.250\.25\) and at0\.980\.98during the weighing; the first recording is not used\.

#### Exact requirement\.

The solver certifies transitivity at every step, and \(A4\) fails at the single step at which the second lift begins\.H⁡\(Γt∣O¯t\)=2H\(\\Gamma\_\{t\}\\mid\\bar\{O\}\_\{t\}\)=2bits from the put\-down to the grasp decision and00afterwards, whileH⁡\(Γts∣O¯t\)H\(\\Gamma^\{s\}\_\{t\}\\mid\\bar\{O\}\_\{t\}\)returns to22bits during the second descent: the imminent sag separates the observation supports, so histories that differ in mass are vacuously compatible\. Expiration is here produced by physical re\-revelation rather than by the end of use\.

#### Learning\.

The gate isSΓ\>0\.9S\_\{\\Gamma\}\>0\.9at every gap step andSG\>0\.9S\_\{G\}\>0\.9at the grasp decision\. Plain training is sufficient on1/81/8and0/80/8seeds \(gap 6, 20\), the task\-informed forecast supervised from the last identification step on3/83/8and2/82/8, and system identification \(class labels at every step\) on8/88/8and7/87/8\. The observed failures occur at memory acquisition:SΓS\_\{\\Gamma\}is constant from the put\-down through the gap to the grasp, typically0\.750\.75with1\.51\.5bits, with two adjacent mass classes sharing a code\. The measured distinction is already incomplete at the gap’s start, with no further loss detected during retention\. Supervising the same forecast from the first step of the episode, which also drops the reveal\-time annotation from the objective, gives8/88/8and7/87/8\(N=448N=448\) and7/87/8and8/88/8\(N=896N=896\), with2\.022\.02–2\.052\.05bits in the gap and identical counts on the independent recordings\. The event\-agnostic objective does not repair acquisition here \(2/82/8,0/80/8; annealed0/80/8\)\.

#### Closed loop\.

No tested compact\-carrier learner solves the task reliably in closed loop \(Table[20](https://arxiv.org/html/2609.25757#A5.T20)\), although full\-history policies do\. A rollout of an offline\-sufficient policy illustrates one failure mode: it executes the weighing itself imprecisely \(the lift reaches100100–165165mm instead of about200200mm and the object is half grasped; action error1111cm\), so no clean sag is produced and the code is written from an observation distribution it never saw\. The auxiliary objectives are associated with worse imitation precision inside the weighing block: the held\-out action error there is0\.0780\.078for the forecast supervised from step 0 and0\.0570\.057for system identification, against0\.0020\.002–0\.0040\.004for plain training and0\.0010\.001for the full\-history policies\. Annealing the supervision lowers it only to0\.0130\.013–0\.0310\.031, costs sufficiency \(4/84/8at gap 6\) and leaves success at0\.080\.08–0\.130\.13\. Two pre\-specified controls examine these limitations\.

*Hybrid rollout:*the scripted expert executes the approach and the weighing while the learned policy observes, updates its own memory with the executed actions, and acts alone from the gap onward\. For the offline\-sufficient policies supervised from step 0 the grasp\-side accuracy rises from0\.300\.30to0\.990\.99at gap 6 \(8 seeds\) and from0\.320\.32to0\.930\.93at gap 20 \(7 seeds; prediction: above0\.80\.8\): the code is written, kept across the gap and read correctly in closed loop once the observation is produced correctly\. Task success rises only to0\.380\.38and0\.130\.13, because the subsequent physical grasp and transport of objects of up to2\.52\.5kg remain imprecise\.

*Longer hold:*with an 8\-step hold the analog read\-out becomes easier \(plain training4/84/8instead of1/81/8; supervised from step 08/88/8\) while closed\-loop success stays low \(0\.200\.20,0\.090\.09\)\. These controls support retention and use of the measured grasp\-side distinction after expert\-led weighing\. Producing informative observations and executing the subsequent physical manipulation remain limitations; the controls do not establish sufficiency of the learned state for complete continuous control\.

Table 20:Weighing task, 8 seeds per cell: sufficient seeds on the training recording / on the independent recording, and closed\-loop success / grasp\-side accuracy over all seeds \(chance0\.250\.25\)\.

### E\.4External Passive T\-maze

We take the Passive T\-maze of[Ni et al\. \(2023\)](https://arxiv.org/html/2609.25757#bib.bib28)without modification from the MIKASA\-Base suite\([Cherepanov et al\., 2026](https://arxiv.org/html/2609.25757#bib.bib38)\): the goal cue appears in the first observation only, the position is ambiguous along the corridor, and the decision is takenLLsteps later\. We clone a scripted oracle \(move rightLLtimes, then turn according to the cue\) from256256recorded episodes and roll the learned policy out in the external environment \(100100episodes; success = goal reward\)\. No rate is compared with theory; the question is whether the supervision of Appendix[D\.1](https://arxiv.org/html/2609.25757#A4.SS1)changes closed\-loop success on a task we did not design\. The compact carrier is the unchangedK=16K=16model \(β=10−3\\beta=10^\{\-3\},50005000steps\); the event\-agnostic head predicts the recorded expert action at a random future offset \(the task is deterministic given the cue, so no forecaster is needed and no annealing is applied\)\. A standard GRU policy with a continuous128128\-dimensional state and no bottleneck is the reference\.

Table 21:Passive T\-maze: seeds \(of 8\) that solve the task in closed loop \(success≥0\.99\\geq 0\.99\); mean success in parentheses\.The128128\-dimensional GRU is reliable throughL=50L=50, whereas the compact carrier benefits from event\-agnostic supervision but still fails atL=100L=100\. Doubling the compact learner’s training budget does not improve itsL=50L=50result \(3/83/8\): runs either capture the cue early and retain it to the junction or collapse to a cue\-free code\. Neither the future\-conditioned−\-RF decoder nor the annealed scaffold extends the horizon in these tested configurations\.

The GRUs with44\- or88\-dimensional states fail at all three tested lengths,L=20,50,100L=20,50,100\. These implementations shrink the input network and action head together with the recurrent state, and their information capacity is not matched to the discrete carrier\. They show that training difficulties also occur in small continuous RNNs; they do not isolate the effects of state size, network size, optimization, or quantization\.

The event\-agnostic objective need not identify the cue when candidate objects are rearranged after the delay: a future action can then be uninformative about the cue without the future observation\. A future\-observation\-conditioned decoder, such as−\-RF, can supply that side information\. We make no architectural claim against the unconstrained GRU\.

### E\.5Certified requirements on community benchmarks

We apply the protocol to two community benchmarks using their environment code unmodified from MIKASA\-Base\([Cherepanov et al\., 2026](https://arxiv.org/html/2609.25757#bib.bib38)\)\. In the bsuite memory chain \(“MemoryLength”\), abb\-bit context is displayed initially, then hidden until a query index appears; the final action repeats the queried bit\. The second task is the Passive T\-maze of Appendix[E\.4](https://arxiv.org/html/2609.25757#A5.SS4)\.

#### Certification and timing\.

We record the external environments’ own observations and oracle actions, retain a trajectory for every hidden value \(2b​b2^\{b\}bcontext–query pairs; two T\-maze goals\), and run the solver on the induced finite POMDP\. It certifies \(A2\), \(A4\), and transitivity on every instance\. With uniform latent weights, the memory\-chain requirement is zero in the first two steps, while the cue is visible\. It becomesbbbits at the first cue\-free step \(step 3\), remainsbbuntil the query, and falls to one bit once the query index is observed\. Only the queried bit then matters\. The requirement is independent ofmemory\_length;H⁡\(GE,t∣Ot\)=0H\(G\_\{E,t\}\\mid O\_\{t\}\)=0until the final decision\. The T\-maze requires one bit during its cue\-free corridor\.

#### Learning and rate convention\.

We clone the oracle with the unchangedK=16K=16sole carrier \(β=10−3\\beta=10^\{\-3\}, 5000 steps, 512 demonstrations\), using plain imitation or the event\-agnostic objective \(a recorded future action at a random offset; no annealing\)\. A seed is sufficient whenSΓ\>0\.9S\_\{\\Gamma\}\>0\.9at every positive\-requirement step and closed\-loop success is at least0\.990\.99in the external environment \(200 episodes\)\. The pre\-specified prediction concerns*mid\-delay*: sufficient codes should match the requirement there within0\.10\.1bit\. Table[22](https://arxiv.org/html/2609.25757#A5.T22)uses the empirical demonstration occupancy for both entropies, giving1\.991\.99and2\.992\.99bits instead of the uniform\-law values22and33\.

Figure 10:bsuite memory chain: mid\-delay rates and learning reliability\.At mid\-delay the certified requirement \(line\) is independent of delay, and every sufficient seed \(dots; both learners\) matches it\. Bars show the number of sufficient seeds \(right axis\)\. This comparison concerns one point in the waiting period; it does not establish minimal rates after the query appears\.Table 22:Certified requirement and learned rate on community benchmarks \(8 seeds per cell\)\. The certified value is the empirical requirement at mid\-delay under the recorded occupancy \(512512episodes, hence1\.991\.99and2\.992\.99\); the learned rate isH^​\(Ct∣Ot\)\\hat\{H\}\(C\_\{t\}\\mid O\_\{t\}\)at the same step over the sufficient seeds of both learners\.All8787sufficient runs \(of192192\) match the certified rate at mid\-delay \(largest measured deviation0\.000\.00bit\)\. This agreement is strongly constrained by the task structure\. For a fixed trained model, the mid\-delay code is a deterministic function of the context; no other episode\-dependent distinction has been observed, and the preceding oracle actions are fixed\. Under the uniform law,H⁡\(Ct∣Ot\)≤bH\(C\_\{t\}\\mid O\_\{t\}\)\\leq b, while exact sufficiency requiresH⁡\(Ct∣Ot\)≥bH\(C\_\{t\}\\mid O\_\{t\}\)\\geq b\. Under empirical occupancy the same argument uses the empirical context entropy\. These benchmarks test transfer of the protocol and acquisition of the required information; Task A and readout provide the comparisons with additional world information\.

#### Acquisition without immediate forgetting\.

Mid\-delay agreement does not extend to the query step\. Once the query index appears, the requirement is about one bit under empirical occupancy\. The2929sufficient one\-bit memory\-chain runs carry0\.9990\.999bit there\. The1919sufficient two\-bit runs retain1\.5141\.514–1\.9901\.990bits \(mean1\.911\.91\), and the1717sufficient three\-bit runs retain2\.3852\.385–2\.9762\.976bits \(mean2\.692\.69\)\. Thus the models can acquire all required context while retaining distinctions that the query has made unnecessary\. Like the surplus under un\-annealed prediction on A′\(§[5](https://arxiv.org/html/2609.25757#S5)\), this separates acquisition from subsequent compression; it does not identify a shared optimization mechanism\.

Learning reliability also decreases with delay \(§[5](https://arxiv.org/html/2609.25757#S5)\): no seed passes for two or three bits at delay100100, and the event\-agnostic objective is at least as reliable as plain imitation in every tested cell\. The scope remains finite symbolic tasks with enumerable hidden variables and the environment’s own observations\.

## Appendix FRepresentation targets of related work

#### Scope of the empirical comparison\.

We compare representation*targets*under one solver rather than re\-implementing the methods that pursue them\. Multi\-step inverse kinematics and agent\-centric state discovery\([Mhammedi et al\., 2023](https://arxiv.org/html/2609.25757#bib.bib21);[Lamb et al\., 2023](https://arxiv.org/html/2609.25757#bib.bib16)\)target control\-endogenous state, and decision\-centric memory compression\([Zou et al\., 2026](https://arxiv.org/html/2609.25757#bib.bib34);[Yamin et al\., 2026](https://arxiv.org/html/2609.25757#bib.bib27)\)targets history distinctions that change a near\-optimal decision; neither is defined for a fixed demonstrator with future observations as decoder side information\. Their objectives must therefore be re\-specified before their rates can be compared with ours, while their exploration, reachability, or language\-model mechanisms would introduce additional implementation differences\. We implement the fixed\-demonstrator forms of the theoretically distinct targets and measure them with the same solver \(last column\); targets not implemented are marked accordingly\.

Table 23:Representation targets in related work \(§[6](https://arxiv.org/html/2609.25757#S6)\)\. Property columns:Fthe target reproduces a fixed demonstrator’s action distribution;Scontinuations are compared over stochastic reachable supports, so future observations act as decoder side information;Eminimality is defined through occupancy\-weighted conditional entropy;Bthe general case yields a non\-transitive compatibility bracket\. Individual properties appear in prior work; the contribution here is their combination\.

相似文章

量化逆强化学习中潜在观测缺失问题

arXiv cs.LG

本文识别了逆强化学习(IRL)中观测缺失的问题,该问题可能导致专家行为看似次优,并提出了一种实用算法,用于量化使专家行为显得最优所需的最小扰动,并在合成任务、癌症治疗模拟和ICU数据上进行了验证。

信念记忆:部分可观测性下的智能体记忆

arXiv cs.AI

本文介绍了 BeliefMem,一种专为大语言模型(LLM)智能体设计的新型记忆范式。该范式通过存储带有概率的多个候选结论来处理部分可观测性问题,并减少自我强化错误。在 LoCoMo 和 ALFWorld 基准测试中的实证评估显示,该方法优于确定性基线模型。