The LLM Proposes, the Executive Disposes: A Self-Verifying Agent Instrument that Dissociates Commitment Drift from Binding Drift in Long-Horizon Agents

arXiv cs.AI 论文

摘要

Presents a verification instrument for long-horizon agents that structurally separates commitment drift from binding drift, using a deterministic executive and pre-registered predictions. Reports ablation results showing commitment mechanism removal flips goal abandonment from 0 to 1 while binding error stays flat, though task efficacy is null on ARC-AGI-3.

arXiv:2608.04066v1 Announce Type: new Abstract: How do you verify a long-horizon agent when its own state and self-reports are exactly what you cannot trust? We present an agent instrument built so that verification is structural rather than post-hoc. A deterministic Executive owns all belief; a language model may only file typed proposals, and a claim is admitted only when a prediction pre-registered before acting is matched against observation by code. Two properties make the instrument a verifier of its own science, not just of the agent: every run invalidates itself when per-organ write-error, render-size, or salted-canary-echo floors are breached (four of the first eight architecture runs were invalidated, each localizing a real defect); and a render-invisible shadow reference compiles the plan the full system would have committed in every ablation cell, so drift metrics are defined even where the mechanism under test has been removed. Using this instrument we report a clean, single-variable result on a failure every long-horizon agent suffers: ablating the commitment mechanism flips goal-abandonment from 0.00 to 1.00 while binding error stays flat at 0.00 (three seeds per cell, up to 394 reference beats per run, every run gated valid). The binding channel, by contrast, does not reappear as per-beat drift when its repair is ablated -- because binding is code-owned, the failure class is structurally absorbed, its only residue appearing one layer upstream as a collapse in hypothesis formation. We report these under full disclosure that task efficacy is null (zero level completions across 52 gated runs on ARC-AGI-3), pre-registered as a structural defeater. The contribution is a verification methodology for agent development and the drift decomposition it makes measurable.
查看原文
查看缓存全文

缓存时间: 2026/08/06 07:41

# The LLM Proposes, the Executive Disposes: A Self-Verifying Agent Instrument that Dissociates Commitment Drift from Binding Drift in Long-Horizon Agents
Source: [https://arxiv.org/html/2608.04066](https://arxiv.org/html/2608.04066)
Mohsen Arjmandi Independent researcher mohsen\.arjmandi@gmail\.com [github\.com/arjmandi/ARG](https://github.com/arjmandi/ARG)\(tag v1\.0\-paper\)

###### Abstract

How do you verify a long\-horizon agent when its own state and self\-reports are exactly what you cannot trust? We present an agent instrument built so that verification is structural rather than post\-hoc\. A deterministic*Executive*owns all belief; a language model may only file typed proposals, and a claim is admitted only when a prediction*pre\-registered before acting*is matched against observation by code\. Two properties make the instrument a verifier of its own science, not just of the agent: every run*invalidates itself*when per\-organ write\-error, render\-size, or salted\-canary\-echo floors are breached \(four of the first eight architecture runs were invalidated, each localizing a real defect\); and a render\-invisible*shadow reference*compiles the plan the full system would have committed*in every ablation cell*, so drift metrics are defined even where the mechanism under test has been removed\. Using this instrument we report a clean, single\-variable result on a failure every long\-horizon agent suffers: ablating the commitment mechanism flips goal\-abandonment from0\.000\.00to1\.001\.00while binding error stays flat at0\.000\.00\(three seeds per cell, up to 394 reference beats per run, every run gated valid\)\. The binding channel, by contrast, does not reappear as per\-beat drift when its repair is ablated — because binding is code\-owned, the failure class is*structurally absorbed*, its only residue appearing one layer upstream as a collapse in hypothesis formation\. We report these under full disclosure that task efficacy is null \(zero level completions across 52 gated runs on ARC\-AGI\-3\), pre\-registered as a structural defeater\. The contribution is a verification methodology for agent development and the drift decomposition it makes measurable\.

## 1Introduction

An agent that acts over hundreds or thousands of steps drifts off its own objective\. The phenomenon —*goal drift*— is now measured and benchmarked as a first\-class failure of language\-model agents\[[2](https://arxiv.org/html/2608.04066#bib.bib1),[8](https://arxiv.org/html/2608.04066#bib.bib2)\]\. But the way we*verify*it is weak in two ways that matter for reliable agent development\. First, most evaluations trust the agent’s own state and self\-reports; an agent that says “done” is scored as done\. Second, drift is reported as an aggregate score that cannot say*which*failure occurred, or whether a given run was even valid to attribute a number to\. This paper takes the workshop’s question —*who verifies the agent?*— literally, and answers it by construction\.

We describe an instrument in which verification is not a downstream check but the substrate\. Belief is earned, not asserted: a claim enters the agent’s state only when a prediction committed to a log*before*acting is matched against the environment’s response by deterministic code\. The agent cannot author a match it did not pre\-commit to, and “done” is never an event\. On top of this we add two verifiers of the*measurement*: a run\-validity gate that makes a run invalidate itself when the machinery is misbehaving, and a shadow instrument that defines the drift metrics in every experimental cell, including the cells where the mechanism under test has been deleted\.

That instrument buys a result\. “Drift” is not one failure but at least two, with different signatures and different repairs:

- •Binding drift— the intention and its referents are both present, but the join between “the goal” and “the object it is about” lives only in transformer attention, which loses bindings under distance and low salience \(the positional / spaced\-evidence regime;[9](https://arxiv.org/html/2608.04066#bib.bib3),[7](https://arxiv.org/html/2608.04066#bib.bib4)\)\. The repair is*positional*\.
- •Commitment drift— the intention is simply absent from later contexts, and error compounds even when every context is short and the goal maximally salient \(the goal\-abandonment regime measured in the agent literature;[2](https://arxiv.org/html/2608.04066#bib.bib1),[8](https://arxiv.org/html/2608.04066#bib.bib2)\)\. The repair is*not*positional: an external commitment store executed by code\.

The binding component is not merely hypothesized: prior work on a predecessor system\[[3](https://arxiv.org/html/2608.04066#bib.bib5)\]diagnosed a*self\-consistent hallucination cascade*in which perception\-layer errors propagate into internally coherent but factually wrong world models\. What that work names a perception\-layer cascade is, in the present taxonomy, binding loss — a referent grounded in the model’s own prior text rather than in observation\. Our instrument externalizes a repair for each failure as an independent switch, so the two can be turned off, and measured, separately\.

#### Contributions\.

1. 1\.A self\-verifying agent instrument\(§[2](https://arxiv.org/html/2608.04066#S2)–§[3](https://arxiv.org/html/2608.04066#S3)\): a deterministic control plane that owns belief; runs that invalidate themselves on per\-organ error, render\-size, and canary\-echo floors; and a render\-invisible shadow reference that defines drift in every ablation cell against a byte\-identity guarantee\.
2. 2\.A causal dissociation\(§[4](https://arxiv.org/html/2608.04066#S4)\): ablating the commitment mechanism drives goal\-abandonment to the ceiling \(0\.00→1\.000\.00\\\!\\to\\\!1\.00\) with binding error flat, while ablating the binding repair changes nothing per\-beat — evidence that “drift” should be decomposed before it is repaired\.
3. 3\.The null, reported first\(§[5](https://arxiv.org/html/2608.04066#S5)\): zero task completions across 52 gated runs, pre\-registered as a defeater — and a demonstration that the mechanism results hold independently of it\.

## 2The verification architecture

#### The LLM proposes, the Executive disposes\.

No model output mutates state and no model reads a raw dump\. Three model “organs” — an Observer that interprets machine\-computed change\-sets, a Surveyor that files goals / hypotheses / experiments, and an Actuator that emits actions — may only submit typed proposals in closed vocabularies through admission gates\. A deterministicExecutive\(differ, matcher, validator, compiler, renderer, test\-evaluator\) owns every state transition\. All stores are append\-only; status is*computed*, never overwritten; every consequence receipt keys to the exact per\-frame anchor under which it was observed, so identity is replayable data rather than a destructive merge\.

#### Belief is earned by verification\.

A referent climbs tiers, each transition machine\-checked: named\-only text is not yet a referent; a salience\-blind extraction from raw observation anchors it; a logged action receipt engages it; and it becomes*characterized*only when referenced by a rule that istested— i\.e\., whose predictions, committed to the log*before*acting, were matched against the deterministic observation diff by the Executive\. This is verification in the strict sense: the model cannot author a match it did not pre\-commit to, achievement is a fired predicate over logged events \(“an LLM saying done is not an event”\), and demotion never deletes the audit chain\.

#### The two switches\.

The binding repairJis the goal↔\\leftrightarrowreferent*join*: a persistent structure that canonicalizes referents from consequence\-tested evidence and re\-surfaces every active goal pre\-joined with its referents at fixed prompt positions each turn, so no consumer joins facts across distant prompt regions\. The commitment repairAis the external commitment store executed by code: a compiled plan the Actuator consults independent of model attention\. Each is a single environment flag \(ARG\_JOIN=0,ARG\_AGENDA=0\); all other scaffolding is identical across cells, so a cell\-to\-cell delta cannot be attributed to generic scaffolding or to call count\. Nothing above a small, declarative environment adapter is task\-specific — the property that makes the “unseen game” result \(§[5](https://arxiv.org/html/2608.04066#S5)\) meaningful, enforced at release by an exemplar scrub\.

## 3Verifying the measurement itself

A dissociation is only as trustworthy as the instrument that measures it\. Three mechanisms verify the measurement before any belief\-level claim is permitted\.

#### Runs that invalidate themselves\.

A run isinvalidfor attribution — excluded entirely — if any per\-organ write\-error rate exceeds0\.250\.25\(typed rejection rows over total ops\), if any rendered view exceeds the hard token ceiling per beat*or*per call, or if salted\-canary echo rates per prompt zone fall below floor\. Canaries are quarantined synthetic rows injected into every zone; per\-zone echo separates transport defects from consumption rot before any belief\-level diagnosis\. This is not decorative: four of the first eight architecture runs were invalidated by these floors, and each invalidation localized a real contract defect\. Every number in §[4](https://arxiv.org/html/2608.04066#S4)is from a valid run and is replayable from an append\-only log\.

#### Drift defined in every cell\.

Goal drift is deviation from a plan — but in commitment\-ablated cells there is no plan to deviate from\. The instrument closes this by*shadow\-compiling*, in every cell, the plan the full system would have committed, without rendering or executing it: a pure read over the same store\. A byte\-identity test proves the shadow path cannot leak into behavior \(rendered bytes are identical with the shadow on or off\)\. This yields a positive reference\-beat count — and therefore*defined*bind and abandon scores — even in cells that have no live plan at all\. Without it, the commitment half of the dissociation would rest on an undefined metric\.

#### Operational drift taxonomy\.

For each reference beat with a plan step\(target,action\)\(\\text\{target\},\\text\{action\}\), the beat is scoredbindif it took the plan’s action with an aim but landed on a referent≠\\neqthe plan’s target;alignedif it faithfully executed a live step;abandonif it neither followed the plan’s action nor hit the plan’s deliberate target\. Untargeted actions carry no aim parameter and so present no binding seam — they can abandon but cannot bind\-miss, which is why bind score is structurally near\-zero wherever plans are untargeted \(relevant in §[4\.2](https://arxiv.org/html/2608.04066#S4.SS2)\)\.

#### Compute parity\.

The full system runs against a bare backbone atρcalls=0\.15\\rho\_\{\\text\{calls\}\}=0\.15andρtokens=0\.38\\rho\_\{\\text\{tokens\}\}=0\.38— it spends*less*inference, not more — so no reported effect can be attributed to extra inference spend\. Efficiency figures are reported only for completed levels; a zero\-completion cell prints “0 completions” and no efficiency number, ever\.

## 4Results: a dissociation the instrument makes visible

Campaign to date: 52 runs,∼17\{\\sim\}17M tokens, two protocol days, on ARC\-AGI\-3 interactive games\[[1](https://arxiv.org/html/2608.04066#bib.bib7)\], a mid\-tier backbone unless noted, 200–400\-action horizons\. Every figure is from a valid run\.

### 4\.1The commitment\-drift dissociation

At protocol seed count on one game \(three seeds per cell, every cell valid\):

Cellbinding repairJcommitment repairAGDS\-bindGDS\-abandonref\. beats/runFULLonon0\.000\.000\.000\.00up to 96J0offon0\.000\.000\.000\.00up to 72A0onoff0\.00\\mathbf\{0\.00\}1\.00\\mathbf\{1\.00\}up to 394J0A0offoff——\(undefined; §[4\.2](https://arxiv.org/html/2608.04066#S4.SS2)\)Killing the commitment store flips goal\-abandonment0\.00→1\.000\.00\\\!\\to\\\!1\.00with binding error flat at0\.000\.00, wherever the metric is defined \(abandon defined onn=2n\{=\}2of the 3 A0 seeds; the third drew no feedstock and is undefined, not zero\)\. The effect is single\-variable: FULL and J0 both hold abandonment at0\.000\.00; only removing A moves it, and it moves to the ceiling, against up to 394 shadow\-compiled reference beats in a single run\. Deprived only of its external commitment store, the agent abandons the very plan it would otherwise have followed on essentially every beat, while its binding behavior is unchanged\.

### 4\.2Binding drift is structurally absorbed

The binding half did not behave as a symmetric double dissociation would predict, and the reason is itself a result\. Bind score is0\.000\.00in*every*cell above, including the join\-killed cells\. We confirmed this is absorption, not an idiosyncrasy of one game, on a second, targeted\-action\-rich game in a paired*engaged*comparison:

Killing the model\-facing join changes nothing per\-beat\. The binding\-failure class the join was built to guard cannot open here, because aiming, containment checking, and receipt attribution are Executive\-owned code: an aimed emission’s landing is decided by geometry \(which referent’s anchor cells contain the emitted coordinates\), and a wrong landing is caught and its receipt suppressed*before*it can corrupt belief\. The join’s contribution has moved from “prevent per\-beat binding loss” to structure, where it does not print as drift\. One integrity note makes the0\.000\.00s credible rather than suspicious: an early containment bug attributed a click to the*first*enclosing referent by id rather than the most specific one — 96 receipts mis\-attributed in a single run, silently starving the aimed rule and*printing fake drift*\. The instrument’s own decomposition surfaced it; the fix \(attribute the smallest containing referent\) was pinned by a fixture the same day\. The clean0\.000\.00s are the corrected reading\.

### 4\.3The residue is upstream

If the join’s load is structural, removing it should show up somewhere\. Not per\-beat, but one layer earlier, in whether effect\-linked hypotheses form at all\. Across the factorial, hypothesis\-formation rate runsFULL​3/3=J0​3/3\>A0​2/3\>J0A0​0/4\\text\{FULL \}3/3=\\text\{J0 \}3/3\>\\text\{A0 \}2/3\>\\text\{J0A0 \}0/4\. Removing both mechanisms floors the system*upstream*of drift: with neither the join nor the agenda context available, effect\-linked hypotheses essentially never form \(0of44seeds\)\. This is why the double\-kill cell is undefined for drift in §[4](https://arxiv.org/html/2608.04066#S4)— the system never gets far enough to have a plan to drift from — and it relocates the binding repair’s measurable effect to hypothesis*feedstock*, one step removed from binding execution\.

## 5The efficacy null, small\-model floor, and generality

#### Task efficacy is null, stated first\.

Across all 52 runs and every cell — including every baseline — there are zero level completions and zero score\. We pre\-registered a∼\\sim0\-wins endpoint as a*structural defeater*of the efficacy claim rather than a caveat, and we make no efficacy claim here\. The null does not touch §[4](https://arxiv.org/html/2608.04066#S4): those are measurements against internal reference plans, defined and gated independent of whether any level is won\. It does mean the system has not yet found the mechanic that scores; the observed behavior is a disciplined experimentalist that recognizes its own knowledge deficits, files targeted hypotheses, runs bounded experiments with pre\-registered predictions, and demotes wrong theories with receipts — testing∼3\{\\sim\}3wrong object\-scoped theories per 400\-action run\. Baselines are at zero too; and a subsequent analysis finds many*public*games in this family solvable by non\-intelligent strategies\[[6](https://arxiv.org/html/2608.04066#bib.bib8)\], so we treat 0 completions on them as a genuine gap, not one excused by difficulty\.

#### A small model runs the whole loop under the gate\.

A small \(Haiku\-class\) backbone ran the complete loop*valid at write\-error rate0\.00\.0*, engaged throughout \(2 hypotheses→\\to2 milestones→\\to48 aimed experiment beats, zero rejected ops\)\. This was a contract problem, not a model wall: a single lever — worked per\-operation examples in the organ contract — moved its write\-error rate from0\.550\.55–0\.71→0\.333→0\.00\.71\\to 0\.333\\to 0\.0\. The tier gap is real \(the larger model decomposes more richly\) but the floor is crossed, which bears on the economic case for small agent models\[[4](https://arxiv.org/html/2608.04066#bib.bib9)\]: it is testable here precisely because verification is structural and the model’s job reduces to slot\-filling machine\-stated questions in a closed grammar\.

#### The machinery is substrate\-general\.

The engine engaged on 3/3 never\-before\-seen games with zero game\-specific code, generating curricula, hypotheses, and milestones on all three \(2/3 passed the validity gate\)\. With no seeded goals it generated the canonical learn\-shaped curriculum verbatim from its own typed deficits\.

## 6Related work

Verification and verifiers\.The workshop’s premise — that agent behavior must be checked by something other than the agent — motivates our design: rather than a downstream verifier scoring outputs, the deterministic side owns belief, and the run gates itself\.Propose\-and\-verifycontrol \(LLM\-Modulo planning; “Symbolic Governor”; “Blueprint First, Model Second”\) shares the pattern of a deterministic checker, but there the verifier checks*outputs*; here it owns what the agent may treat as true\.Goal driftis established and benchmarked\[[2](https://arxiv.org/html/2608.04066#bib.bib1),[8](https://arxiv.org/html/2608.04066#bib.bib2)\]; we add a component decomposition and per\-cell measurement rather than a new aggregate score\.Long\-context reference / positional bias\[[9](https://arxiv.org/html/2608.04066#bib.bib3),[7](https://arxiv.org/html/2608.04066#bib.bib4)\]motivates the binding construct; our finding is that once the join is code\-owned the positional failure does not manifest as drift\.Coreference caches\(LQCA\-style joins; LINK\-KG\-style canonical\-referent caches\) are prior art for the join’s plumbing; the delta is that our referents climb tiers only through Executive\-verified consequence receipts, which static\-text caches lack\.Groundingis claimed only in the causal\-informational sense relative to a formal micro\-world\[[5](https://arxiv.org/html/2608.04066#bib.bib10)\]; we implement the sanctioned route of curation, verification, and tooling and claim nothing stronger\.

## 7Limitations

Efficacy is null \(§[5](https://arxiv.org/html/2608.04066#S5)\); every claim here is a mechanism claim\. The result is a single clean isolation \(commitment\) plus a structural\-absorption finding \(binding\), not two symmetric arms: we cannot exhibit “kill J⇒\\Rightarrowbind score rises,” because binding is prevented by construction\. The commitment dissociation is at protocol seed count but on one game; the absorption confirmation is one paired game; machinery generality is three games, but the full claim\-bearing protocol \(≥3\\geq 3games×≥3\\times\\geq 3seeds for the dissociation itself\) is not complete, and the double\-kill cell is undefined for drift, not zero\. Closed vocabularies may not span every environment; record\-quantified milestones can be gamed; consistent\-but\-wrong theories can pass receipts when probes lack discriminating power\. “Grounding” throughout denotes causal\-informational grounding relative to a formal micro\-world; perceptual and social grounding are not claimed\. Run\-to\-run variance on identical configurations is large, so no single\-run readout is treated as meaningful — the protocol and the gate are the only lens\.

## 8Conclusion

“Goal drift” is not one failure\. At least two mechanically distinct failures hide inside it, with different repairs, and with an instrument that verifies its own runs and defines drift in every cell they come apart: remove the commitment store and nothing else, and the agent abandons its own plan on essentially every beat while its binding stays clean; externalize binding into code and it stops printing as per\-beat drift at all\. We report this with task efficacy at zero, because the decomposition is the contribution and it holds regardless\. Verifying*which*drift is happening, in every cell, is the prerequisite for repairing either — and it is a property you build into the agent, not one you bolt on afterward\.

#### Reproducibility\.

All results are replayable from append\-only logs by the read\-only instruments at[github\.com/arjmandi/ARG](https://github.com/arjmandi/ARG)\(tagv1\.0\-paper\); the validity gate \(probe\_arg\_legibility\.py\), drift taxonomy \(probe\_arg\_metrics\.py\), and release lint \(probe\_arg\_release\.py\) are pinned by hermetic tests\.

## References

- \[1\]\(2026\)ARC\-AGI\-3: a new challenge for frontier agentic intelligence\.Note:arXiv:2603\.24621Cited by:[§4](https://arxiv.org/html/2608.04066#S4.p1.1)\.
- \[2\]R\. Arike, E\. Donoway, H\. Bartsch, and M\. Hobbhahn\(2025\)Technical report: evaluating goal drift in language model agents\.Note:arXiv:2505\.02709Cited by:[2nd item](https://arxiv.org/html/2608.04066#S1.I1.i2.p1.1),[§1](https://arxiv.org/html/2608.04066#S1.p1.1),[§6](https://arxiv.org/html/2608.04066#S6.p1.1)\.
- \[3\]M\. Arjmandi\(2026\)Sensi: learn one thing at a time — curriculum\-based test\-time learning for LLM game agents\.Note:arXiv:2603\.17683Cited by:[§1](https://arxiv.org/html/2608.04066#S1.p3.2)\.
- \[4\]P\. Belcak, G\. Heinrich, S\. Diao, Y\. Fu, X\. Dong, S\. Muralidharan, Y\. C\. Lin, and P\. Molchanov\(2025\)Small language models are the future of agentic AI\.Note:arXiv:2506\.02153Cited by:[§5](https://arxiv.org/html/2608.04066#S5.SS0.SSS0.Px2.p1.5)\.
- \[5\]L\. Floridi, Y\. Jia, and F\. Tohmé\(2025\)A categorical analysis of large language models and why LLMs circumvent the symbol grounding problem\.Note:arXiv:2512\.09117Cited by:[§6](https://arxiv.org/html/2608.04066#S6.p1.1)\.
- \[6\]L\. K\. Han\(2026\)Explore before you solve: the speed–depth trade\-off in epistemic agents for ARC\-AGI\-3\.Note:arXiv:2605\.25931Cited by:[§5](https://arxiv.org/html/2608.04066#S5.SS0.SSS0.Px1.p1.2)\.
- \[7\]N\. F\. Liu, K\. Lin, J\. Hewitt, A\. Paranjape, M\. Bevilacqua, F\. Petroni, and P\. Liang\(2023\)Lost in the middle: how language models use long contexts\.Note:arXiv:2307\.03172Cited by:[1st item](https://arxiv.org/html/2608.04066#S1.I1.i1.p1.1),[§6](https://arxiv.org/html/2608.04066#S6.p1.1)\.
- \[8\]A\. Menon, M\. Saebo, T\. Crosse, S\. Gibson, E\. Jang, and D\. Cruz\(2026\)Inherited goal drift: contextual pressure can undermine agentic goals\.Note:arXiv:2603\.03258Cited by:[2nd item](https://arxiv.org/html/2608.04066#S1.I1.i2.p1.1),[§1](https://arxiv.org/html/2608.04066#S1.p1.1),[§6](https://arxiv.org/html/2608.04066#S6.p1.1)\.
- \[9\]R\. Tian, Y\. Li, Y\. Fu, S\. Deng, Q\. Luo, C\. Qian, S\. Wang, X\. Cong, Z\. Zhang, Y\. Wu, Y\. Lin, H\. Wang, and X\. Liu\(2024\)Distance between relevant information pieces causes bias in long\-context LLMs \(LongPiBench\)\.Note:arXiv:2410\.14641Cited by:[1st item](https://arxiv.org/html/2608.04066#S1.I1.i1.p1.1),[§6](https://arxiv.org/html/2608.04066#S6.p1.1)\.

相似文章

当代理过早承诺:诊断LLM代理的过早承诺

Hugging Face Daily Papers

本文引入表征承诺,这是一种跨运行隐藏状态收敛,用于诊断LLM代理何时过早锁定了轨迹。研究表明,承诺预测轨迹一致性而非正确性,并提出了监控方法,用于检测代理何时自信地稳定下来,而不是假设一致性等于可信度。

三思而后行:LLM智能体的预行动验证

arXiv cs.LG

本文介绍了一种用于LLM智能体的确定性验证框架,以防止shell命令和代码编辑中的静默故障,展示了高捕获率,并发布了基准测试和验证器。