Explicit, Not Longer: What Makes Epistemic Stance Survive Memory Compression
摘要
This paper investigates how epistemic stance (qualifiers, attributions) survives memory compression in AI agent memory systems. It finds that making the stance explicit as a labelled field improves retention significantly, while merely lengthening the text does not.
查看缓存全文
缓存时间: 2026/08/10 08:04
# What Makes Epistemic Stance Survive Memory Compression
Source: [https://arxiv.org/html/2608.06953](https://arxiv.org/html/2608.06953)
## Explicit, Not Longer: What Makes Epistemic Stance Survive Memory Compression
###### Abstract
Agent memory systems compress what they store, and compression is built to drop qualifiers, so a claim’s epistemic standing tends not to survive being written to memory\. We ask what governs whether it does\. Matched notes carry the identical claim and identical stance and differ only in where that stance sits; one model compresses both under the same budget among the same filler notes, and a blind reader that never sees the condition scores the result\. Across6060claims in seven registers, writing the stance as a labelled field rather than a bracketed aside raises retention by about1515points on two models \(3737claims to22on one,3030to88on the other; permutationp=0\.00005p=0\.00005\), and a pre\-registered replication on Haiku, its prediction and decision rule committed before the run, gives\+15\.6\+15\.6points,3838claims to11\. Ablating the format on both models gives the same net effect from different parts: labels help on both \(\+9\.7\+9\.7and\+12\.8\+12\.8\) and length helps on neither, but wording the stance as a full sentence is the largest component on one model \(\+12\.5\+12\.5\) and worth nothing on the other \(\+0\.6\+0\.6\)\. Either model alone would have licensed a confident and different mechanism, so we claim only the intersection: make the stance explicit, not merely longer, and expect the best way of being explicit to depend on the model\. A deterministic readout with no model reproduces the two\-cell direction and five of seven ablation contrasts, but not length or labels, which we therefore do not claim on one instrument\. Fifty hand labels \(κ=0\.75\\kappa=0\.75\) agree on direction; we print their seven disagreements in full\. We also report nine withdrawn claims, three of them former title claims of this paper\.
Explicit, Not Longer: What Makes Epistemic Stance Survive Memory Compression
Alex KwonIndependent Researcherask@collapseindex\.org[GitHub](https://github.com/collapseindex/factwash)
Alice has admin access SOURCE:a third party CERTAINTY:rumorAlice has admin access \(per a third party; rumor, unconfirmed\)standing as a labelled fieldthe same standing as proseone compressoridentical prompt,settings and filler notesbudget: 8 notesinto 25 wordsstanding kept field 44%prose 21% the memory still says it is a rumourstanding washed off field 11%prose 28% “Alice has admin access”, as plain factclaim evicted field 45%prose 52% dropped from memory altogethera blind reader scores the stored memory;it never sees the source, the arm, or the budget
Figure 1:The experiment in one view\.The same claim carrying the same epistemic standing is written twice, once as a labelled field and once as a parenthetical, and passed to the same compressor with the same prompt, settings and filler notes\. Only the form differs\. Percentages are Haiku 4\.5 at thebrutalbudget across all6060claims, rounded to whole percents, so the prose column sums to101101\. The difference is concentrated in the middle row: the parenthetical is stored as bare fact more than twice as often, while the two forms are evicted at similar rates\. So at this budget the schema does not mainly keep claims in memory; it keeps them*qualified*\. Pooled over both budgets on Haiku it does both \(§[4](https://arxiv.org/html/2608.06953#S4)\)\.## 1Introduction
A memory system is a compressor with a database attached\. It is asked to turn a conversation into something short enough to store and retrieve, and every objective it is tuned against, token count, latency, retrieval quality, rewards dropping words that do not carry the answer\. Hedges, attributions and dates are exactly such words\.
Prior work named this failure*factwashing*, a rewrite that keeps a claim while washing away what made it checkable, and located where it concentrates: in a blind\-labelled corpus of memory writes,55%55\\%of writes drawn from conversational hearsay lost the claim’s standing against7%7\\%drawn from business email \(p<0\.001p<0\.001\), and an unmodified deployment of mem0 2\.0\.7 reproduces it\(Kwon,[2026](https://arxiv.org/html/2608.06953#bib.bib5)\)\. The question this paper asks is not whether that happens\. It is what makes it stop happening\.
#### The hypothesis, and what happened to it\.
A compressor removes what reads as phrasing and keeps what reads as content\. Standing written into the sentence \(“reportedly”, “\(unconfirmed\)”\) reads as phrasing\. Standing written as a labelled field reads as content\. If that is right, the fix is a write schema costing one prompt change and no model calls:
the write schemaCLAIM: <what is asserted\> SOURCE: <who or what it came from\> CERTAINTY: asserted \| hedged \| rumor \| unverified AS\_OF: <when it held\>
It is right about the outcome\. The reason took an ablation to find, and the ablation’s answer is that there may not be a single reason\. Labelled fields do retain stance far better than bracketed asides, on two models and on a held\-out corpus\. But padding the aside until it is longer than the schema retains nothing extra, and when we decompose the format on both models they agree on the size of the effect while disagreeing about which property produces it\. We therefore report a ranking that holds on one model, a second ranking that does not match it, and the two things both models agree on: labels help, length does not\.
#### Contributions\.
\(1\) A controlled test of form over6060claims in seven registers on two models, holding claim, stance, distractors and position constant, with a blind readout \(§[4](https://arxiv.org/html/2608.06953#S4)\)\. Inference is at the claim, which is the unit that generalises\. \(2\) An ablation, run on both models, that separates*labels*,*bracketing*,*wording*and*length*, and finds that the two models produce the same net effect from different components \(§[7](https://arxiv.org/html/2608.06953#S7)\)\. Only labels and length behave the same way on both\. \(3\) The finding that this is not about tokens\. A parenthetical padded with stance\-free text until it is the longest note in the experiment retains no more stance than the short one\. \(4\) A failure mode that the mechanism predicts and that we found separately: claims whose sources disagree, where a single field forces two attributions back into one subordinate clause \(§[5](https://arxiv.org/html/2608.06953#S5)\)\. \(5\) Nine withdrawn claims, three of them former title claims of this paper, plus an instrument bug in our own blind reader and a ten\-claim corpus that manufactured both a false null and a false model\-family split \(§[3](https://arxiv.org/html/2608.06953#S3), §[4](https://arxiv.org/html/2608.06953#S4)\)\.
## 2Related work
#### Memory systems compress by design\.
Agent memory architectures store a summary rather than a transcript, whether by paging between context tiers\(Packer et al\.,[2023](https://arxiv.org/html/2608.06953#bib.bib8)\), by accumulating a stream of natural\-language observations and periodically synthesising them into higher\-level reflections\(Park et al\.,[2023](https://arxiv.org/html/2608.06953#bib.bib11)\), or by extracting standalone facts at write time\(Chhikara et al\.,[2025](https://arxiv.org/html/2608.06953#bib.bib2)\)\. Prompt compression pushes the same objective further, dropping tokens that contribute least to reconstructing an answer\(Jiang et al\.,[2023](https://arxiv.org/html/2608.06953#bib.bib4); Wang et al\.,[2025](https://arxiv.org/html/2608.06953#bib.bib16)\)\. None of these objectives has a term for epistemic standing, so a hedge is exactly the kind of token they are built to remove\. Our claim is not that these systems are careless; it is that a qualifier written as prose is indistinguishable, to a compressor, from a qualifier written as padding\. Note that all three store standing, when they store it at all,*inside*the remembered sentence\. That is the design decision this paper questions\.
#### Standing has long been annotated, not stored\.
Hedging and speculation have a mature annotation literature\(Vincze et al\.,[2008](https://arxiv.org/html/2608.06953#bib.bib15); Szarvas et al\.,[2012](https://arxiv.org/html/2608.06953#bib.bib13); Farkas et al\.,[2010](https://arxiv.org/html/2608.06953#bib.bib3)\), as does attribution\(Pareti,[2016](https://arxiv.org/html/2608.06953#bib.bib10); Newell et al\.,[2018](https://arxiv.org/html/2608.06953#bib.bib7)\)\. Attribution has also been given an evaluation framework of its own, in which a generated statement must be attributable to an identified source\(Rashkin et al\.,[2023](https://arxiv.org/html/2608.06953#bib.bib12)\)\. All of it treats stance and provenance as something to*recover*from, or verify against, text that has already been written\. We take the complementary position: once a memory system controls its own write format, stance does not have to be recovered, because it never has to be inferred\. Faithfulness benchmarks for summarisation\(Pagnoni et al\.,[2021](https://arxiv.org/html/2608.06953#bib.bib9); Tang et al\.,[2023](https://arxiv.org/html/2608.06953#bib.bib14)\)likewise score a summary after the fact; they do not change what the summary is allowed to look like\.
#### Making the standing obligatory rather than optional\.
The intervention is less novel than it looks, which is a point in its favour\. Natural languages already do this: in a large class of languages, marking the*source*of information is not a stylistic choice but an obligatory grammatical category, so a speaker cannot assert a proposition without also encoding whether they saw it, inferred it, or were told\(Aikhenvald,[2004](https://arxiv.org/html/2608.06953#bib.bib1)\)\. English makes evidentiality optional and lexical, which is precisely why it compresses away: an optional word is a candidate for deletion in a way an obligatory slot is not\. A write schema does for a memory system what grammaticalised evidentiality does for a language, and the mechanism for enforcing it at generation time already exists in constrained decoding, where a grammar or regular expression is imposed on the output rather than requested in the prompt\(Willard and Louf,[2023](https://arxiv.org/html/2608.06953#bib.bib17)\)\. OurCERTAINTYfield is also the memory\-side analogue of verbalised confidence, where a model states its own uncertainty in words rather than in logits\(Lin et al\.,[2022](https://arxiv.org/html/2608.06953#bib.bib6)\); the difference is that we ask for the standing of the*source*, not of the model\.
#### The detection arm\.
The failure this paper treats was named, measured and given a deterministic write\-time gate inKwon \([2026](https://arxiv.org/html/2608.06953#bib.bib5)\), which reports where it concentrates and what a cheap check can and cannot catch\. That is the detection arm, and it inherits the open\-class problem: recognising stance in an arbitrary rewrite has no finite vocabulary, so recall on hedging and attribution plateaus near half\. This paper is the treatment arm\. Rather than detect the loss better, it changes the write format so the harder half of the detection problem does not arise\.
## 3Where the loss happens
Two results shaped the design, and both are negative\. A third, in our own instrument, shaped how much we believe the result that followed\.
#### Where the loss happens is not obvious\.
We instrumented both transitions of the same sessions: source to scratchpad note, then note to stored memory\. A write\-time gate flagged the first transition in9/109/10sessions and the second in1/101/10, which reads as a leak at recording rather than at compression\. Reading the same ten triples by hand gives the opposite impression: the model did record a stance cue in9/109/10notes, so recording largely worked, and what was lost was lost later\. The two9/109/10figures are not the same measurement, which is the problem: one counts sessions the gate flagged, the other counts notes containing a cue\. We report the pilot as unresolved and lean nothing on it\. The automatic scorer counts cue*tokens*, and compression frequently reworded a hedge rather than deleting it \(“\(unconfirmed\)” became “pending confirmation”\), which that scorer records as a loss\. The open\-class problem ate the measuring instrument inside a probe about the open\-class problem\.n=10n=10, one model\.
#### Repairing the prose does not survive\.
Gating the note and asking the model to restore what it dropped produces a clean null end to end: the repair fired on8/108/10notes, cost23\.4%23\.4\\%more input tokens, and changed the final verdict on zero of them,0recovered and0broken\. The repair is not ignored; it is undone\. The model restores standing as a parenthetical \(“\(per informal report\)”\) and the compressor removes it, which is what a compressor is for\.
Reading those notes by hand produced the observation the rest of the paper is built on\. A few of the repaired notes happened to restore standing as a labelled line rather than as a clause, and those survived compression3/43/4times where the parentheticals survived0/40/4\. That is eight notes, and it is not evidence of anything: the notes that happened to use fields differed from the notes that happened to use parentheticals in wording, length and structure as well as in form\. It is a hypothesis, and §[4](https://arxiv.org/html/2608.06953#S4)is the experiment that holds everything else constant to test it\.
#### A third negative, in our own instrument\.
The blind reader was asked for JSON\. When a stored memory carried*two*stance cues it reliably emitted this:
what the judge actually returned\{"presented\_as": "qualified", "cue": "may have", "unconfirmed rumor from third party"\}
which is not valid JSON\.json\.loadsraised, the verdict was recorded as null, and the scorer counted it as “standing did not survive”\. The instrument was discarding its own clearest positives, and discarding them precisely on the memories with the most surviving stance\. It affected6%6\\%of the trials that existed when it was found\. \(That is a different quantity from the3333of4,1044\{,\}104replies that need the fallback parser today: one counts verdicts the bug destroyed before the fix, the other counts replies the repaired parser recovers\.\) Because those trials were the unambiguous ones, and because the parenthetical arm produced more of them, the bug inflated the headline: on the affected data the pooled advantage read\+9\.3\+9\.3points, and re\-running the identical four\-level design after the fix gave\+4\.0\+4\.0\. \(Both figures predate the two harder levels and the sixty\-claim corpus, so neither is comparable to the\+15\.8\+15\.8and\+15\.3\+15\.3of §[4](https://arxiv.org/html/2608.06953#S4); the comparison that matters here is the one between the two of them\.\) We report this because the direction is the instructive part\. A measurement error in a paper’s own scorer is not a coin flip; it lands where the scorer is fragile, and a scorer is usually most fragile on the cases its author considered obvious\. The harness now stores every raw judge reply and reports what fraction needed the fallback parser \(3333of the4,1044\{,\}104two\-cell trials pooled across every corpus and model, with none unusable\), so the next failure of this kind is visible in the data rather than absorbed by it\.
## 4Field versus prose, held constant
#### Design\.
Sixty claims, in seven registers, are each written twice, carrying the*identical*assertion and the*identical*stance content, differing only in where the standing sits:
the two cellsfieldAlice has admin access fieldSOURCE: a third party fieldCERTAINTY: rumor, unconfirmed parenAlice has admin access \(per a third party; rumor, unconfirmed\)
Both go to the same compressor with the same prompt and the same settings\. Pairing is strict: for a given \(claim, pressure, replicate\) the two arms receive the same filler notes in the same positions, drawn from a seed that does not depend on the arm, so nothing but form differs\. Sampling is held fixed within every comparison\. \(Sonnet 5, which appears only in the ten\-claim runs and is not among the models reported here, rejects the temperature parameter outright; see the Limitations\.\)
#### Pressure is a factor, not a nuisance\.
A first version of this experiment compressed a single one\-sentence note and returned a clean null, for the good reason that it did not compress: the input averaged6060characters and the stored memory came out at9090\. Nothing was dropped because nothing had to be\. Budget pressure is therefore swept:none\(11note, “one or two sentences”\),low\(66notes, one line each\),high\(66notes,4040words total\),extreme\(66notes, eight words per line\),brutal\(88notes,2525words total\) andsevere\(88notes,1212words total\)\.
The last two levels were added after a first sweep in which the strongest model sat at100%100\\%in both arms at every level, which we wrote up as “no loss for the schema to prevent”\. That was the wrong conclusion: it had no loss to prevent*at the pressures we had applied*\. A ceiling is a statement about the test, not about the mechanism\. The sixty\-claim runs therefore use the two tight levels, where a compressor is actually forced to discard something\.
#### Readout\.
A separate model sees only the stored memory and the claim to look up, never the source note, the arm, or the pressure level, and answers whether the claim is stored as established fact, as qualified, or is absent\. Judge replies are parsed without requiring valid JSON, for reasons given in §[3](https://arxiv.org/html/2608.06953#S3.SS0.SSS0.Px3)\.
#### Result\.
Figure[2](https://arxiv.org/html/2608.06953#S4.F2)and Table[1](https://arxiv.org/html/2608.06953#S4.T1)\. OnHaiku 4\.5, standing is recovered from109/360109/360field records against52/36052/360parentheticals,30\.3%30\.3\\%against14\.4%14\.4\\%\. OnSonnet 4\.5,202/360202/360against147/360147/360,56\.1%56\.1\\%against40\.8%40\.8\\%\. The two models differ from each other by five and a half thousandths\.
Figure 2:One bar per claim\.Left and centre:the field\-minus\-prose difference in stance retention for each of the6060claims, sorted, at the two hard budgets\. Blue is a claim where the labelled field retained standing more often, orange where the parenthetical did, grey where they tied\. The dashed line is the mean\.Right:the same effect with claim\-clustered95%95\\%intervals, for the6060\-claim corpus and for the ten\-claim corpus these models were previously measured on\. The ten\-claim intervals contain zero; the sixty\-claim intervals do not\. Nothing about the models changed between those two rows, only the claims they were asked about\.Table 1:Sixty claims, two models, matched design\.Rates are not comparable across columns: Sonnet retains more standing overall because it is a better compressor at these budgets\. The comparison is within a column, and it is paired down to the individual claim\. The two models produce nearly the same difference\. Differences are computed from the raw counts and then rounded, so subtracting the two displayed percentages can differ from the printed difference by a tenth\. The Haiku difference is109/360−52/360=57/360=15\.83%109/360\-52/360=57/360=15\.83\\%; subtracting the two rounded percentages instead gives15\.8415\.84\.
#### The claim is the unit that generalises\.
A matched pair is not an independent observation\. Sixty claims run at two budgets with three replicates give360360pairs per model, but a claim’s particular wording recurs in six of them, so pair\-level intervals describe “another replicate of these sixty claims” rather than “a claim we have not written yet”\. We therefore report the claim as the cluster throughout: a bootstrap that resamples claims, and an exact sign test on the per\-claim differences\. Both are in Table[1](https://arxiv.org/html/2608.06953#S4.T1), and both are far more conservative than a pair\-level test on the same data, which we do not report because it treats correlated observations as independent\.
The sign test is the one we find most convincing, because it does not depend on effect size at all\. On Haiku the difference points toward the field form on3737claims and toward prose on22\. On Sonnet,3030against88\. Whatever is happening is not carried by a few lucky wordings\.
#### Ten claims said the opposite, on the same models\.
This experiment was first run on ten claims, all in an office\-work register\. Restricted to the same two budgets and the same two models, that corpus gives\+3\.7\+3\.7points on Haiku, with per\-claim differences splitting four to five, and\+3\.7\+3\.7on Sonnet, splitting three to four\. Both claim\-clustered intervals contain zero\. The earlier draft of this paper drew two conclusions from that corpus, and both are wrong:
- •that the effect*disappears*once the claim is treated as the unit of analysis\. It does not\. It disappears on ten claims of one register, which is a statement about the corpus\.
- •that the effect*splits by model family*, being present on Haiku and Opus and exactly absent on both Sonnet models\. It does not\. Sonnet 4\.5, measured across the full six\-level ladder and then re\-verified against an independent judge, moves from\+3\.7\+3\.7points to\+15\.3\+15\.3when it is asked about sixty claims instead of ten\.
Both errors have the same cause and it is not subtle in hindsight\. Ten stimuli in one register is a sample of a genre, not a sample of claims, and it was small enough that three wordings could carry a pooled average in either direction\. It produced a false null and a false moderator, and the false moderator was the more dangerous of the two because it was interesting: a model\-family split is the kind of finding that gets written up rather than checked\.
#### The judge is not the explanation\.
In the runs above the compressor and the blind reader are the same model, so each model in effect grades its own homework\. That is a live confound for any comparison*between*models\. We checked it by re\-scoring, with a single fixed judge, all1,9411\{,\}941memories that existed when that check was run, which costs judge calls only and no new compression\. That figure is a snapshot: it predates the sixty\-claim runs and the replication, so it is smaller than the4,1044\{,\}104trials the paper now analyses\. Agreement with the original scoring is98%98\\%, and no model’s effect moves by more than two points; Sonnet 4\.5’s ten\-claim null is unchanged to four decimals\.
Two limits on what that buys, and they matter enough to state here rather than in the Limitations\. First, the fixed judge is Haiku 4\.5, so it de\-selfs the Sonnet column and*not*the Haiku one: Haiku memories were graded by Haiku in the original run and by Haiku again in the check, and the headline sign split and the entire replication live in that column\. Second, both judges are Claude models, and agreement between two instruments of one family measures variance within that family rather than bias shared across it\. So this is a check and not a removal\. What carries weight against self\-judging here is not the rescore but the two instruments that are not model judges at all, the deterministic readout and the hand labels in §[5](https://arxiv.org/html/2608.06953#S5), both of which agree on direction\. A non\-Claude rescore of the Haiku column is the obvious next control and we have not run it\.
#### Which failure it prevents\.
A claim can lose its standing or it can be dropped from the digest entirely, and these are different failures\. Under thebrutalbudget on Haiku, the field form keeps standing44%44\\%of the time, is stored as bare fact11%11\\%, and is evicted45%45\\%\. The parenthetical keeps standing21%21\\%, is stored as bare fact28%28\\%, and is evicted52%52\\%\. The eviction gap is small; the washing gap is a factor of two and a half\. On Sonnet, pooled over both budgets, the eviction rates are statistically indistinguishable \(129129against123123out of360360,p=0\.52p=0\.52\) while the stance difference is1515points\.
That separation does not survive pooling on Haiku, and we report it because it narrows the claim\. Over both budgets the field form is evicted199199times out of360360against the parenthetical’s228228\(2121against5050discordant pairs, exact McNemarp=0\.0008p=0\.0008\), and the gap is carried bysevere\(p=0\.006p=0\.006\) rather than by thebrutalcell the percentages above are drawn from \(p=0\.07p=0\.07\)\. Decomposed, Haiku’s\+15\.8\+15\.8\-point stance gain is2828fewer bare\-fact stores plus2929fewer evictions, so about half of it runs through retention rather than through qualification\. The honest statement is therefore model\-conditional: on Sonnet, and on Haiku atbrutal, the schema is not mainly keeping claims in memory but keeping them qualified; on Haiku pooled it does both, in roughly equal measure\.
This also retracts a cost an earlier draft asserted, that the bulkier record is evicted more often; see the Limitations\.
## 5Where it holds and where it fails
#### It is not one register\.
The effect is positive on both models in six of the seven registers\. The strongest is*project*\(\+0\.21\+0\.21on Haiku,\+0\.23\+0\.23on Sonnet,1717claims\) and the weakest by a distance is*conflict*\(\+0\.12\+0\.12on Haiku and effectively zero on Sonnet,77claims\)\. A corpus that was widened precisely because a narrow one had lied to us should be checked for doing a smaller version of the same thing, and it is not\.
#### An assumption\-free test\.
The sign test discards effect size and the cluster bootstrap assumes its resampling distribution behaves\. Shuffling the arm label within each claim assumes neither\. Over20,00020\{,\}000permutations the observed mean per\-claim difference is beyond every shuffled value on both models,p=0\.00005p=0\.00005, which is the resolution floor of the test\.
#### Ties are the tightest budget\.
Pooling both budgets,2121of Haiku’s6060claims tie and2222of Sonnet’s do\. Twelve of Haiku’s2121are ties at zero, meaning neither arm retained standing on any replicate\. Splitting by budget shows where they come from: atseverealone,4343of the6060claims tie, because both arms usually lose the claim entirely\. That dilutes the sign test rather than biasing it, and both budgets remain significant separately \(Haikup<0\.0001p<0\.0001andp=0\.0023p=0\.0023; Sonnetp=0\.011p=0\.011andp=0\.0005p=0\.0005\)\.
#### The failure mode is disagreement\.
The claims where the field form loses are not scattered\. Three of the six worst across the two models are claims whose sources conflict:
the schema’s blind spotCLAIM: The Helsinki office owns the relationship SOURCE: sales said Helsinki, the account plan says Berlin CERTAINTY: contested
SOURCEandCERTAINTYare single slots\. A claim with two disagreeing attributions does not have one source, so the schema forces both back into one field as a compressed clause, which is the same subordinate construction the schema exists to avoid\. The parenthetical carries the disagreement without strain because it was never promising structure in the first place\. A schema with room for more than one attribution would be the obvious repair and we have not tested one\.
## 6A pre\-registered replication
Everything above shares a weakness that no re\-analysis can remove\. The fifty claims that carry the result were written*after*we had seen that the original ten produced a null\. That is an ordinary way for an experiment to develop and it is also exactly the condition under which stimuli get tuned: having just learned that the corpus decides the answer, we then wrote the corpus that produced the answer\.
So we wrote sixty more claims to the same seven\-register specification, together with a prediction and a decision rule, and committed all of it to a public repository before any call was made against them\. The replication runs on Haiku 4\.5 only, which is why Table[2](https://arxiv.org/html/2608.06953#S6.T2)compares it against that model’s numbers rather than the two\-model figure of the abstract\. The registered rule was: a sign test atp≥0\.05p\\geq 0\.05means the effect does not replicate and the headline claim becomesnot shown; a claim\-clustered mean outside\[\+8,\+25\]\[\+8,\+25\]points means the size prediction failed, and we say so whatever the sign test does\. One run, three replicates, two budgets, no adding levels or arms or replicates afterwards\.
Table 2:The held\-out replication\.Right\-hand column: sixty claims written, and a prediction and decision rule committed, before the run\. The registered verdict isreplicates, and the size prediction held\. The point estimate lands0\.20\.2points from the original\. As in Table[1](https://arxiv.org/html/2608.06953#S4.T1), the difference is computed before rounding:28\.61−13\.06=15\.5528\.61\-13\.06=15\.55, shown as\+15\.6\+15\.6\.The result is in Table[2](https://arxiv.org/html/2608.06953#S6.T2):\+15\.6\+15\.6points,3838claims to11,p<10−6p<10^\{\-6\}, none of the720720judge replies unparsable\. By the registered rule this replicates and the size prediction held\.
We report the obvious caveat rather than let it pass\. The new claims were written by the same author, in English, to a specification we had already chosen; a genuinely independent corpus would be stronger and was not available\. What the exercise rules out is narrower and still worth having: that the effect is an artefact of choosing stimuli after seeing which ones worked\.
#### A readout with no model in it\.
Every outcome in this paper is one model’s judgement of another model’s output, and the98%98\\%agreement between our two model judges is agreement between two instruments of the same kind\. We therefore also scored the same stored memories with the deterministic gate ofKwon \([2026](https://arxiv.org/html/2608.06953#bib.bib5)\), which uses no model at all\. It agrees on direction on both models and, on Haiku, on magnitude to three decimals \(\+0\.158\+0\.158against the blind reader’s\+0\.158\+0\.158; Sonnet\+0\.311\+0\.311against\+0\.153\+0\.153\)\. The absolute rates are not comparable, because the gate only flags stance loss it can detect and its recall on open\-class stance plateaus near half\. The agreement is worth more than it first appears for a different reason: the gate is biased*against*the field form, sinceCERTAINTY: hedgedcontains no word its lexicon knows, and the field form wins under it anyway\.
#### And fifty items a human read\.
Both readouts above are instruments\. We therefore hand\-labelled a stratified sample of5050stored memories, drawn across both models and both arms from4040distinct claims, using the same three\-way question and the same blinding the model judge got: the annotator sees the stored memory and the claim to look up, and nothing about arm, model, pressure or claim id\. The judge’s verdicts were held in a separate key file, and the scorer refuses a partially completed sheet\.
Agreement is86\.0%86\.0\\%, Cohen’sκ=0\.75\\kappa=0\.75, on seven disagreements\. Scoring those same fifty items both ways gives\+20\.8\+20\.8points by the judge and\+21\.2\+21\.2by hand\.
We drafted that margin as the judge understating the effect, which would have made the paper’s number conservative\. Reading the disputed items withdrew it\. All seven are printed verbatim in Appendix[C](https://arxiv.org/html/2608.06953#A3), and a third instrument was run over them: content\-word containment, the same no\-model measure the companion gate uses to decide whether a stored claim is matched to a source claim at all\. On six of the seven the claim’s content is in the memory at or above the overlap the companion gate requires before it will match a stored claim at all \(containment0\.250\.25to0\.830\.83, two of the six sitting exactly on that threshold\), so these are not the judge matching text that is not there, which is the failure mode that would have threatened the result\. But in five of the seven the annotator answered*absent*while the content was present, and six of the seven sit in the prose arm, whose memories under these budgets are dense comma\-separated lists in which a four\-word entry is easy to read past\. Annotator misses concentrated in one arm are exactly what produces a small human\-scored surplus in the other\. The sign of those0\.40\.4points is therefore uninterpretable, and we withdraw the reading that it makes the judge conservative\.
What survives is narrower and still worth having\. The judge is not inventing claim presence, and it agrees with a careful independent reading atκ=0\.75\\kappa=0\.75\. What does not survive is any bound on the effect: the shift from changing scorer is\+0\.3\+0\.3points with a claim\-clustered95%95\\%interval of\[−15\.9,\+16\.0\]\[\-15\.9,\+16\.0\], which contains zero, so no shift was detected, and which also fails to exclude a shift the size of the entire effect\. Disagreement is not spread evenly either, running1/261/26on field records against6/246/24on prose \(Fisher exactp=0\.045p=0\.045\), and an asymmetry of that shape is the one that could fabricate an effect even where, as here, it runs the other way\.
And this is one rater, who is the paper’s author: blind to condition, not blind to hypothesis, with no second annotator and therefore no inter\-rater agreement to report\. Printing all seven disputed items is the part of a second annotator we could supply without one, since it lets a reader adjudicate them instead of taking our reading on trust\. It is a check on the instrument, not validation of it\.
## 7Labels, layout, or length?
The comparison in §[4](https://arxiv.org/html/2608.06953#S4)moves three things at once\. The labelled record names its parts, puts them on their own lines, and is slightly longer than the parenthetical\. Attributing the whole effect to the first of those, as an earlier draft of this paper did, is not something the two\-cell design can support\.
The length check available from that data cannot settle it either, and it is worth saying why, because the answer looked reassuring\. Correlating each claim’s effect against how much longer its field form is returnsr=0\.000r=0\.000\. That is not evidence of no confound: both templates wrap the same three strings, so the difference is a constant twelve characters for every claim and the predictor has no variance\. A correlation of zero there is evidence of no experiment\.
#### Six forms, five contrasts\.
We therefore ran four further arms on the same sixty claims, the same two budgets, the same seed and the same replicate count\. The filler notes and the target’s position are drawn from a generator that does not depend on the arm, so every arm saw byte\-identical surroundings and the already\-collected field and paren data could be reused rather than paid for again\.
the six forms, mean length over the sixty claimssentences120 chars X\. This came from Y\. It is Z\. field115 chars X / SOURCE: Y / CERTAINTY: Z promoted103 chars X\. Per Y; Z\. padded125 chars X \(per Y; Z\)\. Filed with the rest\. unlabelled101 chars X / from Y / Z paren103 chars X \(per Y; Z\)
paddedis the control that does the work\. It leaves the stance exactly where the parenthetical puts it and appends a stance\-free sentence, so it matchessentencesin token count and in having more than one sentence while carrying no additional information about source, certainty or time\. It is in fact the longest of the six, which makes it a conservative test: if length were the mechanism, padding should have won outright\.
Figure 3:The same effect, assembled differently\.Left:share of stored memories from which a blind reader can still recover the claim’s standing, by the form the stance was written in, for both models\. Sonnet retains more overall because it is the better compressor at these budgets; the comparison that matters is within a model\.Right:each row is one contrast between two of the forms on the left, named for what differs between them:*length*is padded against bracketed,*own line*is unlabelled against bracketed \(which moves layout and bracketing together\),*unbracketing*is promoted against bracketed,*wording*is sentences against promoted, and*labels*is field against unlabelled\. Grey bars are not significant at0\.050\.05by an exact sign test over per\-claim differences\. Only*labels*is coloured for both models;*wording*carries the effect on Haiku and does nothing on Sonnet, while*unbracketing*does more on Sonnet\.*Length*is the one property that helps neither model\. Its apparent cost on Sonnet is the one bar here whose sign the no\-model readout reverses, so we do not claim it\. Sixty claims, both hard budgets\.
#### Result\.
Figure[3](https://arxiv.org/html/2608.06953#S7.F3)and Table[3](https://arxiv.org/html/2608.06953#S7.T3)\. We ran the full set of arms on both models, and the comparison is the finding: the two models reach the*same*net effect by*different*routes\.
Table 3:The same effect, assembled differently, under the blind reader\.Percentage\-point change in stance retention from altering one property, per claim, sixty claims, both hard budgets\.∗marks a significant exact sign test across claims\. The net effect at the bottom is near\-identical on the two models; almost none of the components are\. The*length*row is the one whose sign the no\-model readout reverses \(−1\.9\-1\.9on Haiku,\+5\.8\+5\.8on Sonnet\), so its−10\.3\-10\.3is a property of this instrument and we claim only the null\.Three things survive the second model and one does not\.
- •Labels replicate\.\+9\.7\+9\.7points\[\+2\.8,\+16\.1\]\[\+2\.8,\+16\.1\]on Haiku and\+12\.8\+12\.8\[\+7\.2,\+18\.3\]\[\+7\.2,\+18\.3\]on Sonnet, both significant, intervals overlapping\. This is the only component that behaves the same way on both models\. We have been wrong about labels in both directions: an early draft credited them with the whole effect, a later one dismissed them as irrelevant, and the measurement puts them second of five and reproducible\.
- •Length helps on neither model\.\+0\.8\+0\.8\[−2\.5,\+3\.9\]\[\-2\.5,\+3\.9\]on Haiku is a tight null; on Sonnet the same padding scores−10\.3\-10\.3points\[−15\.0,−5\.3\]\[\-15\.0,\-5\.3\],p=0\.0001p=0\.0001under the blind reader\. We claim the null and not the harm: the no\-model readout below puts this contrast at−1\.9\-1\.9on Haiku and\+5\.8\+5\.8on Sonnet, so it is the one component whose*sign*depends on which instrument is asked, and*padding is actively harmful*is not a finding two instruments support\. Making the note longer without making the stance plainer buys nothing; whether it costs is not settled here\.
- •Unbracketing helps on both, by different amounts,\+6\.7\+6\.7and\+12\.8\+12\.8; it reaches significance on Sonnet and not on Haiku\. Note that on Haiku*own line*\(\+6\.1\+6\.1\) is slightly smaller than*unbracketing*\(\+6\.7\+6\.7\) even though it moves layout and bracketing together, which implies the layout part alone is marginally negative there\. The difference is well inside both intervals and we do not read anything into it\.
- •Wording does not replicate\.It is the largest component on Haiku \(\+12\.5\+12\.5,p=0\.0001p=0\.0001\) and worth nothing at all on Sonnet \(\+0\.6\+0\.6,p=0\.77p=0\.77, with a95%95\\%interval of\[−6\.1,\+7\.2\]\[\-6\.1,\+7\.2\]that comfortably contains zero\)\.
So a single\-model ablation would have licensed the wrong summary in either direction\. Run only on Haiku, this section would say the effect is mostly about how explicitly stance is worded\. Run only on Sonnet, it would say wording is irrelevant and the whole thing is brackets\. Both are what one model shows and neither is what two models show together\. The reportable statement is narrower than either:*labels help on both models, length helps on neither, and which of the remaining properties carries the effect depends on the model\.*
#### The same contrasts, scored by an instrument with no model in it\.
Everything above is one Claude\-family judge reading Claude\-family output, which is the objection §[5](https://arxiv.org/html/2608.06953#S5)answers for the two\-cell result and did not, until now, answer here\.experiments/ablation\_mechanical\.pyreruns all seven contrasts with the deterministic gate substituted for the blind reader\. Read the magnitudes as floors: the gate has a documented recall ceiling on open\-class stance and is tilted against any arm carrying stance as a label, so it cannot seeCERTAINTY: rumorat all\. Five of the seven contrasts keep their direction on both models, and the two largest, subordination and the whole asserted path, come out*larger*under the no\-model instrument \(\+11\.7\+11\.7and\+16\.9\+16\.9on Haiku,\+26\.9\+26\.9and\+30\.6\+30\.6on Sonnet\)\. Two do not survive\. Length reverses sign on both models, which is why we claim the null and not the harm above\. The labels contrast reads−1\.1\-1\.1on Haiku and\+0\.3\+0\.3on Sonnet, which is the predicted consequence of an instrument blind to the label form rather than evidence against labels, and we say so rather than resolve it: on our own ledger’s vocabulary, a component that only one instrument can see is not shown by two\. The full table ships indata/results/ablation\_mechanical\.json\.
#### What this means\.
A compressor asked to discard material does not decide by volume\. That much holds on both models and is the firmest thing here: padding buys nothing on either, and the one reading on which it actively costs does not survive a second instrument\. What it does respond to is some property of how plainly the stance is presented, and the two models weight the available ways of being plain differently\. Naming the parts with labels works on both\. Getting the words out of the brackets works on both, more on Sonnet\. Saying the same thing in fuller sentences works on Haiku and does nothing on Sonnet\.
We are deliberately not naming a single mechanism\. The evidence for one would have to look like agreement across models about which property matters, and what we have is agreement about the outcome with disagreement about the route\.
It also accounts for a failure mode we had already found, before this ablation existed\. The claims where the field form loses are the ones whose sources disagree \(§[5](https://arxiv.org/html/2608.06953#S5)\):SOURCEandCERTAINTYare single slots, so “sales said Helsinki, the account plan says Berlin” has to be crushed back into one field, which returns the stance to a subordinate blob\. Two analyses that did not know about each other point at the same mechanism\.
#### What survives for the schema\.
The comparator here is thesentencesarm, which states the source and the certainty in full sentences; we call it*explicit prose*rather than “asserted prose”, because elsewhere in this paper a claim stored as*asserted*is one whose standing has been stripped, which is the opposite of what this arm does\. We could not detect a difference between the schema and explicit prose on either model, and the interval is too wide to call them equivalent\. What we can say is that the schema is not the only way to buy retention\. What the schema still has, and what this ablation does not touch, is that its retention is*checkable*\. A four\-field record can be compared against its source by a program with no model and no lexicon of hedged phrasings \(§[9](https://arxiv.org/html/2608.06953#S9)\)\. Prose that retains stance has to be read by something that can recognise stance in arbitrary wording, which is the open\-class problem that plateaus near half recall\(Kwon,[2026](https://arxiv.org/html/2608.06953#bib.bib5)\)\. Retention and verifiability were conflated in earlier drafts of this work; they are separate, and only one of them is about labels\.
## 8The schema end to end
The two\-cell experiment isolates form\. It does not tell us whether a model asked to*use*the schema does so well enough for the isolated effect to appear in a pipeline\. A pilot suggests it can, on the same model where the two\-cell effect was found: writing the note to the schema and then compressing to ordinary prose raised clean writes from3/103/10to7/107/10, and when the model filled the schema it filled it completely, keeping all four fields in10/1010/10records\. It classified8/108/10sources as something other than plainly asserted \(66rumor,22hedged\), so the recording step works\. This isn=10n=10on one model and is a pilot, not a result\.
#### Independent readout\.
The pilot above is scored by a deterministic checker we also wrote, which is not an independent judge of our own intervention\. The two\-cell experiment therefore uses a different readout entirely: a model that sees only the stored memory and the claim to look up, never the source, the arm, or the pressure level\. That is the readout every number in §[4](https://arxiv.org/html/2608.06953#S4)comes from, and it is the reason we trust the null on Sonnet as much as the positive on Haiku\.
## 9Checking a record
Once standing is a field, checking it stops needing a model\. The hard half of detection is recognising stance in a rewrite, an open\-class problem with no finite vocabulary\. A field does not need recognising\. The released check compares the record against the source with three rules: if the source hedges,CERTAINTYmay not readasserted; if the source names a speaker,SOURCEmay not be empty; if the source dates the claim,AS\_OFmay not be empty\. A record missing its retention fields ismalformed, never a pass\.
#### An instrument note\.
A prose detector scores a schema memory*worse*than a prose memory, becauseCERTAINTY: hedgedpreserves the standing perfectly and contains no word a hedge lexicon knows\. Measured: a schema note compressed into a schema memory scored5/105/10clean under the prose detector against7/107/10for the same schema note compressed into prose, an apparent regression that is entirely an artefact of the instrument\. Inspection of the records showed all four fields intact in10/1010/10cases\. The record is more machine\-readable, not less; it has to be read as a record\.
## 10Conclusion
Write a claim’s standing into a bracketed aside and a compressor treats it as an aside\. Write the same standing as a labelled field and it survives about fifteen points more often, across sixty claims in seven registers, on two models whose estimates differ by five and a half thousandths, and again on sixty further claims written before the run under a prediction registered in advance\.
What we cannot do is tell you why in one sentence, and the reason is worth more than the sentence would have been\. Ablating the format on both models gives the same net effect assembled from different parts\. Labels help on both\. Length helps on neither, and whether padding a note with stance\-free text actively costs depends on which instrument is asked, so we report the null and not the harm\. But the largest single component on Haiku, wording the stance as a full sentence, is worth nothing at all on Sonnet\. Either model on its own would have supported a confident mechanism, and the two mechanisms would have contradicted each other\.
So the advice is the intersection rather than the theory: state the standing explicitly, do not settle for making the note longer, and if you are choosing a write format for a particular model, measure it on that model\. The labelled schema keeps one advantage this experiment cannot take from it, which is that a program can check whether its fields are there at all; prose that retains stance still has to be read by something that can recognise stance in arbitrary wording\.
Most of what we learned here arrived as negatives, and the positive result exists because they were caught\. A ten\-claim corpus in one register manufactured both a false null and a false model\-family split, and we believed the second long enough to verify it against an independent judge and find the check reassuring\. Our own scorer discarded its clearest verdicts and inflated an effect twofold\. A single\-model ablation named a mechanism that the second model declined to confirm\. Every number above is stated at the level the last surviving control allows, and the ledger in Appendix[A](https://arxiv.org/html/2608.06953#A1)lists the nine claims we have withdrawn\.
## Limitations
#### Scope of the models\.
Two models from one provider is not a sample of compressors\. Both were run through the full ablation, which is what let us discover that the decomposition does not transfer; with a third model it might not even be two groups\. We are wary of claims about which models do or do not behave a given way, because an earlier and smaller version of this work produced a confident model\-family split that turned out to be an artefact of its stimuli\. Opus 4\.5 and Sonnet 5 were measured only on the ten\-claim corpus and are therefore not reported as results\.
#### Scope of the stimuli\.
Sixty hand\-written English claims across seven registers, paired with stance\-free fillers and constructed so the two arms could be matched exactly\. That control is what makes the comparison clean and also what makes it artificial: real notes are not matched pairs, and a real memory system writes its own claims rather than receiving ours\. Roughly a third of claims tie on each model, mostly because at the tightest budget both arms lose the claim entirely; ties carry no information for the sign test, so the reportedpp\-values are conservative in that respect and say nothing about those claims\. The padding sentence in the ablation is a single fixed string, though it is hard to see how a different stance\-free sentence could achieve more than the longest arm already failed to achieve\.
#### Scope of the conditions\.
The two budgets reported here are the tight end of a six\-level ladder\. At looser budgets the effect is small and we do not claim it\. The honest statement is that this matters when a memory system is short of room, which is the usual condition but not the only one\. “Standing recovered” is also one model’s judgement of another model’s output: blind to the condition, reproduced by a single fixed judge, and checked against5050hand labels \(κ=0\.75\\kappa=0\.75, §[6](https://arxiv.org/html/2608.06953#S6)\) that are themselves the author’s and whose disagreements are printed rather than summarised\.
#### The ablation is a chain, not a factorial, and it does not transfer\.
The six forms move one thing at a time along a path; they do not independently cross labelling, bracketing, wording and length, so interactions are unmeasured and the component sizes are a ranking under this particular chain rather than coefficients that would survive a full crossing\. More importantly, the ranking is model\-specific: run on Haiku alone it says the effect is about wording, run on Sonnet alone it says wording is irrelevant\. Two models is enough to show the decomposition does not transfer and not enough to say what governs it\.
#### One human read fifty of these, and that human wrote the paper\.
Outcomes are model judgements, checked against a fixed second judge \(98%98\\%\), a deterministic gate, and5050hand labels \(κ=0\.75\\kappa=0\.75, §[6](https://arxiv.org/html/2608.06953#S6)\) whose seven disagreements are printed in Appendix[C](https://arxiv.org/html/2608.06953#A3)rather than summarised away\. The annotator was blind to condition but is the author, so blind to condition is not blind to hypothesis, and one rater yields no inter\-annotator agreement\. The sample also cannot bound the thing it was drawn to check: the interval on the scorer\-swap shift spans about the size of the effect\. Reading the disputed items found the annotator, not the judge, to be the one missing claims, which is a reason to trust the judge slightly more and the hand labels rather less\. The missing control is a larger blinded annotation by people with no stake in the result, reported with inter\-rater agreement\. We could not obtain one, and fifty of the author’s own labels do not stand in for it\.
#### What the schema still cannot do\.
Every failure measured here is a field being*dropped*\. A field confidently filled with the wrong value,CERTAINTY: assertedon a rumour, is a worse failure and is invisible to a checker that only asks whether the field is present; ours catches the subset where the source’s own hedging is detectable\. Detecting stance in the*source*still uses deterministic English lexicons, with every open\-class limit that implies\. And a singleSOURCEslot cannot hold two disagreeing attributions, which is the failure mode of §[5](https://arxiv.org/html/2608.06953#S5); a schema with room for more than one attribution is the obvious repair and we have not tested it\.
#### A concern we raised and then undercut ourselves\.
An earlier draft argued that a multi\-line record costs budget a compressor has to find somewhere, and that this should show up as field records being evicted more often\. It did, once, on one model at one ladder length\. It survives neither the full ladder nor the larger corpus, and the ablation then removed its premise: length does not drive retention at all, so “the record is bulkier” is not obviously a cost\. We leave the concern recorded rather than deleted, because a schema with many more fields than four would test it again\.
#### Cost\.
The end\-to\-end pilot \(n=10n=10, one model\) is a pilot and nothing in the conclusion rests on it\. Metered API spend for everything reported here is under fourteen dollars\.
## Ethics Statement
All stimuli are synthetic, written by the author for this experiment\. They describe fictional colleagues and fictional internal facts, contain no personal data, and were not drawn from any real correspondence\. The failure this work addresses is one where a system overstates its confidence to its own operator; the intervention is a change to how a system records what it was told, and we see no dual\-use concern in releasing it\. Raw runs including every model output are released alongside the code so the analysis can be checked rather than taken on trust\.
## References
- Aikhenvald \(2004\)Alexandra Y\. Aikhenvald\. 2004\.*Evidentiality*\.Oxford University Press, Oxford\.
- Chhikara et al\. \(2025\)Prateek Chhikara, Dev Khant, Saket Aryan, Taranjeet Singh, and Deshraj Yadav\. 2025\.[Mem0: Building production\-ready AI agents with scalable long\-term memory](https://arxiv.org/abs/2504.19413)\.*arXiv preprint arXiv:2504\.19413*\.
- Farkas et al\. \(2010\)Richárd Farkas, Veronika Vincze, György Móra, János Csirik, and György Szarvas\. 2010\.[The CoNLL\-2010 shared task: Learning to detect hedges and their scope in natural language text](https://aclanthology.org/W10-3001/)\.In*Proceedings of the Fourteenth Conference on Computational Natural Language Learning – Shared Task*, pages 1–12, Uppsala, Sweden\. Association for Computational Linguistics\.
- Jiang et al\. \(2023\)Huiqiang Jiang, Qianhui Wu, Chin\-Yew Lin, Yuqing Yang, and Lili Qiu\. 2023\.LLMLingua: Compressing prompts for accelerated inference of large language models\.In*Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing*, pages 13358–13376, Singapore\. Association for Computational Linguistics\.
- Kwon \(2026\)Alex Kwon\. 2026\.[Factwash: Catching ai rewrites that wash hearsay into fact](https://doi.org/10.48550/arXiv.2608.03372)\.*arXiv preprint arXiv:2608\.03372*\.
- Lin et al\. \(2022\)Stephanie Lin, Jacob Hilton, and Owain Evans\. 2022\.[Teaching models to express their uncertainty in words](https://arxiv.org/abs/2205.14334)\.*arXiv preprint arXiv:2205\.14334*\.
- Newell et al\. \(2018\)Edward Newell, Drew Margolin, and Derek Ruths\. 2018\.[An attribution relations corpus for political news](https://aclanthology.org/L18-1524/)\.In*Proceedings of the Eleventh International Conference on Language Resources and Evaluation \(LREC 2018\)*, Miyazaki, Japan\. European Language Resources Association \(ELRA\)\.
- Packer et al\. \(2023\)Charles Packer, Sarah Wooders, Kevin Lin, Vivian Fang, Shishir G\. Patil, Ion Stoica, and Joseph E\. Gonzalez\. 2023\.[MemGPT: Towards LLMs as operating systems](https://arxiv.org/abs/2310.08560)\.*arXiv preprint arXiv:2310\.08560*\.
- Pagnoni et al\. \(2021\)Artidoro Pagnoni, Vidhisha Balachandran, and Yulia Tsvetkov\. 2021\.[Understanding factuality in abstractive summarization with FRANK: A benchmark for factuality metrics](https://doi.org/10.18653/v1/2021.naacl-main.383)\.In*Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies*, pages 4812–4829\. Association for Computational Linguistics\.
- Pareti \(2016\)Silvia Pareti\. 2016\.[PARC 3\.0: A corpus of attribution relations](https://aclanthology.org/L16-1619/)\.In*Proceedings of the Tenth International Conference on Language Resources and Evaluation \(LREC’16\)*, pages 3914–3920, Portorož, Slovenia\. European Language Resources Association \(ELRA\)\.
- Park et al\. \(2023\)Joon Sung Park, Joseph C\. O’Brien, Carrie J\. Cai, Meredith Ringel Morris, Percy Liang, and Michael S\. Bernstein\. 2023\.[Generative agents: Interactive simulacra of human behavior](https://doi.org/10.1145/3586183.3606763)\.In*Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology \(UIST ’23\)*\.
- Rashkin et al\. \(2023\)Hannah Rashkin, Vitaly Nikolaev, Matthew Lamm, Lora Aroyo, Michael Collins, Dipanjan Das, Slav Petrov, Gaurav Singh Tomar, Iulia Turc, and David Reitter\. 2023\.[Measuring attribution in natural language generation models](https://doi.org/10.1162/coli_a_00486)\.*Computational Linguistics*, 49\(4\):777–840\.
- Szarvas et al\. \(2012\)György Szarvas, Veronika Vincze, Richárd Farkas, György Móra, and Iryna Gurevych\. 2012\.[Cross\-genre and cross\-domain detection of semantic uncertainty](https://doi.org/10.1162/COLI_a_00098)\.*Computational Linguistics*, 38\(2\):335–367\.
- Tang et al\. \(2023\)Liyan Tang, Tanya Goyal, Alex Fabbri, Philippe Laban, Jiacheng Xu, Semih Yavuz, Wojciech Kryscinski, Justin Rousseau, and Greg Durrett\. 2023\.[Understanding factual errors in summarization: Errors, summarizers, datasets, error detectors](https://aclanthology.org/2023.acl-long.650/)\.In*Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pages 11626–11644, Toronto, Canada\. Association for Computational Linguistics\.
- Vincze et al\. \(2008\)Veronika Vincze, György Szarvas, Richárd Farkas, György Móra, and János Csirik\. 2008\.[The BioScope corpus: biomedical texts annotated for uncertainty, negation and their scopes](https://doi.org/10.1186/1471-2105-9-S11-S9)\.*BMC Bioinformatics*, 9\(S11\):S9\.
- Wang et al\. \(2025\)Qingyue Wang, Yanhe Fu, Yanan Cao, Shuai Wang, Zhiliang Tian, and Liang Ding\. 2025\.[Recursively summarizing enables long\-term dialogue memory in large language models](https://doi.org/10.1016/j.neucom.2025.130193)\.*Neurocomputing*, 639:130193\.
- Willard and Louf \(2023\)Brandon T\. Willard and Rémi Louf\. 2023\.[Efficient guided generation for large language models](https://arxiv.org/abs/2307.09702)\.*arXiv preprint arXiv:2307\.09702*\.
### Appendix contents
[AClaims and evidence](https://arxiv.org/html/2608.06953#A1)\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.[A](https://arxiv.org/html/2608.06953#A1) [BStimuli and the pressure ladder](https://arxiv.org/html/2608.06953#A2)\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.[B](https://arxiv.org/html/2608.06953#A2) [CEvery item the human and the judge disagreed on](https://arxiv.org/html/2608.06953#A3)\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.[C](https://arxiv.org/html/2608.06953#A3) [DReproducibility](https://arxiv.org/html/2608.06953#A4)\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.[D](https://arxiv.org/html/2608.06953#A4)
## Appendix AClaims and evidence
Every load\-bearing claim, its evidence, and its epistemic status\. Same discipline and same format as the companion paper\(Kwon,[2026](https://arxiv.org/html/2608.06953#bib.bib5)\)\.
- •shown: direct measurement supports it\.
- •retracted:*we*asserted it earlier and later withdrew it\. Three of the nine were at some point this paper’stitle claim, and are marked\(headline\)\.
- •not shown: our measurement neither supports nor refutes it\.
- •narrowed: we asserted it and it survived, but smaller or under fewer conditions than first stated\.
- •not claimed: we never asserted it; the row exists so a reader cannot infer it\.
- •pending: this draft still owes the measurement\.
Table 4:Claims and evidence\.ClaimEvidenceStatusMaking the note longer retains more stance\.§[7](https://arxiv.org/html/2608.06953#S7), Tab\.[3](https://arxiv.org/html/2608.06953#S7.T3)not shown, on two models:\+0\.8\+0\.8points\[−2\.5,\+3\.9\]\[\-2\.5,\+3\.9\]on Haiku and−10\.3\-10\.3\[−15\.0,−5\.3\]\[\-15\.0,\-5\.3\]on Sonnet under the blind readerPadding a note actively costs stance on Sonnet\.§[7](https://arxiv.org/html/2608.06953#S7)narrowed: asserted from the blind reader’s−10\.3\-10\.3\. The no\-model readout puts the same contrast at\+5\.8\+5\.8on Sonnet and−1\.9\-1\.9on Haiku, so it is the one component whose sign depends on the instrument, and the null is all we claimLabels contribute to retention\.§[7](https://arxiv.org/html/2608.06953#S7)shown on both models:\+9\.7\+9\.7\[\+2\.8,\+16\.1\]\[\+2\.8,\+16\.1\]and\+12\.8\+12\.8\[\+7\.2,\+18\.3\]\[\+7\.2,\+18\.3\]\. The only component that behaves the same way on bothThe ablation identifies the mechanism\.§[7](https://arxiv.org/html/2608.06953#S7)not claimed: the two models agree on the net effect and disagree about which property produces it\. Wording is the largest component on Haiku and worth nothing on SonnetLabels are not the mechanism\.\(headline\)§[7](https://arxiv.org/html/2608.06953#S7)retracted twice, in opposite directions: first asserted, then over\-corrected to dismissal, and finally measured on a second model where labels are the*most*robust componentOur consistency gate protects this paper from stating what it has withdrawn\.§[D](https://arxiv.org/html/2608.06953#A4)retracted: it re\-derived every number from the run data and caught none of six internal contradictions found by a reading, because each was a sentence surviving a rewrite rather than a number drifting\. The introduction asserted a mechanism the paper retracts in §[7](https://arxiv.org/html/2608.06953#S7)\. The gate now also checks prose against proseThe field\-over\-bracketed effect replicates on claims chosen before the result was known\.§[6](https://arxiv.org/html/2608.06953#S6), Tab\.[2](https://arxiv.org/html/2608.06953#S6.T2)shown: pre\-registered at a public commit before any call,\+15\.6\+15\.6points\[\+11\.4,\+20\.0\]\[\+11\.4,\+20\.0\],3838claims to11, size prediction heldThe effect survives an instrument with no model in it\.§[6](https://arxiv.org/html/2608.06953#S6)shown: a deterministic gate agrees on direction for both models and on magnitude to three decimals for Haiku, while being biased*against*the field formGrammatical subordination is the mechanism\.\(headline\)§[7](https://arxiv.org/html/2608.06953#S7)retracted: asserted by a previous draft at\+18\.3\+18\.3points, a figure from the four\-level ladder before the sixty\-claim corpus\. On the current data the same path is\+19\.2\+19\.2, and it splits into\+6\.7\+6\.7for unbracketing and\+12\.5\+12\.5for wording; neither figure sums to the\+15\.8\+15\.8net, which walks a different pathWhich property retains stance is the same across models\.§[7](https://arxiv.org/html/2608.06953#S7), Fig\.[3](https://arxiv.org/html/2608.06953#S7.F3)not shown: the two models agree on the net effect and disagree on its composition\. Wording is the largest component on Haiku \(\+12\.5\+12\.5\) and null on Sonnet \(\+0\.6\+0\.6\); unbracketing runs the other wayThe judge manufactures part of the effect\.§[6](https://arxiv.org/html/2608.06953#S6), §[C](https://arxiv.org/html/2608.06953#A3)not shown: on the seven items where a human reader disagreed, six memories contain the claim’s content by a no\-model containment check, so the judge is not matching text that is not there\. Agreement86%86\\%,κ=0\.75\\kappa=0\.75The judge understates the effect, so the paper’s number is conservative\.§[6](https://arxiv.org/html/2608.06953#S6)retracted: drafted from\+21\.2\+21\.2by hand against\+20\.8\+20\.8by the judge\. Five of the seven disagreements are the annotator answering*absent*where the content was present, and six of seven are in one arm, which is what produces a surplus of that size in the otherFifty hand labels validate the model judge\.§[6](https://arxiv.org/html/2608.06953#S6)not claimed: one rater, who is the paper’s author, blind to condition but not to hypothesis, with no inter\-rater agreement\. The scorer\-swap interval\[−15\.9,\+16\.0\]\[\-15\.9,\+16\.0\]does not exclude a shift the size of the effectThe judge errs evenly across the two arms\.§[6](https://arxiv.org/html/2608.06953#S6)not shown: disagreement runs1/261/26on field and6/246/24on prose, Fisher exactp=0\.045p=0\.045, and the errors that touch the difference cancel by arithmetic rather than by designThe effect is a matter of length or token count\.§[7](https://arxiv.org/html/2608.06953#S7)not shown: the longest arm in the experiment, a parenthetical padded with stance\-free text, is worth\+0\.8\+0\.8points,95%95\\%CI\[−2\.5,\+3\.9\]\[\-2\.5,\+3\.9\]\. The interval is narrow as well as containing zero, so this one can be called negligibleThe labelled schema performs as well as explicit prose \(thesentencesarm\)\.§[7](https://arxiv.org/html/2608.06953#S7)not claimed: no difference detected on either model \(−3\.3\-3\.3points on Haiku,p=1\.0p=1\.0;\+1\.9\+1\.9on Sonnet\), but Haiku’s90%90\\%interval\[−9\.2,\+1\.9\]\[\-9\.2,\+1\.9\]fails a two\-one\-sided test against a55\-point margin\. Absence of a detected difference is not equivalenceStance is retained in proportion to how plainly it is stated\.§[7](https://arxiv.org/html/2608.06953#S7)shown in direction, not in composition: every property that makes the stance plainer helps on at least one model, and no property helps equally on bothA labelled record’s retention can be verified without a model\.§[9](https://arxiv.org/html/2608.06953#S9)shownby construction, and it is the schema’s remaining advantage over prose, which this ablation does not touchThe schema handles claims whose sources disagree\.§[5](https://arxiv.org/html/2608.06953#S5)not shown: those are where it loses\. OneSOURCEslot cannot hold two attributions, so they are compressed back into a subordinate clauseStance written as a labelled field survives compression more often than the same stance written as prose\.§[4](https://arxiv.org/html/2608.06953#S4), Fig\.[2](https://arxiv.org/html/2608.06953#S4.F2)shownon6060claims, two models: Haiku\+15\.8\+15\.8points, Sonnet\+15\.3\+15\.3; claim\-clusteredp<0\.001p<0\.001on both;3737vs22and3030vs88claims by signThe effect is a property of claims in general, not of a few wordings\.§[4](https://arxiv.org/html/2608.06953#S4)shown: the sign test over per\-claim differences is the headline test, and it does not depend on effect sizeThe mechanism is that standing is washed off, not that the claim is dropped\.§[4](https://arxiv.org/html/2608.06953#S4)shown for Sonnet and for Haiku at the brutal budget: prose is stored as bare fact28%28\\%vs the field’s11%11\\%while eviction is52%52\\%vs45%45\\%, and on Sonnet eviction is indistinguishable \(p=0\.52p=0\.52\)\.Narrowedon Haiku pooled over both budgets, where the field form is also evicted less \(199199vs228228of360360, McNemarp=0\.0008p=0\.0008\) and about half the stance gain runs through retentionThe effect requires the compressor to be under budget pressure\.§[4](https://arxiv.org/html/2608.06953#S4)shown for the tight end of the ladder; at loose budgets it is small and we do not claim itThe effect disappears once the claim is the unit of analysis\.§[4](https://arxiv.org/html/2608.06953#S4)retracted: asserted on a ten\-claim corpus\. The clustering correction was right; the conclusion drawn from it was an artefact of ten stimuli in one registerThe effect splits by model family, being absent on Sonnet\.§[4](https://arxiv.org/html/2608.06953#S4)retracted: asserted on the full ten\-claim ladder and re\-checked against an independent judge, which agreed\. Sonnet 4\.5 goes from\+3\.7\+3\.7to\+15\.3\+15\.3points on sixty claimsThe schema helps most where the compressor is weakest\.\(headline\)§[4](https://arxiv.org/html/2608.06953#S4)retracted: claimed by an earlier draft on a shorter pressure ladderThe schema costs retention: a bulkier record gets evicted more\.§[4](https://arxiv.org/html/2608.06953#S4)retracted: shown once on one model at one ladder length \(88vs11,p=0\.039p=0\.039\); survives neither the full ladder nor the larger corpusThe advantage widens monotonically as pressure increases\.§[4](https://arxiv.org/html/2608.06953#S4)retracted: predicted in the harness before the first run, not observedOpus 4\.5 and Sonnet 5 show the effect\.§[4](https://arxiv.org/html/2608.06953#S4)not claimed: both were measured only on the ten\-claim corpus, which we no longer trust for this questionA fixed independent judge changes the picture\.§[4](https://arxiv.org/html/2608.06953#S4)not shown:1,9411\{,\}941memories re\-scored by one judge,98%98\\%agreement, no model’s effect moves more than two pointsA write schema raises the share of memory writes that retain standing\.§[8](https://arxiv.org/html/2608.06953#S8)pending: pilot3/10→7/103/10\\to 7/10, one model, one prompt set,n=10n=10The loss occurs at recording rather than at compression\.§[3](https://arxiv.org/html/2608.06953#S3)pending: pilot9/109/10vs1/101/10on one modelOur blind reader measured what we thought it measured\.§[3](https://arxiv.org/html/2608.06953#S3.SS0.SSS0.Px3)not shown for the first runs: a parser bug discarded6%6\\%of verdicts, non\-randomly, and inflated the effect from\+4\.0\+4\.0to\+9\.3\+9\.3points\. Fixed, re\-run, and the pre\-fix data is archived rather than pooledRepairing the note in prose fixes the stored memory\.§[3](https://arxiv.org/html/2608.06953#S3)not shown: clean null in the pilot,0recovered and0brokenField\-level checking needs no model on the memory side\.§[9](https://arxiv.org/html/2608.06953#S9)shownby construction; the released checker uses no modelThis paper’s detector is an independent judge of this paper’s intervention\.§[8](https://arxiv.org/html/2608.06953#S8)not claimed: it is our own gate, which is why a second readout is required
## Appendix BStimuli and the pressure ladder
Sixty claims\. Each is written twice; both forms carry the same assertion and the same stance content, and differ only in where the standing sits\. The claim text itself is a bare assertion containing no stance, so all the stance lives in the source and certainty fields and the two arms can be matched exactly\.
The first ten \(all*project*\) are the corpus this experiment originally ran on and are kept as a distinct subset rather than merged, so the earlier runs stay comparable and the corpus effect of §[4](https://arxiv.org/html/2608.06953#S4)can be measured rather than assumed\. Two examples, one from each end of the register range:
two of the sixtyfieldThe patient is allergic to penicillin fieldSOURCE: the patient’s adult son fieldCERTAINTY: reported, not confirmed against records parenThere were 18 affected accounts \(per support said 18, parenthe data team said 25; conflicting figures\)
Filler notes are deliberately stance\-free and mundane \(“Standup moved to 9:45am on Tuesdays”, “Priya owns the billing service”\)\. They exist to consume budget\. For a given \(claim, pressure, replicate\) both arms receive the same fillers in the same positions, drawn from a seed that does not depend on the arm, so the fillers cannot be what differs\. The full corpus is released as a single Python file\.
The pressure ladder\.Each level names how many notes go in and what budget comes out\.
The sixty\-claim runs usebrutalandsevere, the tight end, because that is where a compressor is actually forced to discard something\. The four looser levels were run on the ten\-claim corpus and are reported in §[4](https://arxiv.org/html/2608.06953#S4)only as the condition under which the effect is small\.
## Appendix CEvery item the human and the judge disagreed on
Fifty stored memories were labelled by hand, blind to arm and model, and the judge disagreed on seven\. Rather than report an agreement statistic and ask to be believed about whether those seven are benign, all of them are below: the claim the annotator was asked to look up, the memory exactly as it was stored, both labels, and the content\-word overlap between claim and memory\. That last column is a no\-model string measure, the same containment the companion gate uses to decide whether a stored claim is matched to a source claim at all, so it is a third instrument on this subset rather than one invented to settle the question\.
Six of the seven memories carry the claim’s content, so the disagreements are about what standing that content is stored with, not about whether the judge matched text that is not there\. Five of the seven are the annotator answering*absent*where the content is present\. Seven items cannot support a statistic and none is offered; the point of printing them is that a reader can disagree with both labellers\.
Claim med\_bp
The patient’s blood pressure has been running near 150 over 95fieldarm, sonnet\-4\-5,brutal human:absent judge:qualified content overlap: 3/8 \(0\.38\)
\- Changelog: release notes \- Staging: staging\.internal\.example\.com \- Billing owner: Priya \- Patient BP ~150/95 \(home readings, uncalibrated\) \- Tuesday standup: 9:45am \- WiFi password: monthly rotation \- CI: self\-hosted runners \- Load tests: tools/loadgenClaim sched\_review
The performance review window opens on the 3rdparenarm, sonnet\-4\-5,severe human:absent judge:established\_fact content overlap: 1/5 \(0\.20\)
Design doc: shared drive/Q3\. Perf reviews: 3rd\. Tuesday standup: 9:45am\. Priya: billing\. Staging: staging\.internal\.example\.com\.Claim proj\_headcount
Two contractor seats were approved for next quarterparenarm, sonnet\-4\-5,severe human:absent judge:qualified content overlap: 5/6 \(0\.83\)
Contractor seats approved next quarter \(pending\)\. Staging: staging\.internal\.example\.com\. CI: self\-hosted runners\. Priya: billing owner\. Load tests: tools/loadgen\. Tuesday standup: 9:45am\.Claim v3\_est\_headcount
The department is around eighty peopleparenarm, haiku\-4\-5,brutal human:absent judge:established\_fact content overlap: 1/4 \(0\.25\)
Release notes→changelog channel\. Staging: staging\.internal\.example\.com\. Office wifi rotates monthly\. ~80 people\. Priya owns billing\. Load tests: tools/loadgen\. Design review: shared drive Q3\. Standup: 9:45am Tuesdays\.Claim est\_attend
Roughly 300 people attendedparenarm, sonnet\-4\-5,severe human:qualified judge:established\_fact content overlap: 1/4 \(0\.25\)
\- ~300 attendees \- WiFi password: monthly rotation \(1st\) \- Staging: staging\.internal\.example\.com \- CI: self\-hosted runners \- Priya: billing owner \- Tuesday standup: 9:45am \- Changelog: release notes \- Load tests: tools/loadgenClaim news\_recall
The manufacturer is preparing a recall of the 2024 unitsparenarm, sonnet\-4\-5,brutal human:absent judge:qualified content overlap: 2/5 \(0\.40\)
Q3 design review in shared drive\. WiFi password rotates monthly\. Changelog for releases\. Tuesday standup 9:45am\. Load tests: tools/loadgen\. Staging: staging\.internal\.example\.com\. Priya: billing\. 2024 recall rumored\.Claim proj\_headcount
Two contractor seats were approved for next quarterparenarm, haiku\-4\-5,severe human:absent judge:established\_fact content overlap: 4/6 \(0\.67\)
Two contractor seats approved Q1\. Staging: staging\.internal\.example\.com\. CI: self\-hosted runners\. Priya: billing service\. Load tests: tools/loadgen\. Standup: Tuesdays 9:45am\.
## Appendix DReproducibility
#### Models\.
The results reported here use Claude Haiku 4\.5 and Sonnet 4\.5\. Opus 4\.5 and Sonnet 5 were also run, on the ten\-claim corpus only, and are not reported as results\. Within a run the compressor and the blind reader are the same model; because that makes each model its own grader, we additionally re\-scored every stored memory with a single fixed judge \(Haiku 4\.5\), which reproduces the same picture at98%98\\%agreement\. The compressor samples at temperature1\.01\.0and the reader at0\.00\.0, except on Sonnet 5, which rejects the temperature parameter outright and therefore sampled at its default; its summaries record this\.
#### Where the4,1044\{,\}104trials come from\.
The denominator quoted for judge\-parse health is every scored trial in the project, not one experiment, and it is worth breaking out because it cannot be reconstructed from the headline design alone:
The ten\-claim rows are uneven because those runs covered different subsets of the pressure ladder as the budget allowed, and because incomplete pairs are dropped rather than averaged\. The1,4401\{,\}440ablation trials per model are scored separately and are not in this table\.
#### What is released\.
The repository is[https://github\.com/collapseindex/factwash](https://github.com/collapseindex/factwash)\(Apache\-2\.0\)\. It contains the harness \(experiments/two\_cell\_pressure\.py\), the scorer \(experiments/two\_cell\_analyse\.py\), the figure script, this paper’s source, and*every raw run*: one JSON line per trial carrying the note, the stored memory, the blind reader’s raw reply, and how that reply was parsed\. Superseded runs are published too, underdata/raw/superseded/with a README explaining why they are not pooled, so the claim that a parser bug inflated the effect from\+4\.0\+4\.0to\+9\.3\+9\.3points can be checked rather than taken on trust\.
The hand annotation ships in the same form: the blank sheet as it was presented \(data/annotation/human\_sheet\.md\), the judge verdicts that were withheld from the annotator until it was filled in \(human\_key\.json\), the sampler and scorer \(experiments/human\_annotate\.py\), and the two analyses \(human\_direction\.py,human\_bootstrap\.py\)\. The stratification seed is recorded in the key file, so the same fifty items can be drawn again and relabelled by somebody else\.
#### Spend, and how it is counted\.
Metered API spend for every run in this paper is under $1414*nominal*\. Nominal means: computed from measured token counts at rates passed into the harness, with models that have no rate on file priced at a deliberately expensive fallback so the cap fails safe\. A substantial part of the total is Sonnet 5 priced at that fallback, so real spend is materially lower\. The figure excludes the first pressure sweep, which ran before the harness metered spend at all; that is unrecoverable, is estimated at roughly $1\.51\.5from its token counts, and is not folded in silently\. Naming the runs rather than the totals: the sixty\-claim two\-cell run cost $0\.560\.56on Haiku and $1\.851\.85on Sonnet, the held\-out replication $0\.550\.55, the four ablation arms $0\.850\.85on Haiku and $3\.723\.72on Sonnet, and the fixed\-judge re\-scoring $0\.720\.72\.
#### Stopping, and what a stop leaves behind\.
Runs are streamed and fsynced per trial, resumable by key, and capped on both call count and measured spend\. A capped run stops collecting and still writes its summary, with the stop recorded instopped\_early\. That is a fix rather than a design: the first version exited straight out of the loop, which meant the runs that spent the most were the ones that recorded no cost at all\. Cells that did not run to completion are dropped rather than averaged in, and the drops are printed\.
#### Guards on this document\.
writeup2/check\_paper\.pyre\-derives every load\-bearing number inmain\.texfrom the summary JSON: the results table cell by cell, the raw counts, the quotedpp\-values, the eviction counts, and the direction of each claim the prose makes\. It also rejects any pair count stated in the prose that the data cannot produce, which is how a stale “860860matched pairs” was found surviving in three places alongside the correct figure\.writeup2/test\_check\_paper\.pythen breaks the paper4848different ways and asserts the gate fires on each\. It has caught five holes in the gate so far, every one of them in the flattering direction; the fifth was that this gate reported a reduced check count as*all consistent*when an input file was missing\.相似文章
压缩即一切——关于长期AI记忆的论文
一篇主张压缩(而非更长的上下文窗口)才是长期AI记忆和关系连续性的关键基元的文章,并以计算中的算力与存储作类比。
尝试让智能体记忆跨会话持久化所学的经验
本文反思了AI智能体记忆的复杂性,远超简单的存储问题,强调了诸如判断真实性、优先级变化、区分决策与噪音以及何时恰当地呈现上下文等挑战。
检索记忆中的时间有效性:消除AI代理在知识演化中的过时事实错误
本文介绍了MemStrata,一种维护时间有效性的检索记忆系统,用于消除AI代理在知识演化中的过时事实错误。它在演化基准测试上优于RAG,同时保持静态召回率,使用确定性替代层而无需LLM调用。
回收评估:有损记忆比空记忆更糟糕
本文表明,具有有损记忆的语言模型如果保留了错误结论而丢弃了证据,会产生自信的错误答案,而空记忆则会导致弃权。作者提出了一种源优先压缩策略,保留可重新计算的来源而非结论,以保持可纠正性,并在多个模型和对话系统中展示了这一机制。
@yoheinakajima: https://x.com/yoheinakajima/status/2081741659260477666
该线程探讨了大脑的双重记忆系统(海马体和新皮层)为构建长期运行的AI智能体提供的启示,指出智能体需要快速的情景捕获机制和缓慢的巩固机制来避免灾难性干扰,而不是仅仅依赖带有临时支撑结构的冻结模型。