Detecting an Effect Is Not Learning to Act on It: A Reward-SNR Floor for LLM Acquisition Agents
Summary
The paper argues that detecting an average effect of an acquired LLM-derived signal is not the same as learning per-instance acquisition policies, and establishes a reward-SNR floor (ρ* ≈ 2.8/√N) below which offline routing is impossible. It introduces Structured Hypothesis Embeddings (SHE) and shows across three datasets that learned per-example acquisition collapses below this floor, recommending design-time regime gates instead.
View Cached Full Text
Cached at: 08/12/26, 08:29 AM
# Detecting an Effect Is Not Learning to Act on It: A Reward–SNR Floor for LLM Acquisition Agents
Source: [https://arxiv.org/html/2608.10441](https://arxiv.org/html/2608.10441)
###### Abstract
Many pipelines can pay a per\-example cost to acquire an auxiliary, model\-derived observation—an LLM’s structured reasoning, a slow oracle, an expensive measurement—and then must decide*when*the acquired signal is worth using\. Our thesis is a distinction that is easy to miss:*detecting*that such a signal helps on average is*not*the same as*learning to act on it*per instance, and a reward–SNR floor governs when the second is even possible\. Even when the acquired signal is genuinely*faithful*and an*in\-sample oracle*that picks the top\-bbexamples by realized reward shows a sizable apparent gain, no*deployable*policy can learn*when*to acquire it: across per\-impression,K=4−64K\{=\}4\{\-\}64cluster, hand\-defined regime, and uplift\-tree granularities, learned routing never beats random, and a matched\-moment i\.i\.d\.\-*noise*placebo reproduces≥100%\\geq\\\!100\\%of the oracle’s apparent gain\. In other words,*the apparent “learnable structure” is order statistics of noise*, not exploitable signal\. We explain this with a single distinction—between*detecting a mean effect*and*learning a per\-instance acquisition policy*—and areward–SNR detectability floor: routing is estimable offline only if the reward effect’s SNRρ=μ/σ\\rho=\\mu/\\sigmaclearsρ⋆\(N\)≈2\.8/N\\rho^\{\\star\}\(N\)\\\!\\approx\\\!2\.8/\\sqrt\{N\}\(equivalentlyN≥Nmin=\(2\.8/ρ\)2N\\geq N\_\{\\min\}=\(2\.8/\\rho\)^\{2\}\), a*necessary*condition we report as such, with a positive control confirming it is a true low\-SNR limit rather than a broken pipeline\. As a concrete instantiation we introduceStructured Hypothesis Embeddings\(SHE\): a frozen LLM decomposes a user history intoKKranked, confidence\-scored, evidence\-grounded intent hypotheses, embedded and fused as an input\-embedding branch of a recommender\. On three public datasets \(MIND, REES46, Amazon\-Beauty\): \(i\) SHE is faithful \(grounded faithfulness\+0\.0705\+0\.0705, distinctiveness2×2\\times, calibratable ECE0\.142→0\.0310\.142\\\!\\to\\\!0\.031\); \(ii\) its downstream*value*is backbone\- and regime\-conditional—significant over an ordered GRU backbone \(\+0\.0114\+0\.0114,95%95\\%CI\[\+0\.0030,\+0\.0209\]\[\+0\.0030,\+0\.0209\]\) yet with a global redundancy gap indistinguishable from zero \(−0\.0005\-0\.0005,\[−0\.0164,\+0\.0150\]\[\-0\.0164,\+0\.0150\]\); and \(iii\)*learned per\-example acquisition collapses at every granularity*because all three datasets sit below the floor \(ρ=0\.048/0\.138/0\.014\\rho\\\!=\\\!0\.048/0\.138/0\.014\)\. The realizable unit is therefore a*design\-time regime gate*, not a learned per\-instance policy; we give an actionable recipe for it\. We release code, a 58\-claim ledger mapping each claim to a script/CSV/figure, and a one\-command reproduction\.
## 1Introduction
A growing number of machine\-learning systems are built around a decision that is usually left implicit:*should we pay to acquire an auxiliary, model\-derived observation for this example, and if so, do we trust it?*The observation might be a large language model’s structured reasoning about an input, a slow but accurate oracle, an additional sensor reading, or a human annotation\. Acquisition has a real cost \(latency, money, compute\), so a natural aspiration is to*learn a policy*—an acquisition agent—that spends the budget only where the auxiliary signal helps\. This aspiration is widespread in active learning, value\-of\-information, learning\-to\-defer, and LLM\-as\-feature pipelines\.
This paper makes a simple but, we argue, under\-appreciated point:*detecting*that the acquired signal helps on average is not the same as*learning to act on it*per instance—and whether that per\-example acquisition policy is even recoverable from offline data is governed by a signal\-to\-noise law that is easy to state and easy to violate\. If the per\-example effect of the acquired signal on the downstream reward has signal\-to\-noise ratioρ=μ/σ\\rho=\\mu/\\sigma, then reliably detecting a policy that conditions on that reward requiresρ\\rhoto exceed a floorρ⋆\(N\)≈2\.8/N\\rho^\{\\star\}\(N\)\\approx 2\.8/\\sqrt\{N\}\. Below the floor, the “obvious” evidence that learning helps—an in\-sample oracle that picks the top\-bbexamples by realized reward—is an artifact of order statistics of noise, not exploitable structure\. We prove the floor with a positive control \(a synthetic signal injected at controllable SNR is recovered by the*same*deployable pipeline onceρ\\rhocrosses the floor\) and we are explicit that the floor is a*necessary*condition, not a claim of impossibility\.
We then ground the abstract “costly observation” in a concrete method,Structured Hypothesis Embeddings\(SHE, §[3](https://arxiv.org/html/2608.10441#S3)\): a frozen LLM turns a user’s interaction history intoKKranked, confidence\-scored, evidence\-grounded hypotheses about the user’s latent intents; these are embedded and fused as an input\-embedding branch of a recommender\. SHE is a faithful, interpretable signal on its own terms \(§[4](https://arxiv.org/html/2608.10441#S4)\), but its downstream value is*backbone\- and regime\-conditional*\(§[5](https://arxiv.org/html/2608.10441#S5)\): it is significant over an ordered sequential backbone yet its global redundancy gap is statistically indistinguishable from zero, with a clean regime split \(absorption in sparse histories, complementarity in long multi\-intent histories\)\. Attempts to*learn when to acquire*SHE fail at every granularity we test \(§[6](https://arxiv.org/html/2608.10441#S6)\), and §[7](https://arxiv.org/html/2608.10441#S7)–[8](https://arxiv.org/html/2608.10441#S8)show why: on all three datasets the reward SNR is below the detectability floor\.
##### An honesty caveat carried throughout\.
The motivating production observation—that structured intent helps most in cold\-start / underdetermined regimes—is*observed, not controlled*\. Our public\-data study is designed to test the*mechanism*\(the SNR floor\) rather than to re\-derive that observation, and we flag every place where an effect is directional\-but\-not\-significant, in\-sample\-only, or below the detectability floor\.
##### Contributions\.
- •A diagnosis: apparent acquisition “learnability” is order statistics of noise\.Across per\-impression, cluster \(K=4−64K\{=\}4\{\-\}64\), regime, and uplift\-tree granularities on three datasets, no deployable policy beats random, and a matched\-moment noise placebo reproduces≥100%\\geq\\\!100\\%of the in\-sample oracle’s apparent gain—so the oracle gap that looks like exploitable structure is not \(§[6](https://arxiv.org/html/2608.10441#S6), §[7](https://arxiv.org/html/2608.10441#S7)\)\.
- •The explanation: mean\-detection≠\\neqpolicy\-learnability, and a reward–SNR detectability floorρ⋆\(N\)≈2\.8/N\\rho^\{\\star\}\(N\)\\approx 2\.8/\\sqrt\{N\}that separates them, with a positive control establishing it is a genuine low\-SNR limit rather than a broken/underpowered pipeline \(§[8](https://arxiv.org/html/2608.10441#S8)\)\.
- •A concrete instantiation, Structured Hypothesis Embeddings: a frozen\-LLM input\-embedding branch of ranked, confidence\-scored, evidence\-grounded hypotheses with a cited\-vs\-non\-cited faithfulness metric; it is*faithful*yet its downstream value is backbone\- and regime\-conditional \(significant over an ordered GRU/SASRec backbone, global redundancy gap indistinguishable from zero\) \(§[3](https://arxiv.org/html/2608.10441#S3)–[5](https://arxiv.org/html/2608.10441#S5)\)\.
- •An actionable prescription: since per\-instance routing is unlearnable below the floor, deploy a*design\-time regime gate*instead; we give a four\-step recipe and validate it at the pooled\-regime level \(§[10](https://arxiv.org/html/2608.10441#S10)\)\.
- •A reproducible artifact: a 58\-claim ledger mapping each claim to a script, CSV and figure, and a one\-command offline reproduction\.
5001000300010000400000\.010\.020\.050\.10\.2undetectable: no learnableper\-instance acquisitiondetectableregionMIND \(Nmin=3403N\_\{\\min\}\{=\}3403,2\.7×2\.7\\timesshort\)Amazon\-Beauty \(Nmin≈40N\_\{\\min\}\{\\approx\}40k,62×62\\timesshort\)REES46 \(powered, but effectnegative\)Sample sizeNNReward SNRρ=μ/σ\\rho=\\mu/\\sigmafloorρ⋆\(N\)=2\.8/N\\rho^\{\\star\}\(N\)=2\.8/\\sqrt\{N\}Figure 1:The reward–SNR detectability floor is the paper’s thesis in one picture\.A costly semantic observation \(here SHE\) can only support a*learned*per\-instance acquisition policy if its downstream reward SNR clearsρ⋆\(N\)=2\.8/N\\rho^\{\\star\}\(N\)\\\!=\\\!2\.8/\\sqrt\{N\}\(necessary mean\-detection condition, §[8](https://arxiv.org/html/2608.10441#S8)\)\. Both content\-rich datasets where SHE is a*faithful*signal sit*below*the floor \(MIND2\.7×2\.7\\times, Amazon\-Beauty62×62\\timesshort ofNminN\_\{\\min\}\), which is why learned acquisition collapses at every granularity\. The one dataset above the floor \(REES46\) is detectable—and there the LLM signal significantly*hurts*downstream AUC\. The bottleneck is not whether the LLM can reason, but whether the downstream reward carries enough signal to learn*when*to acquire that reasoning\.Concretely, we study an*acquisition agent*—a gate that decides, per example and from cheap side\-information only, whether to spend the costly LLM call—and ask when such an agent can be*learned*from offline reward data\. Figure[10](https://arxiv.org/html/2608.10441#S8.F10)states our answer in one picture: a faithful LLM signal is not enough; the downstream reward must clear a reward–SNR floor before any per\-instance acquisition policy is even detectable, and our three datasets sit below it—so the apparent in\-sample oracle gain is order statistics of noise, not learnable structure\. \(We use “agent” for this one\-shot acquisition gate, not a sequential/RL planner\.\)
## 2Problem Setup
Let each exampleii\(an impression / a ranking slate\) have a base predictor using cheap features and an optional*costly observation*oio\_\{i\}obtained by paying a fixed costcc\. Usingoio\_\{i\}changes a downstream rewardRiR\_\{i\}\(here NDCG@10 of the slate\) by a per\-example amountΔi=Ri\(withoi\)−Ri\(without\)\\Delta\_\{i\}=R\_\{i\}\(\\text\{with \}o\_\{i\}\)\-R\_\{i\}\(\\text\{without\}\)\. Writeμ=𝔼\[Δi\]\\mu=\\mathbb\{E\}\[\\Delta\_\{i\}\]andσ=Var\[Δi\]\\sigma=\\sqrt\{\\mathrm\{Var\}\[\\Delta\_\{i\}\]\}, and define the*reward SNR*ρ=μ/σ\\rho=\\mu/\\sigma\. An*acquisition policy*π:xi↦\{0,1\}\\pi:\\;x\_\{i\}\\mapsto\\\{0,1\\\}decides, from cheap side\-informationxix\_\{i\}only, whether to pay foroio\_\{i\}, subject to a budgetb=𝔼\[π\]b=\\mathbb\{E\}\[\\pi\]\. We evaluate policies by the realized system reward at budgetbbagainst a random\-acquisition baseline, with a per\-example bootstrap95%95\\%CI, always out\-of\-fold \(the policy never sees the outcome it is deciding to buy\)\.
Two things are worth separating\. \(a\) The*value*ofoio\_\{i\}if you always acquire it—an average\-treatment question\. \(b\) The*learnability*of*when*to acquire it—a heterogeneous\-policy question\. Our theory \(§[8](https://arxiv.org/html/2608.10441#S8)\) concerns \(b\); our recommender experiments \(§[4](https://arxiv.org/html/2608.10441#S4)–[5](https://arxiv.org/html/2608.10441#S5)\) concern \(a\)\. The two interact: a signal can have real average value yet a per\-example acquisition policy for it can be undetectable\.
##### Datasets\.
We use three public datasets spanning domain×\\timescontent richness \(Table[1](https://arxiv.org/html/2608.10441#S2.T1)\):MIND\[[9](https://arxiv.org/html/2608.10441#bib.bib4)\]\(English news, content\-rich, genuinely multi\-topic histories\),Amazon\-Beauty\(e\-commerce, content\-rich titles but short histories\), andREES46\(e\-commerce sessions, content\-thin:87\.9%87\.9\\%of windows are single\-category\)\. The three differ sharply in history length \(median\|H\|=20/5/short\|H\|=20/5/\\text\{short\}\) and multi\-intent fraction, which is precisely what lets us separate*regime\-dependent*value from a global effect\. All feature tables are cached; no proprietary data is used\.
Table 1:The three public datasets span two axes: domain×\\timescontent richness\.History length, distinct\-category count, and the sparse / multi\-intent slice fractions are recomputed from the cached hypotheses;ρ\\rhois the per\-example reward SNR \(§[7](https://arxiv.org/html/2608.10441#S7)\)\. The spread in history length and multi\-intent fraction is exactly what makes the value*regime\-conditional*\.MIND \(news\)History\(5 items, 3 cats\): \[1\]‘Wheel of Fortune’ guest intro\(tv/tvnews\) \[2\]Helen Hunt hospitalized, car crash\(tv/celebrity\) \[3\]‘Hocus Pocus’ sequel talk\(tv/tvnews\) \[4\]Girl, 7, shot trick\-or\-treating\(news/crime\) \[5\]50 holiday gift ideas under $50\(lifestyle/shop\) Candidates\(click label\): People’s Choice fashion \[0\], 39 party appetizers \[0\], Maren Morris sonogram \[0\],2019 celebrity property roundup \[1\] *Rich text titles; genuine multi\-topic\. Task: rerank candidates by click\.*Amazon\-Beauty \(e\-com, rich\)History\(5 events\): \[1\]view — 18” faux\-locs crochet hair \[2\]view — makeup\-brush cleaner/dryer \[3\]purchase— wax\-warmer waxing kit \[4\]view — \(beauty tool\) \[5\]view — \(beauty tool\) Candidates: 60 ASINs B00HDOZY1G, B00ZGT9O4S, … targetB08DK5D9J5\(*future\_buy=1*\) *Rich product titles \+ view/cart/purchase funnel\. Task: predict next buy\.*REES46 \(e\-com, thin\)Session\(action, category\_code\): \[1\]view — electronics\.smartphone \[2\]view — electronics\.smartphone \[3\]view — electronics\.smartphone \[4\]cart — electronics\.smartphone \[5\]view — electronics\.smartphone Label:*future\_buy*∈\{0,1\}\\in\\\{0,1\\\} *No product text — only category codes; 87\.9% of windows are single\-category\. Task: predict purchase\.* ⇒\\Rightarrowstructurally floors the LLM’s “decompose into facets” step\.
Figure 2:What one raw input record looks like in each dataset\(real, abbreviated\),*before*the frozen LLM produces structured hypotheses\. The three span a2×22\\times 2of domain×\\timescontent\-richness: MIND and Amazon\-Beauty carry rich item text supporting genuine multi\-intent decomposition, whereas REES46 exposes only category codes and is dominated by single\-category sessions\. This contrast is why agent\-side facet metrics are strong on MIND/Amazon but floored on REES46 \(Table[1](https://arxiv.org/html/2608.10441#S2.T1)\)\.
## 3Method: Structured Hypothesis Embeddings \(SHE\)
User historyH=\(e1,…,en\)H=\(e\_\{1\},\\dots,e\_\{n\}\)Frozen LLM\(no grad\)h1h\_\{1\}: celebrity / fitness γ1=0\.66\\gamma\_\{1\}\{=\}0\.66, cite\{3,9,13,14\}\\\{3,9,13,14\\\}h2h\_\{2\}: crime / politics γ2=0\.58\\gamma\_\{2\}\{=\}0\.58, cite\{2,6,9,12\}\\\{2,6,9,12\\\}h3h\_\{3\}: viral food / animals γ3=0\.47\\gamma\_\{3\}\{=\}0\.47, cite\{1,5,10,11,12\}\\\{1,5,10,11,12\\\}Embed𝒆k=ϕ\(hk\)\{\\bm\{e\}\}\_\{k\}=\\phi\(h\_\{k\}\)maxk\\max\_\{k\}γk⋅cos\(𝒄,𝒆k\)\\gamma\_\{k\}\\\!\\cdot\\\!\\cos\(\{\\bm\{c\}\},\{\\bm\{e\}\}\_\{k\}\)Candidate𝒄\{\\bm\{c\}\}Input\-embeddingbranch fusion\[fbase,fB\]\[f\_\{\\text\{base\}\},\\,f\_\{B\}\]Ranker\(NDCG@10\)fbasef\_\{\\text\{base\}\}\(mean\-pool / GRU / SASRec\)
Figure 3:Structured Hypothesis Embedding \(SHE\) pipeline\.A*frozen*LLM decomposes a user history intoK=3K\{=\}3ranked, confidence\-scored \(γk\\gamma\_\{k\}\), evidence\-grounded \(cited history indices\) intent hypotheses\. Each is embedded; a candidate𝒄\{\\bm\{c\}\}scores against the*best\-matching facet*fBmax=maxkγkcos\(𝒄,𝒆k\)f\_\{B\}^\{\\max\}=\\max\_\{k\}\\gamma\_\{k\}\\cos\(\{\\bm\{c\}\},\{\\bm\{e\}\}\_\{k\}\), which is fused as an input\-embedding branch alongside a cheap base backbone \(mean\-pool / GRU / SASRec\)\. The confidencesγk\\gamma\_\{k\}and cited evidence make the branch calibratable and faithfulness\-testable\. Values shown are a real MIND example \(impression 2445, Fig\.[4](https://arxiv.org/html/2608.10441#S4.F4)\)\.Given a user historyH=\(e1,…,en\)H=\(e\_\{1\},\\dots,e\_\{n\}\)\(news reads or product interactions\), a*frozen*LLM is prompted to emitKKintent hypotheses\. Each hypothesishkh\_\{k\}is a short natural\-language statement of a latent interest, and carries \(i\) a calibratable confidenceγk∈\[0,1\]\\gamma\_\{k\}\\in\[0,1\]and \(ii\) a set of*evidence indices*Ek⊆\{1,…,n\}E\_\{k\}\\subseteq\\\{1,\\dots,n\\\}citing the history events that support it\. We useK=3K=3\. This “Scheme B” structured output is contrasted with a “Scheme A” single\-summary baseline\.
Each hypothesis is embedded,𝒆k=ϕ\(hk\)\{\\bm\{e\}\}\_\{k\}=\\phi\(h\_\{k\}\), in a fixed text\-embedding space \(ℓ2\\ell\_\{2\}\-normalized, so similarity is cosine\)\. For a candidate item with embedding𝒄\{\\bm\{c\}\}, the SHE branch produces a small feature vector whose primary coordinate is a confidence\-weighted best\-facet match
fBmax\(𝒄\)=maxk∈\{1,…,K\}γkcos\(𝒄,𝒆k\),f\_\{B\}^\{\\max\}\(\{\\bm\{c\}\}\)\\;=\\;\\max\_\{k\\in\\\{1,\\dots,K\\\}\}\\;\\gamma\_\{k\}\\,\\cos\(\{\\bm\{c\}\},\{\\bm\{e\}\}\_\{k\}\),\(1\)alongsideγ\\gamma\-weighted mean and max variants\. Themax\\max\-over\-facets aggregation is the crux: it lets a candidate match*any*of the user’s disjoint interests rather than a blended average, which is precisely what a single summary \(Scheme A,fA=cos\(𝒄,𝒆summary\)f\_\{A\}=\\cos\(\{\\bm\{c\}\},\{\\bm\{e\}\}\_\{\\text\{summary\}\}\)\) cannot express\. SHE plugs in as an*input\-embedding branch*: the downstream model receives\[fbase\(𝒄\),fB\(𝒄\)\]\[\\,f\_\{\\text\{base\}\}\(\{\\bm\{c\}\}\),f\_\{B\}\(\{\\bm\{c\}\}\)\\,\], wherefbasef\_\{\\text\{base\}\}is the cheap backbone \(mean\-pooled history similarity, ID popularity features, or the state of an ordered sequence model\)\. Fusion is late and the branch is frozen; nothing is back\-propagated into the LLM\.
##### Instantiation ofϕ\\phi\.
The method is agnostic to the choice of text encoderϕ\\phi; we deliberately state all results in terms of the fixed,ℓ2\\ell\_\{2\}\-normalized space rather than a particular model\. In our experimentsϕ\\phiis OpenAItext\-embedding\-ada\-002for MIND and REES46, and a local256256\-dimensional TF\-IDF\+\+TruncatedSVD \(LSA\) space over item titles for Amazon\-Beauty \(an embedding\-API ACL prevented using the same hosted encoder there\)\. Candidate items, hypotheses, and the Scheme\-A summary are always embedded in the*same*space within a dataset, so all comparisons are intra\-space; we never compare cosine values across the ada\-002 and LSA spaces \(§[10](https://arxiv.org/html/2608.10441#S10)\)\. Full encoder, feature\-block, and downstream\-head details are in App\.[B](https://arxiv.org/html/2608.10441#A2)\.
##### Why “structured” and “grounded” matter\.
The evidence indicesEkE\_\{k\}give a*testable*notion of faithfulness \(§[4](https://arxiv.org/html/2608.10441#S4)\): a hypothesis should be more similar to the history events it cites than to those it does not\. The confidencesγk\\gamma\_\{k\}let us calibrate and, in principle, gate\. Both are properties of the*structure*, not of any particular embedding model\.
## 4Agent\-Side Quality of the Hypotheses
MIND impression 2445— a 14\-item, 9\-category news history \(multi\-intent slice\)\. Abbreviated titles:\[1\] free donuts at Shipley; \[2\] Trump offer to British teen’s family; \[3\] Hailey Bieber gym/butt workout; \[4\] emotional nurse photo goes viral; \[5\] Häagen\-Dazs peppermint bark returns; \[6\] Trump sex\-assault claims corroborated; \[7\] Bezos to lose richest\-person crown; \[8\] pollution photos; \[9\] Kevin Spacey not charged; \[10\] baby elephant falls in watering hole; \[11\] MetLife Stadium black cat; \[12\] man killed at Popeyes over chicken sandwich; \[13\] Demi Moore on Ashton Kutcher ”addiction”; \[14\] Kim Kardashian gained 18 lbs\. The frozen LLM returnsthree disjoint, confidence\-ranked, evidence\-groundedhypotheses: ∙\\bulleth1h\_\{1\}\(γ=0\.66\\gamma\{=\}0\.66\):*celebrity / body & fitness / relationships*cites\{3,9,13,14\}\\\{3,9,13,14\\\}∙\\bulleth2h\_\{2\}\(γ=0\.58\\gamma\{=\}0\.58\):*sensational crime & political controversy*cites\{2,6,9,12\}\\\{2,6,9,12\\\}∙\\bulleth3h\_\{3\}\(γ=0\.47\\gamma\{=\}0\.47\):*light viral food & animal\-interest items*cites\{1,5,10,11,12\}\\\{1,5,10,11,12\\\}The facets are*soft, not a hard partition*\(index 9 is cited by bothh1h\_\{1\}andh2h\_\{2\}; index 12 by bothh2h\_\{2\}andh3h\_\{3\}\), and each carries a distinct confidence\. A candidate scores against its*best*matching facet viamaxkγkcos\(𝒄,𝒆k\)\\max\_\{k\}\\gamma\_\{k\}\\cos\(\{\\bm\{c\}\},\{\\bm\{e\}\}\_\{k\}\), so a single\-summary embedding \(Scheme A\) that blends these three interests into one vector cannot express this\.
Figure 4:Qualitative SHE example \(real, unedited\)\.One multi\-topic history yields three grounded intent facets with calibratable confidences — the concrete mechanism behind the\+0\.0705\+0\.0705grounded faithfulness and2×2\\timesdistinctiveness of §[4](https://arxiv.org/html/2608.10441#S4)\.Before asking whether SHE helps a recommender, we ask whether the hypotheses are*good on their own terms*\. All numbers here are computed on MIND \(news, genuine multi\-intent\) and REES46 \(e\-commerce, thin single\-category\); details in Appendix[C](https://arxiv.org/html/2608.10441#A3)\.
##### Faithfulness \(grounded\)\.
We measure the paired difference in cosine similarity between a hypothesis and its*cited*vs\.*non\-cited*history events\. On MIND this paired difference is\+0\.0705\+0\.0705\(95%95\\%CI\[\+0\.068,\+0\.073\]\[\+0\.068,\+0\.073\]\): hypotheses genuinely track the evidence they cite\. \(On REES46, where87\.9%87\.9\\%of windows are single\-category, the difference is≈0\\approx 0, as expected—there is nothing to decompose\.\) Figure[6](https://arxiv.org/html/2608.10441#S4.F6)plots this as a single paired\-Δ\\Deltabar with CI, deliberately*not*as two absolute\-similarity bars, to avoid the common mis\-reading that the non\-cited similarity is zero\.
##### Distinctiveness\.
1−cos¯1\-\\overline\{\\cos\}over hypothesis pairs is0\.2040\.204on MIND versus0\.1040\.104on REES46: on genuinely multi\-intent histories SHE produces disjoint facets \(∼2×\\sim 2\\timesthe separation of the thin\-content dataset\)\.
##### Calibration\.
Raw top\-1 confidence is over\-confident \(ECE0\.1420\.142, Brier0\.1660\.166on REES46\)\. A cross\-fit isotonic map reduces ECE to0\.0310\.031\(−78%\-78\\%; Figure[6](https://arxiv.org/html/2608.10441#S4.F6)\), a standard, honest, deployable post\-hoc fix\. We report per\-bin counts and label the analysis directional at smallNN\.
Figure 5:Agent\-side quality on MIND\. Grounded faithfulness is a paired cited\-minus\-non\-cited cosine difference \(\+0\.0705\+0\.0705\), not an absolute similarity; distinctiveness is2×2\\timesthe thin\-content dataset\.
Figure 6:Confidence is over\-confident but cheaply calibratable: cross\-fit isotonic regression reduces ECE0\.142→0\.0310\.142\\\!\\to\\\!0\.031\(−78%\-78\\%\)\. Bins are equal\-frequency \(∼\\sim83/bin\); small\-NN, directional\.
## 5Downstream Value Is Backbone\- and Regime\-Conditional
Does the SHE branch improve a recommender? The answer depends on*what it is fused onto*\. We first establish the*gradient*\(§[5\.1](https://arxiv.org/html/2608.10441#S5.SS1)\), then a controlled ordered\-vs\-unordered test \(§[5\.2](https://arxiv.org/html/2608.10441#S5.SS2)\)\.
### 5\.1A redundancy gradient over baseline strength
We fuse SHE onto a ladder of increasingly strong base features and measure the\+\+SHE lift in NDCG@10 \(GroupKFold\-by\-impression, class\-balanced pointwise logistic regression, per\-impression bootstrap CI\)\. On MIND the lift descends monotonically as the base gets stronger:L0L\_\{0\}\(subcategory popularity\)\+0\.0161∗\+0\.0161^\{\\ast\},L1L\_\{1\}\(ID pair\)\+0\.0146∗\+0\.0146^\{\\ast\},L2L\_\{2\}\(ID\+\+text\)\+0\.0100∗\+0\.0100^\{\\ast\},L3L\_\{3\}\(strong mean\-pooled text\)\+0\.0094\+0\.0094\(ns\); the sparse slice starts∼3×\\sim 3\\timeshigher at the weak end \(Figure[7](https://arxiv.org/html/2608.10441#S5.F7)\)\. This is the honest core of the “redundancy boundary”: SHE’s marginal value shrinks against a strong content baseline, but does*not*vanish, and is largest exactly where behavioral signal is underdetermined\. A controlled degradation sweep \(Appendix[D](https://arxiv.org/html/2608.10441#A4)\) corroborates the gradient and is explicitly labeled a*diagnostic corruption*, not an achieved real\-world lift\.
Figure 7:Redundancy gradient on MIND\. The\+\+SHE NDCG@10 lift descends monotonically as the base features get stronger \(L0→L3L\_\{0\}\\\!\\to\\\!L\_\{3\}\) and is largest on sparse \(cold\-start\) histories\. The strong\-text rung is not significant; the effect is a gradient, not a chasm\.
### 5\.2A controlled ordered\-vs\-unordered test
To ask specifically whether*ordered sequential access*makes SHE redundant, we run a2×22\\times 2on MIND \(the clean testbed: long, multi\-topic histories, mediann=19n\{=\}19\) holding split, slate, labels and the late\-fusion ranker fixed and varying only two factors: \(unordered mean\-pool vs\. ordered GRU\[[3](https://arxiv.org/html/2608.10441#bib.bib3)\]\)×\\times\(no LLM vs\.\+\+SHE\)\. Because MIND histories are long and multi\-topic, the ordered GRU \(0\.39920\.3992\) is genuinely*stronger*than mean\-pool \(0\.37010\.3701\)—so this is a clean test of ordering, unlike Amazon\-Beauty \(median history55, where the GRU is*weaker*than mean\-pool and thus cannot isolate ordering; we report Amazon in Appendix[E](https://arxiv.org/html/2608.10441#A5)as an additional backbone check, not a clean ordering test\)\.
Table 2:MIND2×22\\times 2\(NDCG@10,N=1263N\{=\}1263;95%95\\%bootstrap CI\)\. The global redundancy gap \(interaction\) is statistically indistinguishable from zero, while the SHE branch remains significant over the*ordered*GRU backbone\. Slice analyses show regime\-dependent absorption vs\. complementarity\.The reading of Table[2](https://arxiv.org/html/2608.10441#S5.T2)is deliberately careful\.Ordered access does not globally absorb the LLM signal: SHE adds a*significant*gain over the ordered GRU \(\+0\.0114\+0\.0114,p=0\.005p\{=\}0\.005\), and the interaction/redundancy gap is*statistically indistinguishable from zero*\(−0\.0005\-0\.0005,p=0\.919p\{=\}0\.919\)\. What is real is a*regime split*\(Figure[8](https://arxiv.org/html/2608.10441#S5.F8)\): the gap is positive on sparse/short histories \(\+0\.033\+0\.033/\+0\.022\+0\.022, absorption—the ordered backbone already captures the little intent present\) and*reverses*on long/multi\-intent histories \(−0\.006\-0\.006/−0\.005\-0\.005, complementarity—SHE contributes semantic structure the sequence model cannot subsume\)\. Thus SHE’s value is a function of*backbone strength*×\\times*history regime*, not of ordering per se\.
Figure 8:Backbone\-conditional value on MIND\. SHE adds value over both an unordered and an ordered backbone; the global redundancy gap is≈0\\approx 0, but slices show absorption in sparse histories and complementarity in long multi\-intent histories\.##### Robustness \(B1–B5\)\.
Five checks defend the finding \(Appendix[F](https://arxiv.org/html/2608.10441#A6)\)\. \(B1\) A 5\-seed sweep: every seed has GRU\>\>mean\-pool and a significant SHE gain over the GRU \(\+0\.011\+0\.011to\+0\.023\+0\.023\)\. \(B2\) A degradation sweep \(full/truncate/shuffle/mean\-pool\): the SHE gain is stable and reliably significant only*over ordered backbones*\. \(B3\) Residualization—projecting SHE onto the sequential state and keeping the residual—retains101%101\\%of the gain \(raw\+0\.0114→\+0\.0114\\toresid\+0\.0115\+0\.0115\); the sequence state explains SHE features withR2≈0\.00R^\{2\}\\\!\\approx\\\!0\.00–0\.010\.01, i\.e\. SHE is largely orthogonal\. \(B4\) A redundancy probe \(predict SHE from the sequence state\) givesR2≤0\.010R^\{2\}\\\!\\leq\\\!0\.010on all slices\. \(B5\) A structurally different ordered backbone,SASRec\[[5](https://arxiv.org/html/2608.10441#bib.bib2)\]\(causal self\-attention\), replicates the pattern: SHE gain over ordered SASRec\+0\.0179\+0\.0179\[\+0\.0076,\+0\.0281\]\[\+0\.0076,\+0\.0281\], same regime split\. We treat SASRec as an*additional ordered\-backbone check, not a uniformly stronger backbone*\(here SASRec0\.37130\.3713is on par with mean\-pool and below the GRU; attention over\-fits atN≈1\.4N\\\!\\approx\\\!1\.4k\)\.
## 6Learning*When*to Acquire Fails at Every Granularity
The recommender results concern the*average*value of always acquiring SHE\. We now ask the acquisition question: can we*learn a per\-example policy*that spends a fixed budget on the impressions where SHE helps most? The answer, across three datasets and the full granularity axis, isno\.
##### Per\-impression\.
A learned out\-of\-fold policy \(predictΔi\\Delta\_\{i\}from cheap features, act out\-of\-sample\) does not beat random acquisition at any budget; neither does an uncertainty or sparse heuristic\. The only “winner” is an*in\-sample*oracle that ranks impressions by realizedΔi\\Delta\_\{i\}—and a matched\-moment i\.i\.d\.\-*noise*placebo reproduces≥100%\\geq 100\\%of that oracle’s apparent gain \(MIND\+0\.0518=\+0\.0518\+0\.0518=\+0\.0518\)\. The oracle is order statistics of noise, not exploitable structure\.
##### Cluster / regime / tree\.
We test representation clusters \(K=4,8,16,32,64K\{=\}4,8,16,32,64via KMeans on the sequence\-state vector\), hand\-defined cross\-product regimes \(sparse×\\timesmulti\-intent×\\timesdeterminacy×\\timesdiversity\), and honest uplift trees, each with per\-fold EB\-shrunk cluster means and label\-free test assignment\.*No tested granularity significantly beats random at any budget*on either MIND or Amazon\-Beauty \(Figure[9](https://arxiv.org/html/2608.10441#S6.F9)\); the granularity curve is flat and near zero\. Crucially, a*positive control*\(Appendix[G](https://arxiv.org/html/2608.10441#A7)\) shows the*same*deployable pipeline*does*recover a synthetic cluster signal once its cluster\-SNR exceeds≈0\.20\\approx 0\.20\(MIND\) /0\.350\.35\(Amazon\); the real data sit at0\.0750\.075/0\.0560\.056, an order of magnitude below\. So the null is a genuine low\-SNR limit, not a broken or underpowered pipeline\. A power analysis confirms the observed\|d\|\|d\|is below each granularity’s own MDE80on both datasets, so we phrase the result as*not detectable at theseNN*, not*impossible*\.
Figure 9:Acquisition collapse across granularity \(MIND\)\. From per\-impression throughK=4K\{=\}4–6464clusters, hand\-defined regimes and uplift trees, no deployable policy significantly beats random \(all CIs cross zero\)\. The in\-sample oracle is largely reproduced by a matched\-noise placebo\.
##### What*is*realizable\.
The regime cell means are sensible and consistent \(dense\-single≈−0\.001\\approx\-0\.001; multi\-intent\+0\.0112\+0\.0112; cold\-start\+0\.0109\+0\.0109\)\. Because pooling averages the reward noise down byn\\sqrt\{n\}, a*design\-time subsystem split*—apply SHE to a cold\-start / multi\-intent subsystem, skip dense\-single—is realizable, even though a*learned per\-example*policy is not\. This is the actionable unit\.
## 7The Mechanism: a Three\-Dataset Reward\-SNR Story
Why does acquisition collapse? Because the per\-example reward effect is buried in noise\. Across three datasets spanning domain×\\timescontent\-richness, the reward SNR is small and the conclusion is identical \(Table[5](https://arxiv.org/html/2608.10441#A9.T5)\):
- •MIND\(content\-rich news\): tiny positive per\-impression lift \(\+0\.0094\+0\.0094\),ρ=0\.048\\rho=0\.048\.
- •REES46\(content\-poor e\-commerce\):≈0\\approx 0/negative per\-window lift \(−0\.0158\-0\.0158; ROC\-AUC0\.840→0\.8330\.840\\\!\\to\\\!0\.833, LLM*hurts*\),ρ=0\.138\\rho=0\.138\.
- •Amazon\-Beauty\(content\-rich e\-commerce\): near\-zero lift \(\+0\.0012\+0\.0012\),ρ=0\.014\\rho=0\.014\(lowest of the three\)\.
In all three, no deployable per\-unit policy beats random and the in\-sample oracle gap is reproduced by a noise placebo\. The acquisition\-limits result is therefore a*mechanism*\(low per\-unit reward SNR\), not a MIND\-specific quirk\. Note we do*not*claim a global backbone redundancy is significant on MIND; the global gap is indistinguishable from zero \(§[5](https://arxiv.org/html/2608.10441#S5)\), while slices suggest regime\-dependent absorption/complementarity\.
## 8The Reward–SNR Detectability Floor
We now state the law that ties the sections together\.
##### Proposition 1 \(reward–SNR detectability floor\)\.
*Consider detecting, at significanceα\\alphaand power1−β1\-\\beta, that the per\-example effectΔi\\Delta\_\{i\}of a costly observation on the reward has positive mean, fromNNexamples with per\-example SNRρ=μ/σ\\rho=\\mu/\\sigma\. Detection requires*
ρ≥ρ⋆\(N\)=z1−α/2\+z1−βN≈2\.8N,N≥Nmin\(ρ\)=\(z1−α/2\+z1−βρ\)2=\(2\.8ρ\)2,\\rho\\;\\geq\\;\\rho^\{\\star\}\(N\)\\;=\\;\\frac\{z\_\{1\-\\alpha/2\}\+z\_\{1\-\\beta\}\}\{\\sqrt\{N\}\}\\;\\approx\\;\\frac\{2\.8\}\{\\sqrt\{N\}\},\\qquad N\\;\\geq\\;N\_\{\\min\}\(\\rho\)=\\left\(\\frac\{z\_\{1\-\\alpha/2\}\+z\_\{1\-\\beta\}\}\{\\rho\}\\right\)^\{2\}=\\left\(\\frac\{2\.8\}\{\\rho\}\\right\)^\{2\},\(2\)*atα=0\.05\\alpha\{=\}0\.05,1−β=0\.81\-\\beta\{=\}0\.8\(soz1−α/2\+z1−β≈1\.96\+0\.84=2\.8z\_\{1\-\\alpha/2\}\+z\_\{1\-\\beta\}\\approx 1\.96\+0\.84=2\.8\)\.*
This is a standard one\-sample mean\-detection bound; its force here is*interpretive*\. A*policy*that learns*when*to acquire must at minimum detect that the conditioned\-on reward has signal; if even the average effect is below the floor, a heterogeneous policy conditioned on the same noisy reward is a fortiori undetectable\. We are explicit that Eq\.[2](https://arxiv.org/html/2608.10441#S8.E2)is a*necessary mean\-detectability condition, not a sufficient policy\-learning / regret bound*, and not an impossibility theorem\.
##### Consistency and dataset placement\.
Two independent routes agree on the floor to within1\.3×1\.3\\times: the closed form givesρ⋆\(1263\)=0\.079\\rho^\{\\star\}\(1263\)=0\.079, while a semi\-synthetic learnability sweep \(inject a feature\-predictable effect at controllable SNR, measure the captured fraction of the tau\-oracle\) locates the threshold near0\.100\.10\. Placing the datasets against Eq\.[2](https://arxiv.org/html/2608.10441#S8.E2)\(Figure[10](https://arxiv.org/html/2608.10441#S8.F10), Table[3](https://arxiv.org/html/2608.10441#S8.T3)\):
Table 3:Datasets against the detectability floor\. We report precise, dataset\-specific gaps rather than an “order of magnitude” shorthand\. “Powered”=1=1means the sample is above the floor\.MIND sits1\.6×1\.6\\timesbelow its SNR floor \(needing∼2\.7×\\sim 2\.7\\timesmore data,Nmin≈3400N\_\{\\min\}\\approx 3400\); Amazon\-Beauty sits7\.9×7\.9\\timesbelow \(needing∼60×\\sim 60\\timesmore data\)\. REES46 is the*one powered*case \(ρ=0\.138\>ρ⋆=0\.126\\rho=0\.138\>\\rho^\{\\star\}=0\.126\)—and there the effect is significantly*negative*\(the LLM branch hurts\)\. This last point matters for honesty: the acquisition collapse is*not*merely a power excuse, because the one dataset with enough power shows the acquired signal is net\-negative, not net\-positive\.
A sufficiency\-side check \(HTE\-SNR, Appendix[H](https://arxiv.org/html/2608.10441#A8)\) closes the “a cleverer heterogeneous policy still wins” loophole: correlation, out\-of\-foldR2R^\{2\}, and a heterogeneity\-SNR against a permutation null are all inside the null band on both datasets—there is no learnable heterogeneity to exploit\.
Figure 10:The reward–SNR detectability floorρ⋆\(N\)≈2\.8/N\\rho^\{\\star\}\(N\)\\approx 2\.8/\\sqrt\{N\}\(80% power\)\. MIND and Amazon\-Beauty sit below the floor; REES46 is above it and its effect is significantly negative\. The floor is a*necessary*condition, not an impossibility theorem\.
## 9Related Work
LLM\-as\-feature for recommendation\.KAR\[[10](https://arxiv.org/html/2608.10441#bib.bib5)\]and RLMRec\[[7](https://arxiv.org/html/2608.10441#bib.bib6)\]augment recommenders with LLM\-derived knowledge or representations\. We differ by \(i\) a*structured, evidence\-grounded, confidence\-scored*hypothesis representation with a testable faithfulness metric, and \(ii\) a focus on*when the signal is \(un\)learnable to acquire*rather than average lift\.Active learning and value of information\[[8](https://arxiv.org/html/2608.10441#bib.bib7)\]learn what to query; we give a detectability floor that governs whether such a policy is estimable offline at all\.Uplift / heterogeneous treatment effects\[[6](https://arxiv.org/html/2608.10441#bib.bib8)\]motivate our per\-exampleΔi\\Delta\_\{i\}estimation and our HTE\-SNR null check\.Selective prediction / learning\-to\-defer\[[1](https://arxiv.org/html/2608.10441#bib.bib9),[4](https://arxiv.org/html/2608.10441#bib.bib11)\]route examples to an abstain/expert option; our acquisition policy is a deferral to a costly LLM observation, and our contribution is the SNR limit on learning that routing\.Calibration\[[2](https://arxiv.org/html/2608.10441#bib.bib10)\]underlies our confidence post\-hoc fix\. Finally, structured LLM hypotheses with confidences were used for*unsupervised*cluster\-geometry scoring in single\-cell gene\-set annotation by our prior HypoGeneAgent\[[11](https://arxiv.org/html/2608.10441#bib.bib1)\]; we reuse ranked hypotheses\+\+confidence but move to a different domain and add a*supervised downstream task*, learned conditional weighting, grounded faithfulness, and the acquisition/SNR analysis\. Net novelty: thereward–SNR detectability floorfor costly semantic acquisition\.
## 10Discussion, Deployment, and Limitations
##### Deployment prescription\.
Do*not*learn per\-instance acquisition from noisy offline rewards\. Use*pre\-specified, design\-time regime gates*: route SHE to \(a\)*cold\-start / underdetermined*histories, where a strong backbone has little to absorb \(absorption\-style value\), and \(b\)*long multi\-intent*histories, where SHE contributes complementary semantic structure \(complementarity value\)\. These are exactly the two regimes where §[5](https://arxiv.org/html/2608.10441#S5)finds value, and they are addressable at design time because pooling averages the reward noise down—unlike a per\-impression policy\. Concretely, a practitioner can follow a four\-step recipe: \(1\) estimate the reward SNRρ\\rhoand check it against the floorNmin=\(2\.8/ρ\)2N\_\{\\min\}=\(2\.8/\\rho\)^\{2\}—ifN<NminN<N\_\{\\min\},*do not*attempt a learned per\-instance router \(it will fit noise order statistics\); \(2\) instead pre\-specify a small number of gates from*cheap, label\-free*slice features \(history length, distinct\-category count, sparsity\), not from the reward; \(3\) enable the costly signal only inside the two value regimes above; \(4\) validate the gate at the*pooled\-regime*level with an out\-of\-fold95%95\\%CI, never per instance\. This turns an unlearnable routing problem into a one\-time subsystem\-placement decision that*is*statistically supported\.
##### Limitations\.
\(1\) The motivating production observation is*observed, not controlled*; our public study tests the mechanism, not that observation\. \(2\) The detectability floor is a necessary mean\-detection condition, not a policy\-learning/regret bound\. \(3\) Downstream hypothesis embeddings on Amazon\-Beauty use a local LSA space \(embedding\-API ACL\), which is internally consistent but not identical to the ada\-002 space used elsewhere; we flag this and avoid cross\-space comparisons \(encoder details in App\.[B](https://arxiv.org/html/2608.10441#A2)\)\. \(4\) SASRec is an additional ordered\-backbone check atN≈1\.4N\\\!\\approx\\\!1\.4k, not a claim of a uniformly stronger backbone\. \(5\) All significance is at the sample sizes reached; several nulls are power\-limited and labeled as such\. Full\-recomputation of a few private on\-pod metrics is labeled*precomputed evidence*in the artifact\.
##### Conclusion\.
Costly semantic observations—LLM structured reasoning being a timely instance—obey a reward–SNR detectability floor\. When a dataset sits below it, learning*when*to acquire is not detectable and its apparent in\-sample gains are noise order statistics; the realizable unit is a design\-time regime gate\. Instantiated as Structured Hypothesis Embeddings, the signal is faithful and its downstream value is backbone\- and regime\-conditional, with a global redundancy gap indistinguishable from zero but a clean absorption / complementarity split\. We hope the floor is a useful, honest yardstick for the growing class of pipelines that pay to think\.
#### Reproducibility Statement
All offline results \(R1–R13, appendix\) are regenerated by a single command,scripts/reproduce\.sh, from committed feature tables \(CPU, no LLM calls\); only from\-scratch hypothesis generation needs an LLM\. A 58\-claim ledger \(results/paper\_claims\.csv\) maps every claim to a script, a persisted CSV, and a figure; Appendix[I](https://arxiv.org/html/2608.10441#A9)reproduces the reviewer\-facing subset\.
##### Author contributions\.
Ying Yuan \(corresponding author,yingyuan238@gmail\.com\) conceived the research question and the core thesis \(that detecting a mean effect is distinct from learning a per\-instance acquisition policy\), formulated the costly\-semantic\- observation problem and the reward–SNR detectability floor, designed and implemented the Structured Hypothesis Embeddings method and the full experimental pipeline \(data processing, hypothesis generation, backbone and acquisition studies, the positive control, and all robustness analyses\), produced all figures and tables, and wrote the manuscript\. Any additional authors and their specific contributions will be recorded here upon joining the project\.
##### Acknowledgments\.
We thank colleagues for discussion and feedback\. This paper uses only publicly available datasets \(MIND, REES46, Amazon\-Beauty\); no proprietary data were used\.
## References
- \[1\]Y\. Geifman and R\. El\-Yaniv\(2017\)Selective classification for deep neural networks\.Advances in Neural Information Processing Systems \(NeurIPS\)\.Cited by:[§9](https://arxiv.org/html/2608.10441#S9.p1.2)\.
- \[2\]C\. Guo, G\. Pleiss, Y\. Sun, and K\. Q\. Weinberger\(2017\)On calibration of modern neural networks\.InInternational Conference on Machine Learning \(ICML\),Cited by:[§9](https://arxiv.org/html/2608.10441#S9.p1.2)\.
- \[3\]B\. Hidasi, A\. Karatzoglou, L\. Baltrunas, and D\. Tikk\(2016\)Session\-based recommendations with recurrent neural networks\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§5\.2](https://arxiv.org/html/2608.10441#S5.SS2.p1.7)\.
- \[4\]H\. Jiang, B\. Kim, M\. Y\. Guan, and M\. Gupta\(2018\)To trust or not to trust a classifier\.Advances in Neural Information Processing Systems \(NeurIPS\)\.Cited by:[§9](https://arxiv.org/html/2608.10441#S9.p1.2)\.
- \[5\]W\. Kang and J\. McAuley\(2018\)Self\-attentive sequential recommendation\.InIEEE International Conference on Data Mining \(ICDM\),Cited by:[§5\.2](https://arxiv.org/html/2608.10441#S5.SS2.SSS0.Px1.p1.13)\.
- \[6\]S\. R\. Künzel, J\. S\. Sekhon, P\. J\. Bickel, and B\. Yu\(2019\)Metalearners for estimating heterogeneous treatment effects using machine learning\.Proceedings of the National Academy of Sciences116\(10\)\.Cited by:[§9](https://arxiv.org/html/2608.10441#S9.p1.2)\.
- \[7\]X\. Ren, W\. Wei, L\. Xia,et al\.\(2024\)Representation learning with large language models for recommendation\.InThe Web Conference \(WWW\),Cited by:[§9](https://arxiv.org/html/2608.10441#S9.p1.2)\.
- \[8\]B\. Settles\(2012\)Active learning\.Morgan & Claypool\.Cited by:[§9](https://arxiv.org/html/2608.10441#S9.p1.2)\.
- \[9\]F\. Wu, Y\. Qiao, J\. Chen,et al\.\(2020\)MIND: a large\-scale dataset for news recommendation\.InAnnual Meeting of the Association for Computational Linguistics \(ACL\),Cited by:[§2](https://arxiv.org/html/2608.10441#S2.SS0.SSS0.Px1.p1.3)\.
- \[10\]Y\. Xi, W\. Liu, J\. Lin,et al\.\(2024\)Towards open\-world recommendation with knowledge augmentation from large language models\.InACM Conference on Recommender Systems \(RecSys\),Cited by:[§9](https://arxiv.org/html/2608.10441#S9.p1.2)\.
- \[11\]Y\. Yuanet al\.\(2025\)HypoGeneAgent: a hypothesis language agent for gene\-set cluster resolution selection using Perturb\-seq datasets\.arXiv preprint arXiv:2509\.09740\.Cited by:[Appendix A](https://arxiv.org/html/2608.10441#A1.p1.1),[§9](https://arxiv.org/html/2608.10441#S9.p1.2)\.
## Appendix ASHE prompt templates and agent configuration
This appendix documents the exact agent configuration used to produce the Structured Hypothesis Embeddings \(SHE\)\. We reuse the structured\-hypothesis format of our prior HypoGeneAgent work\[[11](https://arxiv.org/html/2608.10441#bib.bib1)\]\(a ranked list of natural\-language hypotheses, each with a calibrated confidence and explicit supporting evidence\), but re\-target it from single\-cell gene\-set annotation to per\-user intent over behavioral sequences\. Nothing here is tuned on downstream labels: the language model is a*frozen zero\-shot reasoner*and every prompt below is fixed a priori\.
##### Decoding and model configuration\.
We use two schemes\. Scheme A produces a single blended\-intent summary; Scheme B produces the ranked, evidence\-grounded Top\-3 hypotheses that constitute SHE\. Table[4](https://arxiv.org/html/2608.10441#A1.T4)lists the exact settings\. Both are called through the same credential\-free proxy transport; no fine\-tuning, retrieval, or tool use is involved\.
Table 4:Agent configuration for the two hypothesis\-generation schemes\. Scheme B is the source of the SHE feature block \(App\.[B](https://arxiv.org/html/2608.10441#A2)\)\.
##### Input serialization\.
Each behavioral window is rendered to a numbered plain\-text list, one action per line, so that theevidence\_indicesthe model returns are directly interpretable and can be checked against the cited steps \(the faithfulness measurement in §[4](https://arxiv.org/html/2608.10441#S4)\)\. For e\-commerce \(REES46, Amazon\-Beauty\) each line is`\[i\] <action\>, category: <category\_code\>`; for news \(MIND\) each line is the clicked article title with its category, oldest to newest\. The step index\[i\]is the anchor referenced byevidence\_indices\.
##### Output schema \(Scheme B\)\.
The model must return a single JSON object whosehypothesesarray holds*exactly three*elements, ordered by descending strength/urgency, each with: \(i\)hypothesis— a specific intent statement with a short rationale; \(ii\)confidence— a calibrated probability in\[0,1\]\[0,1\]; and \(iii\)evidence\_indices— the integer input\-step indices that support the hypothesis\. The three fixed rank slots \(with the enforced third\-slot “browsing” fallback below\) are what make SHE a fixed\-width, alignable feature block across users; the four scalar coordinates derived from this block are given in App\.[B](https://arxiv.org/html/2608.10441#A2)\.
##### Boundary rules\.
Three a\-priori rules keep the confidences honest and the facets distinct, and are identical across all datasets in a domain:*\(1\) Weak\-signal fallback*— if the window is uninformative \(e\.g\. all views with no add\-to\-cart, or a very short single\-topic history\), every confidence must be below0\.50\.5and the third hypothesis must be the exact fixed string \(“Just browsing / no clear purchase goal” for e\-commerce, “Casual browsing / no strong topical interest” for news\)\.*\(2\) Strength gate*— only a strong behavioral trigger \(e\.g\. an add\-to\-cart, or, for news, multiple distinct topics\) may push the top hypotheses above0\.70\.7, and the multiple hypotheses must cover genuinely*different*facets rather than restating one topic\.*\(3\) Bias elimination*— the model must never infer absolute gender, age, or region from category names\. These rules are the reason the raw confidences are usable\-but\-overconfident and cheaply post\-hoc calibratable \(ECE0\.142→0\.0310\.142\\to 0\.031, App\.[B](https://arxiv.org/html/2608.10441#A2)\), and the reason Scheme B produces disjoint evidence facets that the distinctiveness metric rewards\.
##### Verbatim prompts\.
The complete system prompts follow\. The serialized window \(above\) is appended as the user turn\.
News — Scheme A \(single summary\)\.
Youareanews\-recommendationandreading\-interestexpert\.Belowisauser’srecent\[readinghistory\]\(newsarticlestheyclicked,oldesttonewest\)\.InONEfree\-textsentenceofatmost30words,giveasinglesharpsummaryoftheuser’scurrentcorereadinginterest\.OutputSTRICTLYaJSONobject:\{"summary":"<yoursummary\>"\}\-\-nomarkdown,noextraexplanation\.
News — Scheme B \(SHE, Top\-3 ranked hypotheses\)\.
Youareatopexpertinnews\-recommendationandreaderpsychology\.Belowisauser’srecent\[readinghistory\]\(clickednewsarticles,oldesttonewest\)\.Basedstrictlyonthereadingevidence,infertheuser’sTop\-3ranked\[Reading\-InterestHypotheses\]\-\-thedistinctthemes/topicstheuserismostlikelytowanttoreadnext\.
OutputSTRICTLYasingleJSONobject:\{"hypotheses":\[\.\.\.\]\}whosearrayholdsEXACTLY3elements,orderedbydescendingstrength\.Nomarkdown\(no‘‘‘json\),noextraexplanation\.Eachelementmustcontainthesefixedkeys:
1\."hypothesis":aspecificreading\-intereststatementwithreasoningaboutthetheme\.
2\."confidence":acalibratedprobabilityscorebetween0\.0and1\.0\.
3\."evidence\_indices":anintegerarrayoftheinputstepindices\(the\[n\]atthestartofeachinputline\)thatsupportthishypothesis\.
\[Hardboundaryrules\]
\-Weak\-signalfallback:ifthehistoryisveryshortorallonenarrowtopic,everyconfidencemustbebelow0\.5,andthe3rdhypothesismustbeexactly"Casualbrowsing/nostrongtopicalinterest"\.
\-Multi\-interestcapture:whenthehistoryspansseveraldistincttopics,thethreehypothesesMUSTcovergenuinelydifferentinterestfacets\(donotrestateonetopic\)\.
\-Biaselimination:neverfabricatetheuser’sabsolutegender,age,orregion\.
E\-commerce — Scheme A \(single summary\)\.
Youareane\-commerceconsumer\-behaviorexpert\.Belowisauser’s\[behaviorsequence\]withinasingleshoppingwindow\.InONEfree\-textsentenceofatmost30words,giveasinglesharpsummaryoftheuser’scurrentcoreshoppingintent\.OutputSTRICTLYaJSONobject:\{"summary":"<yoursummary\>"\}\-\-nomarkdown,noextraexplanation\.
E\-commerce — Scheme B \(SHE, Top\-3 ranked hypotheses\)\.
Youareatopexpertine\-commerceconsumerbehavioranduserpsychology\.Belowisauser’s\[behaviorsequence\]withinasingleshoppingwindow\.Basedstrictlyontheconcretebehavioralevidence,infertheuser’sTop\-3mosturgent\[RankedIntentHypotheses\]\(immediatepurchasemotivations\)\.
OutputSTRICTLYasingleJSONobject:\{"hypotheses":\[\.\.\.\]\}whosearrayholdsEXACTLY3elements,orderedbydescendingurgency\.Nomarkdown\(no‘‘‘json\),noextraexplanation\.Eachelementmustcontainthesefixedkeys:
1\."hypothesis":aspecificintent\-hypothesisstatementwithbehavioral\-motivationreasoning\.
2\."confidence":acalibratedprobabilityscorebetween0\.0and1\.0\.
3\."evidence\_indices":anintegerarrayoftheinputstepindices\(the\[n\]atthestartofeachinputline\)thatsupportthishypothesis\.
\[Hardboundaryrules\]
\-Weak\-signalfallback:iftheinputisALLviews\(view\)withNOadd\-to\-cart,everyconfidencemustbebelow0\.5,andthe3rdhypothesismustbeexactly"Justbrowsing/noclearpurchasegoal"\.
\-Strong\-signaltrigger:onlywhenanadd\-to\-cartactionispresentmaythetoptwohypothesesexceed0\.7confidence\.
\-Biaselimination:neverfabricatetheuser’sabsolutegender,age,orregionfromcategorynames\.
## Appendix BEmbedding and feature\-construction details
This section makes the encoderϕ\\phiand the SHE feature vector fully concrete and reproducible\.
##### Encoders\.
ϕ\\phiis OpenAItext\-embedding\-ada\-002\(d=1536d\{=\}1536, returnedℓ2\\ell\_\{2\}\-normalized so dot product is cosine\) for MIND and REES46, and a local256256\-d TF\-IDF\+\+TruncatedSVD \(LSA\) space fit on item titles for Amazon\-Beauty \(bigram TF\-IDF,≤20\\leq\\\!20k vocab,mindf=2\\min\_\{\\text\{df\}\}\{=\}2, sublinear tf; SVD to256256\-d\), which we thenℓ2\\ell\_\{2\}\-normalize\. Within a dataset every text—candidate items, theKKhypotheses, and the Scheme\-A summary—passes through the*same*ϕ\\phi, so all cosines are intra\-space; we never compare across the ada\-002 and LSA spaces\.
##### What text is embedded\.
An item is embedded from its content string \(MIND: news title; Amazon\-Beauty: product title\)\. A user history is summarized on the backbone side by a*mean\-pool*of its item embeddings \(unordered\) or the final state of a GRU/SASRec over the ordered item embeddings; each hypothesishkh\_\{k\}and the Scheme\-A summary are embedded directly from their LLM\-generated text\. Faithfulness uses the embedded evidence\-category strings the hypothesis cites\.
##### SHE feature block\.
For a candidate𝒄\{\\bm\{c\}\}the frozen branch emits a small fixed\-width feature vector \(no learned parameters inside the branch\):
\[fmax⏟maxkcos\(𝒄,𝒆k\),fmaxγ⏟maxkγkcos\(𝒄,𝒆k\),fmean⏟1K∑kcos\(𝒄,𝒆k\),γk⋆⏟conf\. of best facet\],\\big\[\\;\\underbrace\{f\_\{\\max\}\}\_\{\\max\_\{k\}\\cos\(\{\\bm\{c\}\},\{\\bm\{e\}\}\_\{k\}\)\},\\;\\underbrace\{f\_\{\\max\}^\{\\gamma\}\}\_\{\\max\_\{k\}\\gamma\_\{k\}\\cos\(\{\\bm\{c\}\},\{\\bm\{e\}\}\_\{k\}\)\},\\;\\underbrace\{f\_\{\\text\{mean\}\}\}\_\{\\tfrac\{1\}\{K\}\\sum\_\{k\}\\cos\(\{\\bm\{c\}\},\{\\bm\{e\}\}\_\{k\}\)\},\\;\\underbrace\{\\gamma\_\{k^\{\\star\}\}\}\_\{\\text\{conf\.\\ of best facet\}\}\\;\\big\],wherefmaxγf\_\{\\max\}^\{\\gamma\}is the primary coordinate of Eq\.[1](https://arxiv.org/html/2608.10441#S3.E1)\. This block is concatenated with the backbone scorefbase\(𝒄\)f\_\{\\text\{base\}\}\(\{\\bm\{c\}\}\)\(and, in the Scheme\-A ablation, the single summary matchfA=cos\(𝒄,𝒆summary\)f\_\{A\}\{=\}\\cos\(\{\\bm\{c\}\},\{\\bm\{e\}\}\_\{\\text\{summary\}\}\)\)\.
##### Downstream head and hygiene\.
The concatenated features feed a*late\-fusion*head only—anℓ2\\ell\_\{2\}\-regularized logistic ranker \(C=1\.0C\{=\}1\.0, class\-balanced\) for MIND reranking, and a logistic/Ridge probe for the REES46 acquisition study; nothing is back\-propagated intoϕ\\phior the LLM\. Features are standardized with statistics*fit on training folds only*\(cross\-fit\), and on REES46 the15361536\-d hypothesis embeddings are PCA\-reduced before the low\-NNacquisition probe to avoid overfitting\. All reported cosines and metrics are on theℓ2\\ell\_\{2\}\-normalized vectors above\.
## Appendix CAgent\-side details
Metric is cosine onℓ2\\ell\_\{2\}\-normalized embeddings\. Each impression yields exactlyK=3K\{=\}3Scheme\-B hypotheses\. Faithfulness bootstraps over impressions; distinctiveness averages pairwise1−cos1\-\\cos; calibration uses 5\-fold cross\-fit isotonic regression with equal\-frequency bins\.
## Appendix DControlled degradation sweep
On MIND, injectingN\(0,σ2\)N\(0,\\sigma^\{2\}\)into the base feature and re\-measuring the\+\+SHE lift yields a smooth monotone rise from\+0\.0094\+0\.0094\(ns,σ=0\\sigma\{=\}0\) to\+0\.0775\+0\.0775\(σ=4\\sigma\{=\}4\)\. This is a*diagnostic controlled corruption*illustrating the gradient,*not*an achieved real\-world lift\.
## Appendix EAmazon\-Beauty backbone check
On Amazon\-Beauty \(median history55\) the GRU \(0\.41040\.4104\) is*weaker*than mean\-pool \(0\.47770\.4777\); ordered access does not strengthen the backbone, so Amazon\-Beauty cannot cleanly isolate ordering\. It is reported as an additional backbone\-strength check, consistent with the richness gradient \(SHE significant over the weak ordered GRU, redundant over strong mean\-pool\)\.
## Appendix FBackbone robustness B1–B5
Full tables for the 5\-seed sweep, degradation sweep, residualization \(SHEresid=SHE−Projseq\(SHE\)\\text\{SHE\}\_\{\\text\{resid\}\}=\\text\{SHE\}\-\\mathrm\{Proj\}\_\{\\text\{seq\}\}\(\\text\{SHE\}\), retains101%101\\%\), redundancy probe \(R2≤0\.010R^\{2\}\\leq 0\.010\), and the SASRec second backbone \(\+0\.0179\+0\.0179over ordered SASRec\) are inresults/mind/backbone\_redundancy\_\{2x2,slices,seed\_sweep,residualization,sasrec\}\.csv\.
## Appendix GPositive control for acquisition
Holding the real folds/base/features fixed and replacing the lift with a synthetic cluster signal at controllable cluster\-SNR, the exact deployable policy recovers signal once cluster\-SNR≥0\.20\\geq 0\.20\(MIND\) /0\.350\.35\(Amazon\); real data sit at0\.0750\.075/0\.0560\.056\. Human\-regime and random\-rotated true\-cluster variants shift the threshold but the real data remain far below all of them\.
## Appendix HHTE\-SNR sufficiency check
corr\(s^i,Δi\)\\mathrm\{corr\}\(\\hat\{s\}\_\{i\},\\Delta\_\{i\}\), out\-of\-foldR2R^\{2\}, and a heterogeneity\-SNR against a 200\-draw permutation null are all inside the null band on both datasets \(MINDcorr=−0\.012\\mathrm\{corr\}=\-0\.012,R2=−0\.010R^\{2\}=\-0\.010; Amazoncorr=\+0\.001\\mathrm\{corr\}=\+0\.001,R2=−0\.005R^\{2\}=\-0\.005\): no learnable heterogeneity\.
## Appendix IReviewer\-facing claim table
Table[5](https://arxiv.org/html/2608.10441#A9.T5)is the master cross\-dataset summary\. Each main\-text claim inresults/paper\_claims\.csvlistsclaim\_id, dataset, script, evidence CSV, figure, and whether the public artifact regenerates it \(all R1–R13 offline results: yes; from\-scratch LLM generation: requires gai\-proxy\)\.
Table 5:Master cross\-dataset summary\.d=d=per\-example effect \(instance / best cluster\); MDE=80\{\}\_\{80\}=minimum detectable effect at80%80\\%power; noise\-repro%==fraction of the in\-sample oracle reproduced by a matched\-noise placebo; pos\-ctrl thr==cluster\-SNR at which the pipeline recovers a synthetic signal; real clust\-SNR==measured\.
## Appendix JDatasets, window construction, and preprocessing
All three datasets are public\. We convert each into fixed*windows*— one user history plus a prediction target — and generate hypotheses per window\. Table[6](https://arxiv.org/html/2608.10441#A10.T6)gives the exact construction\. A “window” is a point\-in\-time slice: the history is the observed behavior, and the target is a future click \(MIND candidate slate\), a future purchase \(REES46 session\), or a held\-out next\-item slate \(Amazon\-Beauty\)\. Cold\-start \(*sparse*\) and*multi\-intent*slices are defined by fixed thresholds, not tuned\.
Table 6:Window construction per dataset\. MIND is built with a cohort of 1,600 impressions at stride 5 \(keep 1 of every 5\) and a 40\-candidate slate cap; hypotheses were generated for 1,557 windows \(1,549 valid for both schemes\)\. Amazon\-Beauty: 650 windows \(222 sparse\)\. REES46: 498 session windows\. History length is the number of pre\-target actions; the*sparse*/*multi*slices are the cold\-start and multi\-intent regimes analyzed throughout\.##### Per\-analysisNN\.
The number of*rerankable*windows in a given result can be smaller than the generated count \(windows with an empty valid slate, or missing a valid hypothesis under a scheme, are dropped\)\. We therefore report the exactNNwith each result \(e\.g\. MINDN=1263N\{=\}1263for the ada\-002 late\-fusion study,N=1411N\{=\}1411for the backbone\-conditional LSA study; AmazonN=650N\{=\}650; REES46N=498N\{=\}498\)\. No window is dropped on the basis of its outcome\.
## Appendix KHyperparameters
Every downstream component is deliberately small and CPU\-only; the language model is the sole expensive step \(App\.[L](https://arxiv.org/html/2608.10441#A12)\)\. Table[7](https://arxiv.org/html/2608.10441#A11.T7)lists all settings; none are tuned against the reported test metrics\.
Table 7:Downstream hyperparameters\. The ranker, backbones, and LSA space are fixed a priori; the GRU/SASRec inputs are the frozen LSA item vectors, so only the sequence\-combination parameters are learned\. All CIs use the same impression\-level bootstrap\.
## Appendix LHypothesis\-generation cost and latency
Because the language model is called once per window and never fine\-tuned, the entire method cost is the generation pass; everything downstream \(embedding, ranking, all robustness runs\) is CPU\-seconds\. This asymmetry is exactly what motivates treating a hypothesis as a*costly semantic observation*: the question is not whether to train, but whether it is worth*acquiring*the observation at all\. Table[8](https://arxiv.org/html/2608.10441#A12.T8)reports the measured latency\.
Table 8:Measured generation cost\. Scheme B \(the SHE source\) dominates because of high\-effort reasoning; its latency is variable per window\. End\-to\-end: Amazon\-Beauty’s 650 windows took 284\.8 min on a single worker; MIND’s 1,557 windows were generated across 8 parallel workers \(∼\\sim63 min wall \)\. All downstream analyses run in CPU\-seconds from the cached hypotheses, so results are fully reproducible offline without any further model calls\.##### A design\-time gate trades a small call reduction for equal accuracy\.
Since each acquisition is costly, one might hope to*skip*it on windows unlikely to benefit\. Our floor result \(§[8](https://arxiv.org/html/2608.10441#S8)\) predicts that a*learned per\-instance*router cannot do this reliably; a*design\-time*regime gate \(spend only on the sparse/multi\-intent regimes\) is the prescribed alternative\. On MIND \(N=1263N\{=\}1263, over the strong content baseline\[fhist\]\[f\_\{\\text\{hist\}\}\]\) the gate calls the model on86\.4%86\.4\\%of windows and matches spending everywhere: gate NDCG@10=0\.4554=0\.4554vs\. spend\-everywhere0\.45520\.4552\(difference\+0\.0001\+0\.0001, 95% CI\[−0\.0038,\+0\.0042\]\[\-0\.0038,\+0\.0042\]\)\. Relative to the base ranker the gate is\+0\.0096\+0\.0096\[−0\.0004,\+0\.0194\]\[\-0\.0004,\+0\.0194\]—*not*statistically significant, matching spend\-everywhere \(\+0\.0094\+0\.0094\[−0\.0013,\+0\.0199\]\[\-0\.0013,\+0\.0199\]\)\. We therefore do*not*claim an accuracy gain here: on this strong baseline the design\-time gate’s only realized benefit is a∼\\sim14% reduction in expensive calls at no accuracy loss, and its non\-significance overfhistf\_\{\\text\{hist\}\}is itself consistent with the detectability floor\. The paper’s significant value results are the backbone\-conditional lift over the ordered GRU \(\+0\.0114\+0\.0114,p=0\.005p\{=\}0\.005, §[5](https://arxiv.org/html/2608.10441#S5)\) and the agent\-quality gains \(§[4](https://arxiv.org/html/2608.10441#S4)\);src/regime\_gate\.pyreproduces the numbers above\.Similar Articles
When LLM Reward Design Fails: Diagnostic-Driven Refinement for Sparse Structured RL
This paper frames LLM-generated reward shaping for sparse structured RL as a debugging problem, identifying failure modes like reward flooding and semantic misunderstanding. The authors propose diagnostic-driven iterative refinement, achieving dramatic success rate improvements (e.g., DoorKey-8×8 from 2.3% to 97.6%) compared to one-shot generation.
Spurious Tool Use: When RL Agents Learn the Wrong Reason to Act
This paper studies how reinforcement learning can lead LLM agents to learn spurious tool-use policies based on superficial cues rather than task requirements, and introduces a dense reward method to mitigate this issue.
Representation Without Reward: A JEPA Audit for LLM Fine-Tuning
This paper audits Joint-embedding predictive architectures (JEPA) for LLM fine-tuning on a natural-language-to-regex task, testing twenty-two auxiliary objectives. The results show that hidden-state representation improvements are only weakly coupled to decoded-task accuracy, with no auxiliary surviving family-wise correction.
Multimodal Reward Hacking in Reinforcement Learning
This paper systematically studies reward hacking in reinforcement learning for multimodal LLMs, demonstrating that outcome-only rewards can cause severe failure rates even at large scales, and introducing the Newly Rewarded Failure Rate (NRFR) metric to isolate RL-induced failures.
Rethinking RL for LLM Reasoning: It's Sparse Policy Selection, Not Capability Learning
This paper challenges the assumption that RL teaches new reasoning capabilities to LLMs, arguing instead that it performs sparse policy selection at high-entropy decision points. It introduces ReasonMaxxer, an RL-free method that matches full RL performance with significantly lower training costs.