Fork Where the Model Changes Its Mind: Belief-Shift Branching for Tree-Structured Reinforcement Learning
Summary
The paper proposes belief-shift branching for tree-structured reinforcement learning to enhance step-level credit assignment, demonstrating improved performance in mathematical and coding benchmarks.
View Cached Full Text
Cached at: 09/12/26, 08:22 AM
# Fork Where the Model Changes Its Mind:Belief-Shift Branching for Tree-Structured Reinforcement Learning
Source: [https://arxiv.org/html/2609.11061](https://arxiv.org/html/2609.11061)
Yu LiAffiliation:Salesforce AI ResearchPrafulla Kumar ChoubeyAffiliation:Salesforce AI ResearchJiaxin ZhangAffiliation:Salesforce AI ResearchBecky Xiangyu PengAffiliation:Salesforce AI ResearchQinyuan YeAffiliation:Salesforce AI ResearchKartik NarayanAffiliation:Salesforce AI ResearchCaiwen DingAffiliation:University of MinnesotaSilvio SavareseAffiliation:Salesforce AI ResearchChien\-Sheng WuEmail:[\{bin\.lei,yu\.li,pchoubey,jiaxin\.zhang,becky\.peng\}@salesforce\.com\{qinyuan\.ye,kartik\.narayan,ssavarese,wu\.jason\}@salesforce\.com \{lei00126,dingc\}@umn\.edu](mailto:)Affiliation:Salesforce AI Research
###### Abstract
Tree\-structured rollouts give critic\-free reinforcement learning with verifiable rewards \(RLVR\) step\-level credit: fork a chain at an intermediate point, and sibling outcome differences estimate step value\. Each fork adds sampling cost, so realistic budgets typically allow only a few forks per chain\. A fork placed where the outcome is already largely settled yields siblings that mostly agree and provide almost no credit signal; hence, for a given tree size, where forks are placed largely determines how much step\-level RL can gain\. Most existing mainstream methods place forks by structure, such as fixed lengths, midpoints, and delimiters, or by next\-token entropy\. We formalize fork placement as locating the*pivots*of the chain’s value curve, where the expected outcome turns\. We propose*belief\-shift branching*: read the model’s answer belief at candidate boundaries and fork just before the step where consecutive beliefs diverge most\. Three instantiations, none needing step\-level supervision, span access levels: a black\-box probe, a logit\-lens depth profile, and a learned activation direction, which is fit offline and therefore used only in the validation before RL training\. The signal only*places*forks, and the probe costs about1%1\\%of step compute on mathematics and under5%5\\%on code when it runs inside the rollout engine\. In that validation, against Monte\-Carlo value curves, a belief\-shift signal ranks first in each of the eight model×\\timesbenchmark panels, ahead of entropy, structural, and LLM\-judge baselines\. In RL across three model families and two domains, belief\-shift forking leads every mathematics aggregate, on OLMo\-3\-7B by\+2\.6\+2\.6aggregate and\+2\.9\+2\.9on AIME 2026 over the strongest baseline, and sweeps every OLMo code column, by\+6\.5\+6\.5on LiveCodeBench\-medium\.
## 1Introduction
Figure 1:An OLMo\-3\.1\-32B\-Think chain on AIME 2025; 128 newline boundaries, ten forks per selector;xx\-axis: chain position \(fraction\)\.*Top*: ground\-truthV∗V^\{\\ast\}\(16 MC completions per boundary\); red dotted: the eight value pivots with\|ΔV∗\|≥0\.25\|\\Delta V^\{\\ast\}\|\\\!\\geq\\\!0\.25; shaded: settled region \(V∗≡1V^\{\\ast\}\\\!\\equiv\\\!1\); stars: fork positions\. Flip mass captured by the ten forks,∑forks\|ΔV∗\|\\sum\_\{\\text\{forks\}\}\|\\Delta V^\{\\ast\}\|:Probe\-JS1\.561\.56vs\. entropy0\.440\.44, which spends six forks in the settled region\. Full case: Appendix[F](https://arxiv.org/html/2609.11061#A6)\.Reinforcement learning with verifiable rewards \(RLVR\) trains reasoning models with one scalar per sampled solution\[[11](https://arxiv.org/html/2609.11061#bib.bib9),[33](https://arxiv.org/html/2609.11061#bib.bib8)\]\. Group\-relative methods copy that scalar onto every token, mis\-crediting at both ends: a flawed step inside a luckily\-correct solution is reinforced in full, and a prompt whose rollouts all fail yields zero gradient\[[40](https://arxiv.org/html/2609.11061#bib.bib10),[36](https://arxiv.org/html/2609.11061#bib.bib6)\]\. Tree\-structured rollouts repair this without a critic: fork the trajectory at an intermediate position, and the sibling continuations’ verified outcome difference is a counterfactual estimate of the forked step’s advantage\[[14](https://arxiv.org/html/2609.11061#bib.bib12),[39](https://arxiv.org/html/2609.11061#bib.bib40),[20](https://arxiv.org/html/2609.11061#bib.bib17)\]\. The The catch is cost: each fork adds rollout tokens, so realistic budgets typically allow only a few forks per chain\.
Where should those scarce forks go? In critic\-free tree RLVR this is an open design choice: the tree already converts sibling outcomes into step credit, so what remains is deciding which step the siblings should re\-roll\. Current selectors make this choice without asking what the step*decides*: TreePO and TreeRPO fork on a fixed token grid\[[20](https://arxiv.org/html/2609.11061#bib.bib17),[39](https://arxiv.org/html/2609.11061#bib.bib40)\], earlier step methods cut at delimiters\[[21](https://arxiv.org/html/2609.11061#bib.bib3)\], TreeRL and FR3E fork at next\-token uncertainty peaks, and ARPO at post\-tool entropy rises\[[14](https://arxiv.org/html/2609.11061#bib.bib12),[43](https://arxiv.org/html/2609.11061#bib.bib18),[6](https://arxiv.org/html/2609.11061#bib.bib35)\]\. Uncertainty is the natural candidate, but it tracks freedom of*wording*rather than of*outcome*: entropy is inflated by interchangeable paraphrases\[[17](https://arxiv.org/html/2609.11061#bib.bib37)\], and it reads near zero for any confidently written step, right or wrong, whereas an answer probe moves when the step changes the model’s own answer\. Consistent with this, TreeRL’s surprisal\-based forking \(its “entropy”\) beats*random*forking by2\.12\.1pass\-rate points in its own sampling ablation, and TreePO’s static probability\-guided branching allocation loses to uniform\.
We propose belief\-shift branching: track the model’s belief over the final answer along the chain and fork where that belief shifts most between consecutive candidate boundaries\. We instantiate it with three reads spanning access levels—from output\-only sampling to internal activations—none needing step\-level supervision\. PROBE\-JS is a black\-box read in the spirit of a pop quiz: at each boundary it interrupts the model with a∼\\sim10\-token probe \(“so what is your final answer?”\), elicits the answer distribution, and scores the boundary by the Jensen–Shannon divergence between consecutive distributions, so that a sharp jump marks a pivot or a mistake; it needs only sampling access, at the cost of a short extra generation per boundary\. LENS\-SHIFT is a white\-box read—an X\-ray rather than a quiz—that asks the same question without making the model speak: given the final answer, the logit lens reads how strongly each layer already commits to it, yielding a depth profile at every boundary, and a large change in this profile between boundaries marks a belief shift, at no extra generation\. Belief\-Shift Vector \(BSV\) is a radar calibrated in advance: by contrasting activations under high confidence with those under doubt, it fits offline, before RL, a single activation direction for “confidence shift”, so scoring a boundary is just a projection of its hidden state onto this axis\. Since BSV measures belief shift most directly, we use it to test whether, if the model’s true belief shift could be measured directly, it would track the ground\-truth value curve more accurately\.
The probe costs about1%1\\%of a training step on mathematics and under5%5\\%on code when served by the rollout engine \(Section[3\.6](https://arxiv.org/html/2609.11061#S3.SS6)\), and the signal only*places*forks\. Credit still comes from verified sibling outcomes, so a miscalibrated belief wastes a fork but cannot corrupt training\. Figure[1](https://arxiv.org/html/2609.11061#S1.F1)contrastsProbe\-JSwith entropy on a real chain against a Monte\-Carlo ground\-truth value curve: given ten forks each, the probe’s forks capture3\.6×3\.6\\timesthe value movement of entropy’s \(∑\|ΔV∗\|\\sum\|\\Delta V^\{\\ast\}\|:1\.561\.56vs\.0\.440\.44\), and entropy wastes six forks in the settled region whereV∗≡1V^\{\\ast\}\\\!\\equiv\\\!1\.
*Before*RL, we score selectors against Monte\-Carlo value curves: a belief\-shift signal ranks first in each of the eight model×\\timesbenchmark panels, ahead of entropy, structural, and LLM\-judge baselines, and transfers zero\-shot to an out\-of\-domain benchmark \(Section[4](https://arxiv.org/html/2609.11061#S4)\)\.*During*RL, four arms that differ only in the fork criterion span three model families and two domains under matched budgets: a belief\-shift arm leads every mathematics aggregate \(\+2\.9\+2\.9on OLMo\-3\-7B AIME 2026 over the strongest baseline\) and sweeps every OLMo code column \(Section[6](https://arxiv.org/html/2609.11061#S6)\)\. Ablations cover the tree budget\(M,k\)\(M,k\), policy updater, probe read lengthLL, andLens\-Shiftread\-out, all on OLMo\-3\-7B mathematics, plus model scale on Qwen3\-4B versus Qwen3\-8B \(Section[7](https://arxiv.org/html/2609.11061#S7)\)\.
## 2Background and Related Work
#### Trajectory\-level RLVR and its blind spots\.
Group\-relative RLVR\[[33](https://arxiv.org/html/2609.11061#bib.bib8),[40](https://arxiv.org/html/2609.11061#bib.bib10)\]gives every token of a response the same advantageA^i∝Ri−R¯\\hat\{A\}\_\{i\}\\propto R\_\{i\}\-\\bar\{R\}\. Two pathologies follow: final\-answer rewards leave far more correct solutions with flawed reasoning than process feedback\[[36](https://arxiv.org/html/2609.11061#bib.bib6)\], and all\-tie prompts yield zero gradient\[[40](https://arxiv.org/html/2609.11061#bib.bib10)\]\. Process supervision improves on outcome\-only training in reasoning settings\[[37](https://arxiv.org/html/2609.11061#bib.bib4),[41](https://arxiv.org/html/2609.11061#bib.bib7),[32](https://arxiv.org/html/2609.11061#bib.bib38)\], with up to6×6\\timesbetter RL sample efficiency from process advantage verifiers\[[32](https://arxiv.org/html/2609.11061#bib.bib38)\]\.
#### Three rollout structures, one common input\.
Absent any learned value estimator \(whether an explicit critic or an implicit one recovered from outcome\-trained log\-ratios\[[41](https://arxiv.org/html/2609.11061#bib.bib7),[5](https://arxiv.org/html/2609.11061#bib.bib42)\]\), step values must come from extra sampling: \(a\) Monte\-Carlo completions from step prefixes\[[16](https://arxiv.org/html/2609.11061#bib.bib13),[37](https://arxiv.org/html/2609.11061#bib.bib4),[12](https://arxiv.org/html/2609.11061#bib.bib39)\]; \(b\) MCTS\-style trees, in practice used to label process supervision offline or alongside learned value models\[[10](https://arxiv.org/html/2609.11061#bib.bib31),[42](https://arxiv.org/html/2609.11061#bib.bib14),[25](https://arxiv.org/html/2609.11061#bib.bib5),[4](https://arxiv.org/html/2609.11061#bib.bib15),[8](https://arxiv.org/html/2609.11061#bib.bib16)\]; \(c\) on\-policy branching trees whose node statistics directly yield RL advantages without any value model\[[14](https://arxiv.org/html/2609.11061#bib.bib12),[39](https://arxiv.org/html/2609.11061#bib.bib40),[20](https://arxiv.org/html/2609.11061#bib.bib17),[6](https://arxiv.org/html/2609.11061#bib.bib35)\], our setting\. All three consume the same input: boundaries where a chain may be cut, of which realistic budgets afford only a few per chain\.
Table 1:Existing fork\-point selectors\.Few read the outcome at fork time: grids and delimiters follow surface structure, surprisal and entropy track wording freedom, and the LM judge reads correctness but needs an external generation call per step\.
#### Existing fork selectors are mostly outcome\-agnostic proxies\.
LetV∗\(t\)=Pr\[correct∣x,y≤t\]V^\{\\ast\}\(t\)=\\Pr\[\\text\{correct\}\\mid x,y\_\{\\leq t\}\]; sibling comparison atttestimates the local change ofV∗V^\{\\ast\}, so the ideal selector targets the \(unobservable\) pivots ofV∗V^\{\\ast\}\. Practice approximates them with proxies that never read the outcome \(Table[1](https://arxiv.org/html/2609.11061#S2.T1)\)\.*Delimiters*: PRM800K fine\-tunes the generator to emit newline steps\[[21](https://arxiv.org/html/2609.11061#bib.bib3)\], and surface tokens do not typically mark true decision points\[[23](https://arxiv.org/html/2609.11061#bib.bib41)\]\.*Fixed\-length grids*\[[20](https://arxiv.org/html/2609.11061#bib.bib17),[39](https://arxiv.org/html/2609.11061#bib.bib40)\]fork mid\-expression; TreePO’s static probability\-guided branching allocation over its fixed segments loses to uniform \(their §4\.4\)\.*Next\-token uncertainty*\(TreeRL’s sampled\-token surprisal, FR3E’s entropy, ARPO’s post\-tool entropy\[[14](https://arxiv.org/html/2609.11061#bib.bib12),[43](https://arxiv.org/html/2609.11061#bib.bib18),[6](https://arxiv.org/html/2609.11061#bib.bib35)\]\) measures wording, not outcome \(Section[1](https://arxiv.org/html/2609.11061#S1)\), and beats*random*forking by2\.12\.1pass\-rate points in TreeRL’s own sampling ablation at the same tree configuration, with about8%8\\%fewer generated tokens\. SPO\[[12](https://arxiv.org/html/2609.11061#bib.bib39)\]already moves Monte\-Carlo cutpoints off a fixed grid, placing them at low\-probability tokens on the premise that a boundary is worth sampling only where the value may change; we share the premise, but its read is again the probability of the sampled token, whereas ours is the model’s belief over the answer\.*LM judges*are offline and costly\[[19](https://arxiv.org/html/2609.11061#bib.bib34)\], and learned rewards are exploitable during RL\[[11](https://arxiv.org/html/2609.11061#bib.bib9)\];*agent turns*\[[44](https://arxiv.org/html/2609.11061#bib.bib32),[30](https://arxiv.org/html/2609.11061#bib.bib33)\]costKTK^\{T\}rollouts if branched exhaustively overTTturns\.
## 3Belief\-Shift Branching
#### Template\.
A fork\-point selector is a scoring rule over candidate boundaries: given a chainyywith poolℬ\(y\)\\mathcal\{B\}\(y\)\(Section[3\.1](https://arxiv.org/html/2609.11061#S3.SS1)\), it assigns a scorests\_\{t\}to each boundaryt∈ℬ\(y\)t\\in\\mathcal\{B\}\(y\)and forks at the top scorert^=argmaxtst\\hat\{t\}=\\arg\\max\_\{t\}s\_\{t\}\. Our selectors all take one form: read the model’s current*answer belief*at each boundary, and score the segment between consecutive boundaries by how far it moved that belief,st−=D\(bt−,bt\)s\_\{t^\{\-\}\}=D\(b\_\{t^\{\-\}\},b\_\{t\}\), wheret−t^\{\-\}is the boundary immediately precedingttinℬ\(y\)\\mathcal\{B\}\(y\)andDDis an instantiation\-specific measure of belief change \(a divergence or a profile\-change norm; the learned read of Section[3\.4](https://arxiv.org/html/2609.11061#S3.SS4)instead projects the single boundary statehth\_\{t\}onto a fitted belief\-shift direction\)\. Note the*fork\-before*attribution: the belief change is caused by the segment\(t−,t\]\(t^\{\-\}\\\!,t\], so the score lands ont−t^\{\-\}and siblings re\-roll exactly the belief\-moving step\. Three instantiations vary only in howbtb\_\{t\}is read:Probe\-JSneeds sampling access andLens\-Shiftreads hidden states, while BSV amortizes the read into a direction fit offline, which confines it to pre\-RL validation—testing whether, if the model’s true belief shift could be measured directly, it would track the ground\-truth value curve more accurately \(Section[3\.4](https://arxiv.org/html/2609.11061#S3.SS4)\)\.
### 3\.1Candidate step boundaries
Mid\-word forks yield siblings that differ for trivial lexical reasons, so candidates are restricted to*step boundaries*: positions after a chosen delimiter, here the line break \(\\n\), the natural step separator in both tasks we test \(mathematics and code\)\. At mostC=16C\{=\}16evenly spaced candidates are scored per chain; a chain whose span has no boundary is cut at the span midpoint instead, so it still forks\. Only a chain that cannot fork at all \(fewer than two tokens past the last cut, or no response budget left for a continuation\) is replaced by a fresh root chain \(Section[3\.5](https://arxiv.org/html/2609.11061#S3.SS5)\), so every arm always produces its full tree\.
Figure 2:The three belief reads:Probe\-JScompares consecutive elicited answer distributions \(JS\);Lens\-Shifttakes theℓ2\\ell\_\{2\}change of the teacher\-forced answer’s per\-layer log\-prob profile; BSV projectshth\_\{t\}onto an offline\-fit surge/steady/drop direction \(pre\-RL only\)\.
### 3\.2Probe\-JS: black\-box belief probing
At boundaryttwe append a short elicitation suffixeethat asks for the final answer \(“So the final answer is…”;E≈8E\{\\approx\}8–99tokens\) to the prefixy<ty\_\{<t\}and read the next\-token distribution at the end of the suffix: the beliefptp\_\{t\}is the top\-KK\(K=20K\{=\}20\) log\-probabilities ofπθ\(⋅∣x,y<t,e\)\\pi\_\{\\theta\}\(\\cdot\\mid x,y\_\{<t\},e\), i\.e\. the model’s distribution over the first answer token\. Instantiating the template score with the Jensen–Shannon divergence\[[22](https://arxiv.org/html/2609.11061#bib.bib2)\]:
st−=JS\(pt−∥pt\)=12DKL\(pt−∥m\)\+12DKL\(pt∥m\),m=pt−\+pt2\.s\_\{t^\{\-\}\}\\;=\\;\\mathrm\{JS\}\\\!\\left\(p\_\{t^\{\-\}\}\\,\\\|\\,p\_\{t\}\\right\)\\;=\\;\\tfrac\{1\}\{2\}D\_\{\\mathrm\{KL\}\}\\\!\\left\(p\_\{t^\{\-\}\}\\\|\\,m\\right\)\+\\tfrac\{1\}\{2\}D\_\{\\mathrm\{KL\}\}\\\!\\left\(p\_\{t\}\\\|\\,m\\right\),\\qquad m=\\tfrac\{p\_\{t^\{\-\}\}\+p\_\{t\}\}\{2\}\.\(1\)Unlike KL, the mixturemmmakes the score symmetric,JS\(p∥q\)=JS\(q∥p\)\\mathrm\{JS\}\(p\\\|q\)=\\mathrm\{JS\}\(q\\\|p\), and bounded,0≤JS≤log20\\leq\\mathrm\{JS\}\\leq\\log 2: segments are ranked by movement alone, on one scale, with no reference endpoint, whereasDKL\(pt−∥pt\)→∞D\_\{\\mathrm\{KL\}\}\(p\_\{t^\{\-\}\}\\\|\\,p\_\{t\}\)\\to\\inftyaspt\(v\)→0p\_\{t\}\(v\)\\to 0for anyvvwithpt−\(v\)\>0p\_\{t^\{\-\}\}\(v\)\>0, so under top\-KKtruncation a single tail token that drops out of one support can own the argmax; the two\-sided form also yields the bound of Eq\.[9](https://arxiv.org/html/2609.11061#S3.E9)\. The truncation is an engineering constraint, not a modeling choice: vLLM\-class engines and OpenAI\-style APIs return at most the top\-KKnext\-token log\-probabilities, with2020the default ceiling, henceK=20K\{=\}20\. WithStS\_\{t\}the returned support atttandS=St−∪StS=S\_\{t^\{\-\}\}\\\!\\cup S\_\{t\}, each side is floored and renormalized onSS,
p~t\(v\)∝\{pt\(v\),v∈St,12minu∈Stpt\(u\),v∈S∖St,\\tilde\{p\}\_\{t\}\(v\)\\;\\propto\\;\\begin\{cases\}p\_\{t\}\(v\),&v\\in S\_\{t\},\\\\\[1\.0pt\] \\tfrac\{1\}\{2\}\\,\\min\_\{u\\in S\_\{t\}\}p\_\{t\}\(u\),&v\\in S\\setminus S\_\{t\},\\end\{cases\}\(2\)and Eq\.[1](https://arxiv.org/html/2609.11061#S3.E1)is evaluated on\(p~t−,p~t\)\(\\tilde\{p\}\_\{t^\{\-\}\},\\tilde\{p\}\_\{t\}\)\. Each probe rides the rollout’s prefix cache, forwarding only the suffix plus theLLdecoded tokens \(L=1L\{=\}1for mathematics,1616for code; ablated in Section[7](https://arxiv.org/html/2609.11061#S7)\)\. For code the suffix isSo the final solution is:followed by an openingpythonfence, so the “answer” is the opening of the solution program\. ForL\>1L\{\>\}1the engine samplesLLtokens after the suffix at temperature1\.01\.0and returns the top\-KKdistribution at each position; Eq\.[1](https://arxiv.org/html/2609.11061#S3.E1)is applied position by position between consecutive boundaries and averaged over theLLpositions, soL=1L\{=\}1is the special case above\. Position one is the exactL=1L\{=\}1read; positions22toLLare conditioned on that boundary’s own sampled read\-out, so the score adds a sampled component on top of theL=1L\{=\}1signal\. The same read defines theLLablation of Section[7](https://arxiv.org/html/2609.11061#S7)\.
### 3\.3Lens\-Shift: white\-box belief reading
With activation access, beliefs are read from inside the forward pass\. The logit lens\[[28](https://arxiv.org/html/2609.11061#bib.bib1)\]defines per\-layer predictionsP\(ℓ\)\(⋅\)=softmax\(WULN\(h\(ℓ\)\)\)P^\{\(\\ell\)\}\(\\cdot\)=\\mathrm\{softmax\}\\\!\\big\(W\_\{U\}\\,\\mathrm\{LN\}\(h^\{\(\\ell\)\}\)\\big\), whereh\(ℓ\)h^\{\(\\ell\)\}is the residual stream after layerℓ\\ellandLN\\mathrm\{LN\},WUW\_\{U\}are the model’s own final norm and unembedding\. After the prefixy<ty\_\{<t\}and the elicitation suffix of Section[3\.2](https://arxiv.org/html/2609.11061#S3.SS2), we teacher\-force the chain’s committed answera=\(a1,…,an\)a=\(a\_\{1\},\\dots,a\_\{n\}\)\(its final answer,nntokens\) and record how strongly each intermediate layer already encodes that answer \(no sampling is involved\), giving a*depth profile*and its change score
bt\(ℓ\)=∑i=1nlogP\(ℓ\)\(ai\|x,y<t,e,a<i\),st−shape=∥bt\(⋅\)−bt−\(⋅\)∥2,b\_\{t\}^\{\(\\ell\)\}\\;=\\;\\sum\_\{i=1\}^\{n\}\\log P^\{\(\\ell\)\}\\\!\\left\(a\_\{i\}\\,\\middle\|\\,x,\\,y\_\{<t\},\\,e,\\,a\_\{<i\}\\right\),\\qquad s^\{\\mathrm\{shape\}\}\_\{t^\{\-\}\}=\\bigl\\\|\\,b\_\{t\}^\{\(\\cdot\)\}\-b\_\{t^\{\-\}\}^\{\(\\cdot\)\}\\bigr\\\|\_\{2\},\(3\)withxxthe prompt andbt\(⋅\)b\_\{t\}^\{\(\\cdot\)\}the vector collectingbt\(ℓ\)b\_\{t\}^\{\(\\ell\)\}over all probed layers\. The profile traces the answer’s formation across depth, so its change localizes*where*a step moved the belief;sshapes^\{\\mathrm\{shape\}\}fires when that path reorganizes at any depth\. Two scalar read\-outs,
st−height=\|bt¯−bt−¯\|,st−depth=\|ℓt∗−ℓt−∗\|,ℓt∗=min\{ℓ:\|bt\(ℓ\)−bt\(ℓmax\)\|≤ε⋅rangebt\(⋅\)\},s^\{\\mathrm\{height\}\}\_\{t^\{\-\}\}=\\bigl\|\\overline\{b\_\{t\}\}\-\\overline\{b\_\{t^\{\-\}\}\}\\bigr\|,\\quad s^\{\\mathrm\{depth\}\}\_\{t^\{\-\}\}=\\bigl\|\\ell^\{\\ast\}\_\{t\}\-\\ell^\{\\ast\}\_\{t^\{\-\}\}\\bigr\|,\\quad\\ell^\{\\ast\}\_\{t\}=\\min\\Bigl\\\{\\ell:\\,\\bigl\|b^\{\(\\ell\)\}\_\{t\}\-b^\{\(\\ell\_\{\\max\}\)\}\_\{t\}\\bigr\|\\leq\\varepsilon\\cdot\\operatorname\{range\}b^\{\(\\cdot\)\}\_\{t\}\\Bigr\\\},\(4\)\(b¯\\overline\{b\}: profile mean over the deepest half of the layers;ℓmax\\ell\_\{\\max\}: last layer;ε=0\.1\\varepsilon\{=\}0\.1\) read the settled belief’s strength and its lock\-in depth\. All RL arms usesshapes^\{\\mathrm\{shape\}\}\(sheights^\{\\mathrm\{height\}\}/sdepths^\{\\mathrm\{depth\}\}: the read\-out ablation of Section[7](https://arxiv.org/html/2609.11061#S7)\); the cost is one teacher\-forced forward per candidate, no generation\.
### 3\.4Belief\-shift vector: a learned read for pre\-RL validation
BSV is not an RL fork selector but a validation instrument: it tests whether forking at the model’s true belief shifts, read as directly as activation access allows, reconstructs the ground\-truth value curve more accurately than the proxies of Table[1](https://arxiv.org/html/2609.11061#S2.T1)\. To make that read as direct as possible, we amortize the per\-boundary belief read into one dot product \(Figure[2](https://arxiv.org/html/2609.11061#S3.F2)\), fit once, offline, on a*frozen*policy\. A held\-out chain is cut at a boundaryttand continued in three LLM\-drafted continuation regimes \(Appendix[A\.1](https://arxiv.org/html/2609.11061#A1.SS1)\), indexedg∈\{\+,0,−\}g\\in\\\{\+,0,\-\\\}:*surge*\(the continuation commits to the answer\),*steady*\(reasoning proceeds with no change of belief\), and*drop*\(it retracts the answer or expresses doubt\)\. Per layerLLand regimegg, the residual stream is averaged over aWa=16W\_\{a\}\{=\}16\-token window after the cut and aWb=16W\_\{b\}\{=\}16\-token window before it and differenced, then the steady\-regime delta is subtracted:
Δg\(L\)=𝔼\(t,c\):g\(c\)=g\[h¯\[t,t\+Wa\)\(L\)\(c\)−h¯\[t−Wb,t\)\(L\)\],u±\(L\)=Δ±\(L\)−Δ0\(L\)\.\\Delta^\{\(L\)\}\_\{g\}=\\mathop\{\\mathbb\{E\}\}\_\{\(t,c\):\\,g\(c\)=g\}\\Bigl\[\\bar\{h\}^\{\(L\)\}\_\{\[t,\\,t\+W\_\{a\}\)\}\(c\)\-\\bar\{h\}^\{\(L\)\}\_\{\[t\-W\_\{b\},\\,t\)\}\\Bigr\],\\quad u^\{\(L\)\}\_\{\\pm\}=\\Delta^\{\(L\)\}\_\{\\pm\}\-\\Delta^\{\(L\)\}\_\{0\}\.\(5\)Δ0\\Delta\_\{0\}removes the progress shared by all continuations, leaving pure belief\-change axesu±u\_\{\\pm\}; their antagonism picks the layer label\-free, and scoring is one projection of the last prefix token:
L∗=argminLcos\(u\+\(L\),u−\(L\)\),st=⟨u−\(L∗\),ht\(L∗\)⟩\.L^\{\\ast\}=\\arg\\min\_\{L\}\\cos\\bigl\(u^\{\(L\)\}\_\{\+\},u^\{\(L\)\}\_\{\-\}\\bigr\),\\qquad s\_\{t\}=\\bigl\\langle u^\{\(L^\{\\ast\}\)\}\_\{\-\},\\,h^\{\(L^\{\\ast\}\)\}\_\{t\}\\bigr\\rangle\.\(6\)The fit is also what keeps BSV out of the RL arms: Eq\.[5](https://arxiv.org/html/2609.11061#S3.E5)averages activation deltas of the policy that generated them, so once training movesθ\\thetathe direction is off\-policy, scored against activations the fit never saw, and refreshing it means re\-running the fork/group/average pipeline at training cadence, a cost far beyond the probes’\. We therefore use BSV only where the policy is frozen, as pre\-RL evidence that forking at belief shifts reconstructs the true value curve more accurately than every baseline \(Section[4](https://arxiv.org/html/2609.11061#S4)\); the two fit\-free reads carry the RL arms \(Appendix[E](https://arxiv.org/html/2609.11061#A5)\)\.
### 3\.5Integration into tree\-structured RLVR
We plug the selectors into a one\-fork\-per\-chain instantiation of TreeRL’s on\-policy tree scheme\[[14](https://arxiv.org/html/2609.11061#bib.bib12)\]: per prompt,M=4M\{=\}4chains are sampled and each forks at its top\-scoring boundary intok=2k\{=\}2sibling continuations, giving 12 scored rollouts per prompt \(a chain that cannot fork at all, under two tokens or out of response budget, is replaced by a fresh root chain, so each prompt always emits 12 rollouts\)\. Once a chain has finished generating, the selector scores its boundaries \(forProbe\-JS, with independent one\-token probe requests on the finished chain’s prefix\), and the siblings are ordinary engine requests on that prefix: with automatic prefix caching the prefix’s KV blocks are reused and only the continuation is generated, which is what the rollout term of Eq\.[8](https://arxiv.org/html/2609.11061#S3.E8)charges\. On the training side the 12 rows are processed as separate sequences, prefix included; theTtrainT\_\{\\text\{train\}\}of Appendix[B](https://arxiv.org/html/2609.11061#A2)is measured that way\. A verifier scores each leaf, lightly shaped by a DAPO\-style overlong penalty\[[40](https://arxiv.org/html/2609.11061#bib.bib10)\]\. Advantages follow TreeRL’s global\-plus\-local form: for a node with subtree\-mean returnVV, sibling\-group meanμ\\mu, tree meanV¯\\bar\{V\}, and\|ℒ\|\|\\mathcal\{L\}\|descendant leaves,
A=\[\(V−V¯\)\+\(V−μ\)\]/\|ℒ\|,A\\;=\\;\\bigl\[\(V\-\\bar\{V\}\)\+\(V\-\\mu\)\\bigr\]\\big/\\sqrt\{\|\\mathcal\{L\}\|\},\(7\)broadcast to the segment’s tokens\. The optimizer is the standard RLVR composite \(group\-relative policy gradient, asymmetric clipping\(0\.2,0\.28\)\(0\.2,0\.28\), token\-mean aggregation, dynamic sampling\[[33](https://arxiv.org/html/2609.11061#bib.bib8),[40](https://arxiv.org/html/2609.11061#bib.bib10),[24](https://arxiv.org/html/2609.11061#bib.bib11)\]\), identical across arms\. Baselines:Midpointcuts the remaining span in half;Entropy\(adapted from TreeRL, which ranks single tokens\) forks where the segment\-mean sampled\-token surprisal−logπ\(yt\)\-\\log\\pi\(y\_\{t\}\)peaks over the same candidate pool\.
### 3\.6Cost and a probe\-value bound
Table 2:Probe overhead: FLOP upper bounds atC=16C\{=\}16candidates and1010/2525tokens per engine probe \(L=1L\{=\}1/1616; cells read mathematics / code\), with measured mean chain lengths \(Nemotron’s mathematicsT¯\\bar\{T\}inferred from∼660\{\\sim\}660k rollout tokens per step\)\.#### Cost\.
LetNNbe the parameter count,PPthe prompts per step,T¯\\bar\{T\}the mean chain length in tokens, andMM,kkthe root chains and siblings per cut of Section[3\.5](https://arxiv.org/html/2609.11061#S3.SS5)\. With inference at2N2NFLOPs/token, rollout costsFroll≈2NPM\[1\+k\(1−c\)\]T¯F\_\{\\mathrm\{roll\}\}\\\!\\approx\\\!2N\\,PM\\,\[1\{\+\}k\(1\{\-\}c\)\]\\,\\bar\{T\}\(each chain plus itskksiblings, which continue from a cut at fractionccof the chain,c≈0\.5c\\approx 0\.5\)\. ProbingCCcandidates with suffix lengthEEand read lengthLLon the engine’s prefix cache costsFProbe\-JS=2NPMC\(E\+L\)F\_\{\\textsc\{Probe\-JS\}\}=2N\\,PMC\(E\{\+\}L\), giving
FProbe\-JSFroll=C\(E\+L\)\[1\+k\(1−c\)\]T¯→T¯→∞0\.\\frac\{F\_\{\\textsc\{Probe\-JS\}\}\}\{F\_\{\\mathrm\{roll\}\}\}=\\frac\{C\(E\{\+\}L\)\}\{\[1\{\+\}k\(1\{\-\}c\)\]\\,\\bar\{T\}\}\\;\\xrightarrow\[\\;\\bar\{T\}\\to\\infty\\;\]\{\}\\;0\.\(8\)The probe term carries only the fixed probe lengthE\+L≪T¯E\{\+\}L\\ll\\bar\{T\}, so the probe is a vanishing fraction of rollout, smaller still of the full step, and*shrinks*as chains lengthen\. Equation[8](https://arxiv.org/html/2609.11061#S3.E8)is a per\-chain upper bound \(CCis a cap on the candidates actually scored, and every generated prompt is probed\); Table[2](https://arxiv.org/html/2609.11061#S3.T2)reports the probe’s share of a full training step under the accounting of Appendix[B](https://arxiv.org/html/2609.11061#A2)\. We report FLOPs rather than wall\-clock because the arms are not wall\-clock comparable: the probe path and the serving engine differ across models \(Appendix[B](https://arxiv.org/html/2609.11061#A2)\)\.Lens\-Shiftreadsnnanswer tokens instead ofLLand with prefix reuse costs the same order; ours re\-forwards each prefix on a co\-located copy \(the engine hides intermediate layers\),CρT¯/2C\\rho\\bar\{T\}/2token\-forwards per chain \(Appendix[B](https://arxiv.org/html/2609.11061#A2)\)\.
#### Theoretical analysis: what belief shift can see\.
LetV∗\(t\)V^\{\\ast\}\(t\)be the true value of the prefixy≤ty\_\{\\leq t\}\(the success probability of the policy’s own continuations, the Monte\-Carlo quantity of Section[4](https://arxiv.org/html/2609.11061#S4)\) andVe\(t\)=pt\(a∗\)V\_\{e\}\(t\)=p\_\{t\}\(a^\{\\ast\}\)the*probe value*: the mass the elicited belief of Section[3\.2](https://arxiv.org/html/2609.11061#S3.SS2)puts on the first tokena∗a^\{\\ast\}of the correct answer when the suffixeeforces an answer attt\(forL\>1L\{\>\}1, averaged over theLLread positions\)\. The two are not equal \(VeV\_\{e\}makes the model answer*now*;V∗V^\{\\ast\}lets it keep reasoning\), soVeV\_\{e\}is an approximation ofV∗V^\{\\ast\}whose fidelity we measure rather than assume \(reconstruction of Monte\-CarloV∗V^\{\\ast\}, Section[4](https://arxiv.org/html/2609.11061#S4)\)\. For the probe value the guarantee is exact\. WithTV\(p,q\)=12∑a\|p\(a\)−q\(a\)\|=maxA\|p\(A\)−q\(A\)\|\\mathrm\{TV\}\(p,q\)=\\tfrac\{1\}\{2\}\\sum\_\{a\}\|p\(a\)\-q\(a\)\|=\\max\_\{A\}\|p\(A\)\-q\(A\)\|and Pinsker’s inequality on both halves of Eq\.[1](https://arxiv.org/html/2609.11061#S3.E1),JS≥12TV2\\mathrm\{JS\}\\geq\\tfrac\{1\}\{2\}\\mathrm\{TV\}^\{2\}, hence
\|Ve\(t\)−Ve\(t−\)\|≤TV\(pt−,pt\)≤2JS\(pt−∥pt\)=2st−\\bigl\|V\_\{e\}\(t\)\-V\_\{e\}\(t^\{\-\}\)\\bigr\|\\;\\leq\\;\\mathrm\{TV\}\\\!\\left\(p\_\{t^\{\-\}\},p\_\{t\}\\right\)\\;\\leq\\;\\sqrt\{2\\,\\mathrm\{JS\}\\\!\\left\(p\_\{t^\{\-\}\}\\\|\\,p\_\{t\}\\right\)\}\\;=\\;\\sqrt\{2\\,s\_\{t^\{\-\}\}\}\(9\)\(proof in Appendix[C](https://arxiv.org/html/2609.11061#A3)\), i\.e\. theProbe\-JSscore lower\-bounds the movement of the probe value: a near\-zero score cannot hide a movedVeV\_\{e\}\. How closelyVeV\_\{e\}tracksV∗V^\{\\ast\}is measured, not assumed, by the reconstruction results of Section[4](https://arxiv.org/html/2609.11061#S4)\.
## 4Pre\-RL validation: do the signals find the true value pivots?
End\-to\-end RL is a confounded instrument for judging a branching signal: optimizer noise, seed variance, and the interaction between reward shaping and advantage estimation can all mask the signal’s own effect\. To isolate that effect, we first measure every selector directly against the ground\-truth value curve, before any training\.
#### Setup\.
The probe suites are AIME 2025–2026 and GPQA\-Diamond\[[31](https://arxiv.org/html/2609.11061#bib.bib26)\]\. For each of four probe models \(Gemma\-4\-E4B, OLMo\-3\-7B\-Think, Gemma\-4\-31B, and OLMo\-3\.1\-32B\-Think\[[35](https://arxiv.org/html/2609.11061#bib.bib22),[29](https://arxiv.org/html/2609.11061#bib.bib20)\]; Hugging Face ids in Table[5](https://arxiv.org/html/2609.11061#A1.T5)\) we sample one chain per problem at temperature0\.70\.7, cut it at every newline, and evenly subsample the cuts to at mostm=128m\{=\}128candidate boundariest1,…,tmt\_\{1\},\\dots,t\_\{m\}\(all inference settings: Table[6](https://arxiv.org/html/2609.11061#A1.T6), Appendix[A](https://arxiv.org/html/2609.11061#A1)\)\.
#### Ground truth and metric\.
The value curveV∗\(tj\)=Pr\[correct∣y≤tj\]V^\{\\ast\}\(t\_\{j\}\)=\\Pr\[\\text\{correct\}\\mid y\_\{\\leq t\_\{j\}\}\]is estimated at every boundary byNMC=16N\_\{\\mathrm\{MC\}\}\{=\}16independent Monte\-Carlo completions of the prefix, scored by the verifier\. Each selector then ranks themmboundaries by its own score; we keep its topB=10B\{=\}10, reconstructV^\\hat\{V\}by linear interpolation through the kept points, and charge the selector the reconstruction errorE=∑j\|V∗\(tj\)−V^\(tj\)\|E=\\sum\_\{j\}\\lvert V^\{\\ast\}\(t\_\{j\}\)\-\\hat\{V\}\(t\_\{j\}\)\\rvert\(lower is better\), which is low exactly when the chosen boundaries bracket the true pivots\. Flat\-V∗V^\{\\ast\}problems are excluded; a dynamic program overV∗V^\{\\ast\}gives the per\-panel oracle floor\.
#### Selectors compared\.
Ours: BSV \(vectors fit on held\-out HMMT\[[2](https://arxiv.org/html/2609.11061#bib.bib44)\]; layer by the label\-freecos\\coscriterion of Section[3\.4](https://arxiv.org/html/2609.11061#S3.SS4)\),Probe\-JS, andLens\-Shift\(shape/height\)\. BSV’s direction would need re\-fitting at every policy update, so it is a pre\-RL instrument only; the two fit\-free reads carry the RL arms\. Baselines: a*black\-box LLM judge*\(GPT\-5\.5 rates step importance from the chain alone\), the only baseline with content access, plus entropy, newline, uniform, and random placement\.
Figure 3:Pre\-RL signal quality\(lower is better\): reconstruction error*above the DP oracle floor*on AIME \(top\) and the held\-out GPQA\-Diamond transfer \(bottom\), four probe models\. A belief\-shift selector \(colored\) ranks first in each panel; BSV bars use thecos\\cos\-selected layer \(GPQA: fit out\-of\-domain on math\); exact values in Table[7](https://arxiv.org/html/2609.11061#A5.T7)\.
#### Results\.
Figure[3](https://arxiv.org/html/2609.11061#S4.F3)shows excess error over the oracle floor \(exact values: Table[7](https://arxiv.org/html/2609.11061#A5.T7)\)\.*\(i\)*A belief\-shift signal takes the top rank in all eight model×\\timesbenchmark panels\.*\(ii\)*Entropy hovers at the level of blind newline placement\. The LLM judge, which reads the chain, matches the fit\-free reads on AIME and trails every belief\-shift read on GPQA \(6\.0426\.042vs\.5\.9385\.938–5\.9885\.988\), at the cost of a generation call per step\.*\(iii\)*The signal transfers: on held\-out GPQA\-Diamond, BSV fit on*out\-of\-domain math*beats every baseline on average \(5\.9355\.935vs\.6\.0416\.041\) and in three of four panels, and edges the in\-domain fit \(5\.9995\.999\): the belief direction is not benchmark\-specific\.
BSV with thecos\\cos\-selected layer, fit only to*locate*belief shifts \(outcome labels enter its fitting set only to balance correct and wrong endings within every regime, never as a target; Appendix[A\.1](https://arxiv.org/html/2609.11061#A1.SS1)\), beats every baseline in seven of the eight panels and on both averages, so localizing the shift suffices for a faithfulV^\\hat\{V\}\. Its remaining choice, which layer to project, is also label\-free: the antagonism criterionL∗=argminLcos\(u\+\(L\),u−\(L\)\)L^\{\\ast\}=\\arg\\min\_\{L\}\\cos\(u^\{\(L\)\}\_\{\+\},u^\{\(L\)\}\_\{\-\}\)correlates with reconstruction quality on every model \(Figure[6](https://arxiv.org/html/2609.11061#A5.F6), appendix\), so directions, layer, and score all come from the model’s own rollouts\.
## 5RL Experimental Setup
#### Models\.
Three open substrates spanning architectures:Qwen3\-4B\-Base\[[38](https://arxiv.org/html/2609.11061#bib.bib19)\],OLMo\-3\-7B\(SFT checkpoint\)\[[29](https://arxiv.org/html/2609.11061#bib.bib20)\], andNemotron\-Nano\-9B\-v2\-Base\(hybrid Mamba–Transformer\)\[[3](https://arxiv.org/html/2609.11061#bib.bib21)\]\.
#### Domains and data\.
*Mathematics*: training on DAPO\-Math\-17k\[[40](https://arxiv.org/html/2609.11061#bib.bib10)\]; validation on OlympiadBench \(the 674\-problem English text\-only open\-ended maths split, mean@1\)\[[13](https://arxiv.org/html/2609.11061#bib.bib24)\], AIME 2026 \(30 problems, avg@16\)\[[27](https://arxiv.org/html/2609.11061#bib.bib23)\], and Omni\-MATH\-500 \(a 500\-problem subset we sample from Omni\-MATH, mean@1\)\[[9](https://arxiv.org/html/2609.11061#bib.bib25)\], plus their unweighted mean \(*Agg\.*\)\.*Competitive programming*: training on DeepCoder\-24K\[[26](https://arxiv.org/html/2609.11061#bib.bib28)\]\(21\.6K problems after removing overlap with the LiveCodeBench\-v6 window\); validation on the 175 problems added in LiveCodeBench\-v6 \(43/52/80 easy/medium/hard, avg@8\)\[[15](https://arxiv.org/html/2609.11061#bib.bib27)\]with a binary hidden\-test pass reward\.
#### Protocol\.
Four arms per model and domain,Midpoint,Entropy,Probe\-JS, andLens\-Shift\(shape\), share the tree construction of Section[3\.5](https://arxiv.org/html/2609.11061#S3.SS5), the data order, reward, and optimizer, and differ only in the fork criterion\. All arms are critic\-free \(GRPO\-style\); a learned value model is a separate design axis and is not compared here\. Nor do we re\-test the tree itself: TreeRL\[[14](https://arxiv.org/html/2609.11061#bib.bib12)\]established that its tree scheme beats chain RL under matched rollout budgets, and every arm here shares that scheme, so the comparison isolates a single variable, where the forks are placed\. Validation \(accuracy on the three suites\) runs every 20 steps; each arm is reported at its*single best\-aggregate checkpoint*\(step in gray; Appendix[A](https://arxiv.org/html/2609.11061#A1)reports every aggregate under two other selection rules\)\. Qwen arms are budget\-matched to exactly 680 steps; OLMo and Nemotron arms are capped at a common per\-group budget, within which each arm may stop earlier \(per\-arm step counts in Appendix[A](https://arxiv.org/html/2609.11061#A1)\)\. Training uses verl\[[34](https://arxiv.org/html/2609.11061#bib.bib29)\]with vLLM rollouts\[[18](https://arxiv.org/html/2609.11061#bib.bib30)\]on a single8×8\\timesH200 node \(temperature1\.01\.0,lr=10−6\\mathrm\{lr\}=10^\{\-6\},16×12=19216\\times 12=192rollouts per step after dynamic filtering\)\. Complete hyperparameters in Appendix[A](https://arxiv.org/html/2609.11061#A1)\.
## 6RL Results
Figure 4:Per\-benchmark validation accuracy\(%\), OLMo\-3\-7B \(rows 1–2\) and Qwen3\-4B \(rows 3–4\), mathematics and code\. Solid: 3\-point moving average; faint: raw validations every2020steps;⋆\\star= best\-aggregate checkpoint*within the shown window*\(value at the left\); Table[3](https://arxiv.org/html/2609.11061#S6.T3)follows the full\-budget protocol of Appendix[A](https://arxiv.org/html/2609.11061#A1), so its checkpoint can lie beyond the window\. Windows: OLMo first280280/440440steps, Qwen steps100100–680680/100100–600600\(math / code; warm\-up omitted\)\.Table 3:RL results: accuracy \(%\) at each arm’s*single best\-aggregate checkpoint*\(step in gray\), so every row is read at a single validation step\. Mathematics: OlympiadBench, AIME 2026, Omni\-MATH\-500 and their mean\. Code \(DeepCoder\-24K→\\toLiveCodeBench\-v6\): avg@8 on the easy / medium / hard splits and their mean\. Qwen arms are budget\-matched to 680 steps; Nemotron code arms share a240240\-step cap\.1/2= best / second best per column within each model block \(ties share a medal\)\. Shaded rows: our belief\-shift arms\.### 6\.1Mathematics
Figure[4](https://arxiv.org/html/2609.11061#S6.F4)shows the trajectories behind the OLMo and Qwen rows of Table[3](https://arxiv.org/html/2609.11061#S6.T3)\. A belief\-shift arm takes the aggregate lead on every model:Probe\-JSon OLMo\-3\-7B \(23\.323\.3vs\.20\.720\.7for the strongest baseline\) and Nemotron\-9B \(45\.145\.1vs\.40\.940\.9\),Lens\-Shifton Qwen3\-4B \(29\.629\.6vs\.28\.728\.7\)\. On OLMo\-3\-7B,Probe\-JSlifts AIME 2026 from12\.312\.3\(strongest baseline\) to15\.215\.2; entropy branching*hurts*Nemotron \(40\.040\.0aggregate, below midpoint\), in line with the pre\-RL finding that entropy does not track value \(Section[4](https://arxiv.org/html/2609.11061#S4)\)\. A belief\-shift arm tops 11 of the 12 mathematics columns\.
### 6\.2Code
On competitive programming \(Table[3](https://arxiv.org/html/2609.11061#S6.T3), right\),Probe\-JSwins every OLMo\-3\-7B column \(34\.934\.9vs\.31\.031\.0for the best baseline,\+6\.5\+6\.5on LCB\-medium; Figure[4](https://arxiv.org/html/2609.11061#S6.F4), row 2\)\. Qwen3\-4B mirrors mathematics:Probe\-JSposts the best aggregate \(32\.232\.2vs\.31\.831\.8\) and the best or joint\-best score per split\. Under the common240240\-step cap on Nemotron,Midpointleads andProbe\-JSis second on*every*column \(64\.064\.0vs\.64\.564\.5; hard29\.429\.4vs\.30\.230\.2\)\.
#### Why the gains differ across models\.
A fork criterion pays off through three factors:*ranking*quality,*fork contrast*\(do siblings diverge?\), and the size of the candidate*pool*\. In an offline diagnostic \(setup in Appendix[D](https://arxiv.org/html/2609.11061#A4)\), siblings forked from the RL\-trained Qwen policy reached the*same*final answer63%63\\%of the time \(∼5%\{\\sim\}5\\%for the base model\): the ranking stays healthy \(top\- vs\. bottom\-tercile JS boundaries:55%55\\%vs\.24%24\\%sibling disagreement\), but even a well\-placed fork compares a solution with its near\-clone, so little contrast is left to convert and the gain shrinks to under a point\. OLMo sits at the opposite extreme \(siblings rarely agree; peak belief shifts1\.4×1\.4\\timeslarger\), turning the same ranking into the margins of Table[3](https://arxiv.org/html/2609.11061#S6.T3)\.
## 7Ablations
The four table ablations run on one high\-contrast slice,OLMo\-3\-7B math\(Probe\-JSin panels \(a\)–\(c\) of Table[4](https://arxiv.org/html/2609.11061#S7.T4),Lens\-Shiftin panel \(d\)\), varying one axis per panel; a fifth, model scale, compares Qwen3\-4B with Qwen3\-8B \(Figure[5](https://arxiv.org/html/2609.11061#S7.F5)\)\. Two findings stand out: shrinking the tree to\(2,1\)\(2,1\)costs3\.23\.2aggregate points yet still*edges the full\-budgetMidpointarm with a third of its rollouts*\(\+0\.6\+0\.6\), and readingL=4L\{=\}4belief tokens instead of one nearly*doubles*the edge overMidpoint\(\+→\+7\.3\+3\.8\\\!\\to\\\!\+7\.3; AIME→22\.515\.2\\\!\\to\\\!22\.5\)\. The longer read keeps the exact one\-token belief at position one and adds three positions sampled along the model’s own answer, where multi\-digit answers separate\. This ablation postdates the main arms, so the mathematics rows of Table[3](https://arxiv.org/html/2609.11061#S6.T3)are not back\-filled: every mathematicsProbe\-JSrow keeps the pre\-ablation default \(L=1L\{=\}1\); the codeProbe\-JSarms readL=16L\{=\}16throughout \(Section[3\.2](https://arxiv.org/html/2609.11061#S3.SS2)\)\. TheLens\-Shiftread\-outs land within0\.70\.7aggregate of each other, so the RL arms keep the shape read\. Replacing the composite updater of Section[3\.5](https://arxiv.org/html/2609.11061#S3.SS5)with standard GRPO \(symmetric clipping, sequence\-mean loss\), tree and fork rule held fixed, lowers the aggregate by8\.18\.1points \(→15\.223\.3\\\!\\to\\\!15\.2\), so we keep the composite throughout\.
Table 4:Ablations\(OLMo\-3\-7B math, % at the single best\-aggregate checkpoint, step in gray; all arms follow the OLMo anchor protocol: full\-length candidate pool\)\. Per panel:gray= default\-treeMidpoint\(reference forΔ\\DeltaAgg\);blue/orange= the defaultProbe\-JS/Lens\-Shiftarm of Table[3](https://arxiv.org/html/2609.11061#S6.T3); unshaded = variants \(Probe\-JSin \(a\)–\(c\),Lens\-Shiftin \(d\)\)\. Panel \(c\) raises probe cost from1\.3%1\.3\\%to1\.6%1\.6\\%of a step\.\(a\)Rollout budget: tree\(M,k\)\(M,k\), rollouts/prompt
\(b\)Policy\-update algorithm
\(c\)Probe read lengthLL
\(d\)Lens\-Shiftread\-out
#### Model scale \(Qwen3\-4B→\\toQwen3\-8B\)\.
Figure[5](https://arxiv.org/html/2609.11061#S7.F5)repeats the four Qwen mathematics arms on Qwen3\-8B\-Base, recipe and680680\-step budget unchanged\.Lens\-Shifttakes the aggregate gold at both sizes andProbe\-JSthe silver at 8B \(a shared silver at 4B,28\.728\.7tied with entropy\), and their margin overMidpointgrows with scale:Lens\-Shift\+→\+1\.6\+1\.0\\\!\\to\\\!\+1\.6,Probe\-JS\+→\+1\.1\+0\.1\\\!\\to\\\!\+1\.1\.
Figure 5:Model scale: aggregate mathematics accuracy \(mean of the three suites, %\) for the four arms on Qwen3\-4B \(left\) and Qwen3\-8B \(right\), same recipe and680680\-step budget, steps100100–680680\.
## 8Conclusion
We proposed belief\-shift branching: fork where the model changes its mind\. Before training we scored selectors against Monte\-Carlo value curves on four probe models and two benchmarks, AIME and GPQA\-Diamond, and a belief\-shift read ranks first in all eight panels\. In RL the two fit\-free reads lead every mathematics aggregate, and the probe sweeps the OLMo code columns at about3%3\\%overhead\. Belief\-shift branching thus asks one question, where does the model change its mind, and spends the tree’s forks there; it needs only rollout access and verifier labels, and drops into any tree RLVR pipeline as is\.
### AI use statement
In this work, we used generative AI tools for writing assistance \(editing prose andLaTeXformatting\), for plotting and analysis scripts, and for experiment\-orchestration scripts \(job scheduling and monitoring\)\. The research ideas, method design, and experimental conclusions are the authors’ own; all AI\-assisted code was reviewed and tested by the authors, and all numbers reported in the paper were produced by our training and evaluation pipeline and verified against raw logs\. Two uses go beyond assistance and are part of the method or its evaluation: the surge/steady/drop continuations used to fit the BSV direction \(Appendix[A\.1](https://arxiv.org/html/2609.11061#A1.SS1)\) were drafted by Claude Opus 4\.8 from the probe models’ own rollouts and adversarially checked by a second Claude Opus 4\.8 pass before acceptance, and the black\-box LLM\-judge baseline of the pre\-RL validation \(Section[4](https://arxiv.org/html/2609.11061#S4)\) is GPT\-5\.5\. Neither enters any RL training run\. We take responsibility for the final content of this work, including text, claims, and artifacts produced with the aid of generative AI\.
### Ethics statement
This work studies reinforcement learning for mathematical and coding reasoning on publicly available models and datasets\. It involves no human subjects and no personally identifiable data\. Improved reasoning ability carries the usual dual\-use considerations of stronger language models; our method does not introduce risks beyond those of standard RLVR training\.
### Reproducibility statement
Section[5](https://arxiv.org/html/2609.11061#S5)specifies models, datasets, tree configuration, and optimization hyperparameters; Appendix[A](https://arxiv.org/html/2609.11061#A1)lists the complete configuration, including the boundary\-pool construction and probe settings\. All training uses publicly released base models and public datasets\. Source code will be released upon publication\.
## References
- \[1\]R\. Agarwal, M\. Schwarzer, P\. S\. Castro, A\. Courville, and M\. G\. Bellemare\(2021\)Deep reinforcement learning at the edge of the statistical precipice\.InAdvances in Neural Information Processing Systems \(NeurIPS\),External Links:2108\.13264,[Link](https://arxiv.org/abs/2108.13264)Cited by:[Appendix A](https://arxiv.org/html/2609.11061#A1.SS0.SSS0.Px8.p1.1)\.
- \[2\]M\. Balunovic, J\. Dekoninck, I\. Petrov, N\. Jovanović, and M\. Vechev\(2026\)Matharena: evaluating llms on uncontaminated math competitions\.Advances in Neural Information Processing Systems38\.Cited by:[§4](https://arxiv.org/html/2609.11061#S4.SS0.SSS0.Px3.p1.1)\.
- \[3\]A\. Basant, A\. Khairnar, A\. Paithankar, A\. Khattar, A\. Renduchintala, A\. Malte, A\. Bercovich, A\. Hazare, A\. Rico, A\. Ficek,et al\.\(2025\)Nvidia nemotron nano 2: an accurate and efficient hybrid mamba\-transformer reasoning model\.arXiv preprint arXiv:2508\.14444\.Cited by:[§5](https://arxiv.org/html/2609.11061#S5.SS0.SSS0.Px1.p1.1)\.
- \[4\]G\. Chen, M\. Liao, C\. Li, and K\. Fan\(2024\)Alphamath almost zero: process supervision without process\.Advances in Neural Information Processing Systems37,pp\. 27689–27724\.Cited by:[§2](https://arxiv.org/html/2609.11061#S2.SS0.SSS0.Px2.p1.1)\.
- \[5\]G\. Cui, L\. Yuan, Z\. Wang, H\. Wang, Y\. Zhang, J\. Chen, W\. Li, B\. He, Y\. Fan, T\. Yu,et al\.\(2025\)Process reinforcement through implicit rewards\.arXiv preprint arXiv:2502\.01456\.Cited by:[§2](https://arxiv.org/html/2609.11061#S2.SS0.SSS0.Px2.p1.1)\.
- \[6\]G\. Dong, H\. Mao, K\. Ma, L\. Bao, Y\. Chen, Z\. Wang, Z\. Chen, J\. Du, H\. Wang, F\. Zhang,et al\.\(2026\)Agentic reinforced policy optimization\.InInternational Conference on Learning Representations,Vol\.2026,pp\. 16981–17017\.Cited by:[§1](https://arxiv.org/html/2609.11061#S1.p2.1),[§2](https://arxiv.org/html/2609.11061#S2.SS0.SSS0.Px2.p1.1),[§2](https://arxiv.org/html/2609.11061#S2.SS0.SSS0.Px3.p1.1),[Table 1](https://arxiv.org/html/2609.11061#S2.T1.4.5.1)\.
- \[7\]X\. Du, Y\. Yao, K\. Ma, B\. Wang, T\. Zheng,et al\.\(2025\)SuperGPQA: scaling LLM evaluation across 285 graduate disciplines\.InAdvances in Neural Information Processing Systems \(NeurIPS\), Datasets and Benchmarks Track,External Links:[Link](https://arxiv.org/abs/2502.14739)Cited by:[Table 7](https://arxiv.org/html/2609.11061#A5.T7)\.
- \[8\]X\. Feng, Z\. Wan, M\. Wen, S\. M\. McAleer, Y\. Wen, W\. Zhang, and J\. Wang\(2023\)Alphazero\-like tree\-search can guide large language model decoding and training\.arXiv preprint arXiv:2309\.17179\.Cited by:[§2](https://arxiv.org/html/2609.11061#S2.SS0.SSS0.Px2.p1.1)\.
- \[9\]B\. Gao, F\. Song, Z\. Yang, Z\. Cai, Y\. Miao, Q\. Dong, L\. Li, C\. Ma, L\. Chen, Z\. Tang,et al\.\(2025\)Omni\-math: a universal olympiad level mathematic benchmark for large language models\.InInternational Conference on Learning Representations,Vol\.2025,pp\. 100540–100569\.Cited by:[§5](https://arxiv.org/html/2609.11061#S5.SS0.SSS0.Px2.p1.1)\.
- \[10\]X\. Guan, L\. L\. Zhang, Y\. Liu, N\. Shang, Y\. Sun, Y\. Zhu, F\. Yang, and M\. Yang\(2025\)RStar\-math: small llms can master math reasoning with self\-evolved deep thinking\.arXiv preprint arXiv:2501\.04519\.Cited by:[§2](https://arxiv.org/html/2609.11061#S2.SS0.SSS0.Px2.p1.1)\.
- \[11\]D\. Guo, D\. Yang, H\. Zhang, J\. Song, P\. Wang, Q\. Zhu, R\. Xu, R\. Zhang, S\. Ma, X\. Bi,et al\.\(2025\)DeepSeek\-r1 incentivizes reasoning in llms through reinforcement learning\.Nature645\(8081\),pp\. 633–638\.Cited by:[§1](https://arxiv.org/html/2609.11061#S1.p1.1),[§2](https://arxiv.org/html/2609.11061#S2.SS0.SSS0.Px3.p1.1)\.
- \[12\]Y\. Guo, L\. Xu, J\. Liu, D\. Ye, and S\. Qiu\(2026\)Segment policy optimization: effective segment\-level credit assignment in rl for large language models\.Advances in Neural Information Processing Systems38,pp\. 114399–114431\.Cited by:[§2](https://arxiv.org/html/2609.11061#S2.SS0.SSS0.Px2.p1.1),[§2](https://arxiv.org/html/2609.11061#S2.SS0.SSS0.Px3.p1.1)\.
- \[13\]C\. He, R\. Luo, Y\. Bai, S\. Hu, Z\. Thai, J\. Shen, J\. Hu, X\. Han, Y\. Huang, Y\. Zhang,et al\.\(2024\)Olympiadbench: a challenging benchmark for promoting agi with olympiad\-level bilingual multimodal scientific problems\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 3828–3850\.Cited by:[§5](https://arxiv.org/html/2609.11061#S5.SS0.SSS0.Px2.p1.1)\.
- \[14\]Z\. Hou, Z\. Hu, Y\. Li, R\. Lu, J\. Tang, and Y\. Dong\(2025\)Treerl: llm reinforcement learning with on\-policy tree search\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 12355–12369\.Cited by:[§1](https://arxiv.org/html/2609.11061#S1.p1.1),[§1](https://arxiv.org/html/2609.11061#S1.p2.1),[§2](https://arxiv.org/html/2609.11061#S2.SS0.SSS0.Px2.p1.1),[§2](https://arxiv.org/html/2609.11061#S2.SS0.SSS0.Px3.p1.1),[Table 1](https://arxiv.org/html/2609.11061#S2.T1.4.4.1),[§3\.5](https://arxiv.org/html/2609.11061#S3.SS5.p1.1),[§5](https://arxiv.org/html/2609.11061#S5.SS0.SSS0.Px3.p1.1)\.
- \[15\]N\. Jain, A\. Gu, W\. Li, F\. Yan, T\. Zhang, S\. Wang, A\. Solar\-Lezama, K\. Sen, and I\. Stoica\(2025\)Livecodebench: holistic and contamination free evaluation of large language models for code\.InInternational Conference on Learning Representations,Vol\.2025,pp\. 58791–58831\.Cited by:[§5](https://arxiv.org/html/2609.11061#S5.SS0.SSS0.Px2.p1.1)\.
- \[16\]A\. Kazemnejad, M\. Aghajohari, E\. Portelance, A\. Sordoni, S\. Reddy, A\. Courville, and N\. L\. Roux\(2024\)Vineppo: refining credit assignment in rl training of llms\.arXiv preprint arXiv:2410\.01679\.Cited by:[§2](https://arxiv.org/html/2609.11061#S2.SS0.SSS0.Px2.p1.1),[Table 1](https://arxiv.org/html/2609.11061#S2.T1.4.3.1)\.
- \[17\]L\. Kuhn, Y\. Gal, and S\. Farquhar\(2023\)Semantic uncertainty: linguistic invariances for uncertainty estimation in natural language generation\.arXiv preprint arXiv:2302\.09664\.Cited by:[§1](https://arxiv.org/html/2609.11061#S1.p2.1)\.
- \[18\]W\. Kwon, Z\. Li, S\. Zhuang, Y\. Sheng, L\. Zheng, C\. H\. Yu, J\. Gonzalez, H\. Zhang, and I\. Stoica\(2023\)Efficient memory management for large language model serving with pagedattention\.InProceedings of the 29th symposium on operating systems principles,pp\. 611–626\.Cited by:[§5](https://arxiv.org/html/2609.11061#S5.SS0.SSS0.Px3.p1.1)\.
- \[19\]X\. Lai, Z\. Tian, Y\. Chen, S\. Yang, X\. Peng, and J\. Jia\(2024\)Step\-dpo: step\-wise preference optimization for long\-chain reasoning of llms\.arXiv preprint arXiv:2406\.18629\.Cited by:[§2](https://arxiv.org/html/2609.11061#S2.SS0.SSS0.Px3.p1.1),[Table 1](https://arxiv.org/html/2609.11061#S2.T1.4.6.1)\.
- \[20\]Y\. Li, Q\. Gu, Z\. Wen, Z\. Li, T\. Xing, S\. Guo, T\. Zheng, X\. Zhou, X\. Qu, W\. Zhou,et al\.\(2025\)Treepo: bridging the gap of policy optimization and efficacy and inference efficiency with heuristic tree\-based modeling\.arXiv preprint arXiv:2508\.17445\.Cited by:[§1](https://arxiv.org/html/2609.11061#S1.p1.1),[§1](https://arxiv.org/html/2609.11061#S1.p2.1),[§2](https://arxiv.org/html/2609.11061#S2.SS0.SSS0.Px2.p1.1),[§2](https://arxiv.org/html/2609.11061#S2.SS0.SSS0.Px3.p1.1),[Table 1](https://arxiv.org/html/2609.11061#S2.T1.4.2.1)\.
- \[21\]H\. Lightman, V\. Kosaraju, Y\. Burda, H\. Edwards, B\. Baker, T\. Lee, J\. Leike, J\. Schulman, I\. Sutskever, and K\. Cobbe\(2024\)Let’s verify step by step\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 39578–39601\.Cited by:[§1](https://arxiv.org/html/2609.11061#S1.p2.1),[§2](https://arxiv.org/html/2609.11061#S2.SS0.SSS0.Px3.p1.1),[Table 1](https://arxiv.org/html/2609.11061#S2.T1.4.3.1)\.
- \[22\]J\. Lin\(1991\)Divergence measures based on the shannon entropy\.IEEE Transactions on Information theory37\(1\),pp\. 145–151\.Cited by:[§3\.2](https://arxiv.org/html/2609.11061#S3.SS2.p1.1)\.
- \[23\]Y\. Liu, J\. Lu, Z\. Chen, C\. Qu, J\. K\. Liu, C\. Liu, Z\. Cai, Y\. Xia, L\. Zhao, J\. Bian,et al\.\(2025\)Adaptivestep: automatically dividing reasoning step through model confidence\.arXiv preprint arXiv:2502\.13943\.Cited by:[§2](https://arxiv.org/html/2609.11061#S2.SS0.SSS0.Px3.p1.1)\.
- \[24\]Z\. Liu, C\. Chen, W\. Li, P\. Qi, T\. Pang, C\. Du, W\. S\. Lee, and M\. Lin\(2025\)Understanding r1\-zero\-like training: a critical perspective\.InConference on Language Modeling \(COLM\),Cited by:[§3\.5](https://arxiv.org/html/2609.11061#S3.SS5.p1.2)\.
- \[25\]L\. Luo, Y\. Liu, R\. Liu, S\. Phatale, M\. Guo, H\. Lara, Y\. Li, L\. Shu, Y\. Zhu, L\. Meng,et al\.\(2024\)Improve mathematical reasoning in language models by automated process supervision\.arXiv preprint arXiv:2406\.06592\.Cited by:[§2](https://arxiv.org/html/2609.11061#S2.SS0.SSS0.Px2.p1.1)\.
- \[26\]M\. Luo, S\. Tan, R\. Huang, A\. Patel, A\. Ariyak, Q\. Wu, X\. Shi, R\. Xin, C\. Cai, M\. Weber,et al\.\(2025\)Deepcoder: a fully open\-source 14b coder at o3\-mini level\.Notion Blog1\.Cited by:[§5](https://arxiv.org/html/2609.11061#S5.SS0.SSS0.Px2.p1.1)\.
- \[27\]Mathematical Association of America\(2026\)American invitational mathematics examination \(AIME\)\.Note:[https://maa\.org/maa\-invitational\-competitions/](https://maa.org/maa-invitational-competitions/)Annual invitational high\-school mathematics competitionCited by:[§5](https://arxiv.org/html/2609.11061#S5.SS0.SSS0.Px2.p1.1)\.
- \[28\]nostalgebraist\(2020\)Interpreting GPT: the logit lens\.Note:LessWrong blog postExternal Links:[Link](https://www.lesswrong.com/posts/AcKRB8wDpdaN6v6ru/interpreting-gpt-the-logit-lens)Cited by:[§3\.3](https://arxiv.org/html/2609.11061#S3.SS3.p1.1)\.
- \[29\]T\. Olmo, A\. Ettinger, A\. Bertsch, B\. Kuehl, D\. Graham, D\. Heineman, D\. Groeneveld, F\. Brahman, F\. Timbers, H\. Ivison,et al\.\(2025\)Olmo 3\.arXiv preprint arXiv:2512\.13961\.Cited by:[§4](https://arxiv.org/html/2609.11061#S4.SS0.SSS0.Px1.p1.1),[§5](https://arxiv.org/html/2609.11061#S5.SS0.SSS0.Px1.p1.1)\.
- \[30\]P\. Putta, E\. Mills, N\. Garg, S\. Motwani, C\. Finn, D\. Garg, and R\. Rafailov\(2024\)Agent q: advanced reasoning and learning for autonomous ai agents\.arXiv preprint arXiv:2408\.07199\.Cited by:[§2](https://arxiv.org/html/2609.11061#S2.SS0.SSS0.Px3.p1.1)\.
- \[31\]D\. Rein, B\. L\. Hou, A\. C\. Stickland, J\. Petty, R\. Y\. Pang, J\. Dirani, J\. Michael, and S\. R\. Bowman\(2023\)Gpqa: a graduate\-level google\-proof q&a benchmark\.arXiv preprint arXiv:2311\.12022\.Cited by:[§4](https://arxiv.org/html/2609.11061#S4.SS0.SSS0.Px1.p1.1)\.
- \[32\]A\. Setlur, C\. Nagpal, A\. Fisch, X\. Geng, J\. Eisenstein, R\. Agarwal, A\. Agarwal, J\. Berant, and A\. Kumar\(2025\)Rewarding progress: scaling automated process verifiers for llm reasoning\.InInternational Conference on Learning Representations,Vol\.2025,pp\. 60808–60838\.Cited by:[§2](https://arxiv.org/html/2609.11061#S2.SS0.SSS0.Px1.p1.1)\.
- \[33\]Z\. Shao, P\. Wang, Q\. Zhu, R\. Xu, J\. Song, X\. Bi, H\. Zhang, M\. Zhang, Y\. Li, Y\. Wu,et al\.\(2024\)Deepseekmath: pushing the limits of mathematical reasoning in open language models\.arXiv preprint arXiv:2402\.03300\.Cited by:[§1](https://arxiv.org/html/2609.11061#S1.p1.1),[§2](https://arxiv.org/html/2609.11061#S2.SS0.SSS0.Px1.p1.1),[§3\.5](https://arxiv.org/html/2609.11061#S3.SS5.p1.2)\.
- \[34\]G\. Sheng, C\. Zhang, Z\. Ye, X\. Wu, W\. Zhang, R\. Zhang, Y\. Peng, H\. Lin, and C\. Wu\(2025\)Hybridflow: a flexible and efficient rlhf framework\.InProceedings of the Twentieth European Conference on Computer Systems,pp\. 1279–1297\.Cited by:[§5](https://arxiv.org/html/2609.11061#S5.SS0.SSS0.Px3.p1.1)\.
- \[35\]G\. Team, S\. E\. Abd, V\. Aggarwal, R\. Algayres, A\. Andreev, O\. Bachem, I\. Ballantyne, C\. Brick, V\. Cărbune, M\. Casbon,et al\.\(2026\)Gemma 4 technical report\.arXiv preprint arXiv:2607\.02770\.Cited by:[§4](https://arxiv.org/html/2609.11061#S4.SS0.SSS0.Px1.p1.1)\.
- \[36\]J\. Uesato, N\. Kushman, R\. Kumar, F\. Song, N\. Siegel, L\. Wang, A\. Creswell, G\. Irving, and I\. Higgins\(2022\)Solving math word problems with process\-and outcome\-based feedback\.arXiv preprint arXiv:2211\.14275\.Cited by:[§1](https://arxiv.org/html/2609.11061#S1.p1.1),[§2](https://arxiv.org/html/2609.11061#S2.SS0.SSS0.Px1.p1.1)\.
- \[37\]P\. Wang, L\. Li, Z\. Shao, R\. Xu, D\. Dai, Y\. Li, D\. Chen, Y\. Wu, and Z\. Sui\(2024\)Math\-shepherd: verify and reinforce llms step\-by\-step without human annotations\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 9426–9439\.Cited by:[§2](https://arxiv.org/html/2609.11061#S2.SS0.SSS0.Px1.p1.1),[§2](https://arxiv.org/html/2609.11061#S2.SS0.SSS0.Px2.p1.1)\.
- \[38\]A\. Yang, A\. Li, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Gao, C\. Huang, C\. Lv,et al\.\(2025\)Qwen3 technical report\.arXiv preprint arXiv:2505\.09388\.Cited by:[§5](https://arxiv.org/html/2609.11061#S5.SS0.SSS0.Px1.p1.1)\.
- \[39\]Z\. Yang, Z\. Guo, Y\. Huang, X\. Liang, Y\. Wang, and J\. Tang\(2025\)Treerpo: tree relative policy optimization\.arXiv preprint arXiv:2506\.05183\.Cited by:[§1](https://arxiv.org/html/2609.11061#S1.p1.1),[§1](https://arxiv.org/html/2609.11061#S1.p2.1),[§2](https://arxiv.org/html/2609.11061#S2.SS0.SSS0.Px2.p1.1),[§2](https://arxiv.org/html/2609.11061#S2.SS0.SSS0.Px3.p1.1),[Table 1](https://arxiv.org/html/2609.11061#S2.T1.4.2.1)\.
- \[40\]Q\. Yu, Z\. Zhang, R\. Zhu, Y\. Yuan, X\. Zuo, Y\. Yue, W\. Dai, T\. Fan, G\. Liu, L\. Liu,et al\.\(2026\)Dapo: an open\-source llm reinforcement learning system at scale\.Advances in Neural Information Processing Systems38,pp\. 113222–113244\.Cited by:[§1](https://arxiv.org/html/2609.11061#S1.p1.1),[§2](https://arxiv.org/html/2609.11061#S2.SS0.SSS0.Px1.p1.1),[§3\.5](https://arxiv.org/html/2609.11061#S3.SS5.p1.1),[§3\.5](https://arxiv.org/html/2609.11061#S3.SS5.p1.2),[§5](https://arxiv.org/html/2609.11061#S5.SS0.SSS0.Px2.p1.1)\.
- \[41\]L\. Yuan, W\. Li, H\. Chen, G\. Cui, N\. Ding, K\. Zhang, B\. Zhou, Z\. Liu, and H\. Peng\(2024\)Free process rewards without process labels\.arXiv preprint arXiv:2412\.01981\.Cited by:[§2](https://arxiv.org/html/2609.11061#S2.SS0.SSS0.Px1.p1.1),[§2](https://arxiv.org/html/2609.11061#S2.SS0.SSS0.Px2.p1.1)\.
- \[42\]D\. Zhang, S\. Zhoubian, Z\. Hu, Y\. Yue, Y\. Dong, and J\. Tang\(2024\)Rest\-mcts\*: llm self\-training via process reward guided tree search\.Advances in Neural Information Processing Systems37,pp\. 64735–64772\.Cited by:[§2](https://arxiv.org/html/2609.11061#S2.SS0.SSS0.Px2.p1.1)\.
- \[43\]T\. Zheng, T\. Xing, Q\. Gu, T\. Liang, X\. Qu, X\. Zhou, Y\. Li, Z\. Wen, C\. Lin, W\. Huang,et al\.\(2025\)First return, entropy\-eliciting explore\.arXiv preprint arXiv:2507\.07017\.Cited by:[§1](https://arxiv.org/html/2609.11061#S1.p2.1),[§2](https://arxiv.org/html/2609.11061#S2.SS0.SSS0.Px3.p1.1),[Table 1](https://arxiv.org/html/2609.11061#S2.T1.4.5.1)\.
- \[44\]Y\. Zhou, A\. Zanette, J\. Pan, S\. Levine, and A\. Kumar\(2024\)ArCHer: training language model agents via hierarchical multi\-turn rl\.External Links:2402\.19446Cited by:[§2](https://arxiv.org/html/2609.11061#S2.SS0.SSS0.Px3.p1.1)\.
## Appendix AFull configuration
#### Models\.
Table[5](https://arxiv.org/html/2609.11061#A1.T5)maps every model name used in the paper to its Hugging Face checkpoint\. The RL policies and the pre\-RL probe models are separate checkpoints\.
Table 5:Models used in the paper, by Hugging Face id\.Name in textRoleHugging Face idOLMo\-3\-7BRL policy \(SFT\)allenai/Olmo\-3\-7B\-Instruct\-SFTQwen3\-4BRL policy \(base\)Qwen/Qwen3\-4B\-BaseQwen3\-8BRL policy, scale ablationQwen/Qwen3\-8B\-BaseNemotron\-9BRL policy \(base, hybrid Mamba\)nvidia/NVIDIA\-Nemotron\-Nano\-9B\-v2\-BaseGemma\-4\-E4Bpre\-RL probe modelgoogle/gemma\-4\-E4B\-itOLMo\-3\-7B\-Thinkpre\-RL probe modelallenai/Olmo\-3\-7B\-ThinkGemma\-4\-31Bpre\-RL probe modelgoogle/gemma\-4\-31B\-itOLMo\-3\.1\-32B\-Thinkpre\-RL probe model; Figure[1](https://arxiv.org/html/2609.11061#S1.F1)chainallenai/Olmo\-3\.1\-32B\-Think
#### Tree and rollout\.
TreeRL keep\-original scheme:M=4M\{=\}4root chains per prompt, one criterion\-chosen cut per chain,k=2k\{=\}2fresh sibling continuations per cut; every generation runs to completion; a chain with no boundary in its span is cut at the span midpoint, and a chain that cannot fork at all \(under two tokens, or no response budget left for a continuation\) is refilled with a fresh root chain, so each prompt emits exactly 12 rollouts\. Rollout: vLLM, temperature1\.01\.0, top\-pp1\.01\.0, prefix caching on; max prompt / response length2,0482\{,\}048/24,00024\{,\}000tokens\.
#### Optimization\.
Group batching: 32 prompts generated, 16 kept by DAPO dynamic sampling \(metric: accuracy\),16×12=19216\\times 12=192rollouts per step \(==one mini\-batch\)\. AdamW,lr=10−6\\mathrm\{lr\}=10^\{\-6\}\(10 warmup steps\), weight decay0\.10\.1, gradient clip1\.01\.0\. Loss: asymmetric PPO clipping\(0\.2,0\.28\)\(0\.2,0\.28\), token\-mean aggregation, no entropy bonus, no KL term, no advantage std\-normalization\. Reward: verifier output plus DAPO soft overlong penalty \(buffer4,0964\{,\}096tokens, factor1\.01\.0\)\. Advantage: TreeRL global\+\+local \(Eq\.[7](https://arxiv.org/html/2609.11061#S3.E7)\) with\|ℒ\|\\sqrt\{\|\\mathcal\{L\}\|\}reweighting\. Ulysses sequence parallel 8, gradient checkpointing, FlashAttention\-2;8×8\\timesH200 per run\. Validation every 20 steps, decoding at temperature1\.01\.0, top\-pp0\.950\.95\. Checkpoints were saved every 50 steps on the code arms and the OLMo ablations and every 100 steps on the Qwen math arms; the OLMo and Nemotron math arms ran with checkpointing disabled and are reported from their validation logs\.
#### Signal configuration\.
Boundary pool: positions after a line break; at mostC=16C\{=\}16evenly spaced candidates per chain \(candidate range by substrate: see*Base versus SFT substrates*below\)\.Probe\-JSreads the top\-K=20K\{=\}20next\-token log\-probabilities after a short answer\-eliciting suffix\.Lens\-Shiftuses the shape read\-out in all RL arms \(the height read\-out of the ablation averages the deepest50%50\\%of layers\), full\-attention layers only\. TheProbe\-JSandLens\-Shiftarms reduce the vLLM memory fraction from0\.80\.8to0\.50\.5to co\-locate the scorer\. BSV \(pre\-RL only, Eqs\.[5](https://arxiv.org/html/2609.11061#S3.E5)–[6](https://arxiv.org/html/2609.11061#S3.E6)\): vectors fit on held\-out HMMT problems with LLM\-drafted \(Claude Opus 4\.8\) surge/steady/drop continuations \(Appendix[A\.1](https://arxiv.org/html/2609.11061#A1.SS1)\); windowsWb=Wa=16W\_\{b\}\{=\}W\_\{a\}\{=\}16tokens\.
#### Base versus SFT substrates\.
Two arm\-wide settings differ by substrate, and both address the same failure mode: a base\-model policy \(Qwen3\-4B\-Base, Nemotron\-9B\-Base\) produces short, repetitive, or unstructured chains early in training, whereas the SFT checkpoint \(OLMo\-3\-7B\) writes well\-formed steps from step 0\. \(i\)*Candidate range\.*On base models the belief reads score only the firstρ=0\.5\\rho\{=\}0\.5of each chain; on OLMo they score the full chain \(ρ=1\\rho\{=\}1\)\. Uncapped, a base chain’s elicited belief barely moves across repeated or garbled blocks \(JS≈0\\mathrm\{JS\}\\approx 0\), and the only shift sits on the few tail tokens where an answer happens to appear, pinning the argmax to the end of the chain where a fork can no longer change the outcome; the cap restricts the same ranking to positions a sibling can still overturn\. \(ii\)*Advantage resolution\.*On OLMo the segment advantage of Eq\.[7](https://arxiv.org/html/2609.11061#S3.E7)is subdivided width\-proportionally across sub\-segments; on base models it is broadcast at segment resolution\. With short base chains, subdivision diluted per\-token gradients by roughly the number of sub\-segments and stalled early arms, so those arms keep segment\-resolution credit\.
#### Provenance notes\.
On Nemotron\-9B \(a hybrid\-Mamba architecture\), serving11\-token probe reads through the rollout engine interleaved with long generations destabilized the engine and collapsed theProbe\-JSarm’s response lengths; that arm instead serves probes through a co\-located HF forward\. Checkpoint selection takes each arm’s single best\-validation\-aggregate checkpoint, applied identically to all arms\.
#### Per\-arm budgets\.
Each model×\\timesdomain group is trained until aggregate accuracy stops improving appreciably; all arms in a group share that step cap and are reported at their best\-aggregate checkpoint within it\. OLMo math \(cap 460\):Midpoint460,Entropy400,Lens\-Shift420,Probe\-JS340\. OLMo code \(cap 520\):Midpoint480,Entropy540 \(truncated to the cap\),Lens\-Shift440,Probe\-JS600 \(truncated to the cap\)\. Nemotron math \(cap 420\) and code \(cap 240\) are within 60 steps across arms\. On mathematics the belief\-shift arms run the fewest steps; on OLMo codeProbe\-JSruns longest, and at the tighter440440\-step cap shared by all four arms it still leads \(33\.133\.1vs\.31\.531\.5for the next\-best arm\)\.
#### Robustness to checkpoint selection\.
Selection protocols differ across the RLVR literature \(most papers leave it unspecified\), and a maximum taken over training is a known source of optimistic bias\[[1](https://arxiv.org/html/2609.11061#bib.bib43)\]\. We therefore recomputed every arm under three rules: \(i\) the reported single best\-aggregate checkpoint; \(ii\) each benchmark’s own best checkpoint \(three different models per row\); and \(iii\) the checkpoint at a fixed common budget for all arms in a group\. A belief\-shift arm holds the mathematics aggregate lead on all three models under every rule, andProbe\-JSsweeps the OLMo code columns under \(i\) and \(ii\)\. Relative to the reported rule \(i\), rule \(ii\) adds\+0\.52\+0\.52aggregate points on average over the 24 arms\. Relative to the fixed\-budget rule \(iii\), it adds\+1\.98\+1\.98for the baselines and\+1\.71\+1\.71for our arms, so peak selection favors the baselines and does not manufacture the effect\. Several sub\-point margins change order across rules while the headline leads do not: the OLMo mathematics gold under \(iii\) \(Lens\-Shift21\.221\.2vs\.Probe\-JS21\.221\.2\), the Qwen mathematics silver, the Qwen code aggregate \(four arms within0\.80\.8; entropy leads under \(iii\)\), and the Nemotron code silver and bronze\.
#### Pre\-RL validation\.
Table[6](https://arxiv.org/html/2609.11061#A1.T6)lists the inference and scoring settings of the pre\-RL protocol of Section[4](https://arxiv.org/html/2609.11061#S4); chains and Monte\-Carlo completions are generated with the same vLLM build as the RL rollouts\.
Table 6:Pre\-RL validation settings \(Section[4](https://arxiv.org/html/2609.11061#S4)\)\.
### A\.1How the surge / steady / drop continuations are built
Each fitting sample is a shared prefixPPplus LLM\-drafted continuationscgc^\{g\}, one per regimeg∈\{\+,0,−\}g\\in\\\{\+,0,\-\\\}\(two for drop\)\. The prefix is never written by hand: it is one of the model’s own rollouts on a held\-out fitting problem \(Table[6](https://arxiv.org/html/2609.11061#A1.T6)\), cut at a newline where the model has just committed to its current track, either a correct intermediate step \(initial valuevi=Hv\_\{i\}\{=\}H\) or the key wrong step behind a confident wrong answer \(vi=Lv\_\{i\}\{=\}L\)\. Forvi=Lv\_\{i\}\{=\}Lthe cut precedes any second\-guessing by the model itself, so the only belief move in the window is the injected one\. Continuations sharePP, match the voice of the chain, run one to four sentences, and differ only in the designed belief move:
- •Surge\(g=\+g\{=\}\{\+\}\): the model grows more confident and commits to the answer already on its track \(“Yes, this is clearly …; I am now certain the answer is …”\); the value direction is unchanged,vf=viv\_\{f\}\{=\}v\_\{i\}\.
- •Steady\(g=0g\{=\}0\): a matter\-of\-fact continuation that advances the derivation with no change of confidence \(“Continuing methodically, the next step is …”\);vf=viv\_\{f\}\{=\}v\_\{i\}\. This is the baseline thatΔ0\\Delta\_\{0\}in Eq\.[5](https://arxiv.org/html/2609.11061#S3.E5)subtracts\.
- •Drop\(g=−g\{=\}\{\-\}\): the model voices doubt and re\-examines the step; the value may or may not flip\. From a wrong track \(vi=Lv\_\{i\}\{=\}L\) we write a*catch*that diagnoses the actual error and reaches the gold answer \(vf=Hv\_\{f\}\{=\}H\) and a*suppressed doubt*that raises the concern and then talks itself back onto the wrong answer \(vf=Lv\_\{f\}\{=\}L\)\. From a correct track \(vi=Hv\_\{i\}\{=\}H\) we write a*doubt\-then\-confirm*that re\-checks and keeps the step \(vf=Hv\_\{f\}\{=\}H\) and a*derail*that introduces a plausible real mistake, mined where possible from a different failing rollout on the same problem \(vf=Lv\_\{f\}\{=\}L\)\.
Every prefix thus yields one surge, one steady, and two drop continuations, filling the eight admissible cells of theg×vi×vfg\\times v\_\{i\}\\times v\_\{f\}design \(surge and steady cells withvf≠viv\_\{f\}\\neq v\_\{i\}are excluded by construction\)\. The belief shift must live in the content, the stated conclusion or confidence has to change, not in a surface marker such as “wait” or “hmm”; catch and derail continuations must be mathematically valid\. Each continuation also carries two verbatim quotes marking where the belief move begins and where the new conclusion is stated \(for steady, a representative stretch\); they serve the checker below, while the activations themselves are pooled over the fixedWb=Wa=16W\_\{b\}\{=\}W\_\{a\}\{=\}16\-token windows of Eq\.[5](https://arxiv.org/html/2609.11061#S3.E5)\. Continuations were drafted by Claude Opus 4\.8 from a per\-item work file \(the real rollout, the gold answer, and, forvi=Lv\_\{i\}\{=\}L, the wrong answer\), then passed through an adversarial checker, a second Claude Opus 4\.8 pass, that verified cut\-marker uniqueness, exact quote matching, cell consistency, and the correctness of every catch and derail before a sample was accepted\.
## Appendix BCompute\-cost derivation
One mathematics probe appends the elicitation suffix \(99tokens for Nemotron’s tokenizer,88for OLMo’s and Qwen’s; we charge1010throughout\) and reads one position \(L=1L\{=\}1\):1010token\-forwards, riding the engine’s prefix cache\. Every generated prompt is probed and forked: trees are built for allPgen=32P\_\{\\text\{gen\}\}\{=\}32prompts, and dynamic sampling then keepsPtrain=16P\_\{\\text\{train\}\}\{=\}16of the groups for the update, so probe and rollout tokens scale withPgenP\_\{\\text\{gen\}\}and trained tokens withPtrainP\_\{\\text\{train\}\}\. Per step, withMMchains andCCcandidates per chain:
Tprobe\\displaystyle T\_\{\\text\{probe\}\}=Pgen⋅M⋅C⋅\(Lprobe\+1\)=32⋅4⋅16⋅10=20,480token\-forwards,\\displaystyle=P\_\{\\text\{gen\}\}\\cdot M\\cdot C\\cdot\(L\_\{\\text\{probe\}\}\{\+\}1\)=32\\cdot 4\\cdot 16\\cdot 10=20\{,\}480\\ \\text\{token\-forwards\},Troll\\displaystyle T\_\{\\text\{roll\}\}≈Pgen\[MT¯\+MkT¯\(1−c\)\]=8PgenT¯\(cut fractionc≈0\.5\),\\displaystyle\\approx P\_\{\\text\{gen\}\}\\bigl\[M\\bar\{T\}\+M\\,k\\,\\bar\{T\}\(1\-c\)\\bigr\]=8\\,P\_\{\\text\{gen\}\}\\bar\{T\}\\quad\(\\text\{cut fraction \}c\\approx 0\.5\),giving339,712339\{,\}712\(OLMo,T¯=1,327\\bar\{T\}\{=\}1\{,\}327\),1,065,9841\{,\}065\{,\}984\(Qwen,T¯=4,164\\bar\{T\}\{=\}4\{,\}164\), and the measured∼660\{\\sim\}660k \(Nemotron,T¯≈2,600\\bar\{T\}\\approx 2\{,\}600inferred\) rollout tokens\. With forward=2N=2Nand forward\+\+backward=6N=6NFLOPs/token \(NNcancels\),Fprobe/Froll=6\.0%F\_\{\\text\{probe\}\}/F\_\{\\text\{roll\}\}=6\.0\\%\(OLMo\),1\.9%1\.9\\%\(Qwen\),3\.1%3\.1\\%\(Nemotron\), andFprobe/Fstep=2NTprobe/\(2NTroll\+6NTtrain\+2NTprobe\)=1\.3%F\_\{\\text\{probe\}\}/F\_\{\\text\{step\}\}=2NT\_\{\\text\{probe\}\}/\(2NT\_\{\\text\{roll\}\}\+6NT\_\{\\text\{train\}\}\+2NT\_\{\\text\{probe\}\}\)=1\.3\\%/0\.45%0\.45\\%/0\.70%0\.70\\%, withTtrainT\_\{\\text\{train\}\}the median number of tokens actually trained per step \(prompt and response over the 192 kept rollouts, medians over each model’s fullMidpointrun, as isT¯\\bar\{T\}:389389k /1,1581\{,\}158k /752752k\)\. The code arms readL=16L\{=\}16tokens after a99\-token suffix,2525token\-forwards per probe, soTprobe=32⋅4⋅16⋅25=51,200T\_\{\\text\{probe\}\}=32\\cdot 4\\cdot 16\\cdot 25=51\{,\}200\. With the code arms’ measuredT¯\\bar\{T\}/TtrainT\_\{\\text\{train\}\}\(1,1791\{,\}179/415415k OLMo,760760/312312k Qwen,6,5576\{,\}557/1,9771\{,\}977k Nemotron; medians over eachMidpointcode run, Qwen over its final logged window, steps650650–680680\),Fprobe/Froll=17%F\_\{\\text\{probe\}\}/F\_\{\\text\{roll\}\}=17\\%/26%26\\%/3\.1%3\.1\\%andFprobe/Fstep=3\.2%F\_\{\\text\{probe\}\}/F\_\{\\text\{step\}\}=3\.2\\%/4\.3%4\.3\\%/0\.67%0\.67\\%\(second entry of each cell in Table[2](https://arxiv.org/html/2609.11061#S3.T2)\)\.C=16C\{=\}16is a cap, so all of these are upper bounds\.
We do not report wall\-clock: the arms are not comparable on that axis\. On OLMo and Qwen the probe runs inside the vLLM engine on its prefix cache, whereas on Nemotron it runs on a co\-located Hugging Face copy: mixing one\-token probe requests into that hybrid\-Mamba model’s serving batches perturbed its rollouts \(a degenerate length collapse within about6060steps that an otherwise identical probe\-free run did not show\), so the belief read was taken off\-engine\.
#### Lens\-Shiftcost\.
The serving engine exposes only final\-layer log\-probabilities, so the depth profile is read on a co\-located Hugging Face copy of the policy, and that copy re\-forwards each candidate’s full prefix rather than reusing a cache, an engineering choice rather than a property of the signal\. With prefix reuse the read would costC\(E\+n\)C\(E\{\+\}n\)token\-forwards per chain, the same order as the probe\. Without it the cost isCρT¯/2C\\rho\\bar\{T\}/2prefix tokens per probed chain\.
## Appendix CProof of the probe\-value bound
#### Proof of Eq\.[9](https://arxiv.org/html/2609.11061#S3.E9)\.
Throughout,ppandqqare probability distributions on a finite answer set𝒜\\mathcal\{A\}; in the scored instance they are the floored, renormalized top\-KKbeliefsp~t−,p~t\\tilde\{p\}\_\{t^\{\-\}\},\\tilde\{p\}\_\{t\}of Eq\.[2](https://arxiv.org/html/2609.11061#S3.E2), so the bound holds exactly for the quantities the score is computed from\.
*Step 1: the two forms of total variation\.*LetS=\{a∈𝒜:p\(a\)≥q\(a\)\}S=\\\{a\\in\\mathcal\{A\}:\\,p\(a\)\\geq q\(a\)\\\}\. Because∑a\(p\(a\)−q\(a\)\)=0\\sum\_\{a\}\\bigl\(p\(a\)\-q\(a\)\\bigr\)=0,
∑a∈S\(p\(a\)−q\(a\)\)=∑a∉S\(q\(a\)−p\(a\)\),so12∑a\|p\(a\)−q\(a\)\|=p\(S\)−q\(S\)\.\\sum\_\{a\\in S\}\\bigl\(p\(a\)\-q\(a\)\\bigr\)=\\sum\_\{a\\notin S\}\\bigl\(q\(a\)\-p\(a\)\\bigr\),\\qquad\\text\{so\}\\qquad\\tfrac\{1\}\{2\}\\sum\_\{a\}\\bigl\|p\(a\)\-q\(a\)\\bigr\|=p\(S\)\-q\(S\)\.For an arbitrary eventA⊆𝒜A\\subseteq\\mathcal\{A\}, split it alongSS:
p\(A\)−q\(A\)\\displaystyle p\(A\)\-q\(A\)=∑a∈A∩S\(p\(a\)−q\(a\)\)−∑a∈A∖S\(q\(a\)−p\(a\)\)\\displaystyle=\\sum\_\{a\\in A\\cap S\}\\bigl\(p\(a\)\-q\(a\)\\bigr\)\-\\sum\_\{a\\in A\\setminus S\}\\bigl\(q\(a\)\-p\(a\)\\bigr\)≤∑a∈A∩S\(p\(a\)−q\(a\)\)≤∑a∈S\(p\(a\)−q\(a\)\)=p\(S\)−q\(S\),\\displaystyle\\leq\\sum\_\{a\\in A\\cap S\}\\bigl\(p\(a\)\-q\(a\)\\bigr\)\\;\\leq\\;\\sum\_\{a\\in S\}\\bigl\(p\(a\)\-q\(a\)\\bigr\)=p\(S\)\-q\(S\),dropping the non\-negative second sum and then enlargingA∩SA\\cap StoSS\(every summand overSSis non\-negative\)\. Applying the same to the complement,q\(A\)−p\(A\)=p\(Ac\)−q\(Ac\)≤p\(S\)−q\(S\)q\(A\)\-p\(A\)=p\(A^\{c\}\)\-q\(A^\{c\}\)\\leq p\(S\)\-q\(S\)\. Hence\|p\(A\)−q\(A\)\|≤p\(S\)−q\(S\)\|p\(A\)\-q\(A\)\|\\leq p\(S\)\-q\(S\)for everyAA, with equality atA=SA=S:
TV\(p,q\):=12∑a\|p\(a\)−q\(a\)\|=maxA⊆𝒜\|p\(A\)−q\(A\)\|\.\\mathrm\{TV\}\(p,q\)\\;:=\\;\\tfrac\{1\}\{2\}\\sum\_\{a\}\\bigl\|p\(a\)\-q\(a\)\\bigr\|\\;=\\;\\max\_\{A\\subseteq\\mathcal\{A\}\}\\bigl\|p\(A\)\-q\(A\)\\bigr\|\.The singletonA=\{a∗\}A=\\\{a^\{\\ast\}\\\}is a special case,\|p\(a∗\)−q\(a∗\)\|≤TV\(p,q\)\|p\(a^\{\\ast\}\)\-q\(a^\{\\ast\}\)\|\\leq\\mathrm\{TV\}\(p,q\); withp=pt−p=p\_\{t^\{\-\}\},q=ptq=p\_\{t\}andVe\(t\)=pt\(a∗\)V\_\{e\}\(t\)=p\_\{t\}\(a^\{\\ast\}\)this is the first inequality of Eq\.[9](https://arxiv.org/html/2609.11061#S3.E9)\.
*Step 2: distance to the midpoint\.*Form=12\(p\+q\)m=\\tfrac\{1\}\{2\}\(p\+q\), pointwise\|p\(a\)−m\(a\)\|=12\|p\(a\)−q\(a\)\|=\|q\(a\)−m\(a\)\|\|p\(a\)\-m\(a\)\|=\\tfrac\{1\}\{2\}\|p\(a\)\-q\(a\)\|=\|q\(a\)\-m\(a\)\|, soTV\(p,m\)=TV\(q,m\)=12TV\(p,q\)\\mathrm\{TV\}\(p,m\)=\\mathrm\{TV\}\(q,m\)=\\tfrac\{1\}\{2\}\\mathrm\{TV\}\(p,q\)\.
*Step 3: Pinsker on both halves\.*Pinsker’s inequality,DKL\(p∥m\)≥2TV\(p,m\)2D\_\{\\mathrm\{KL\}\}\(p\\\|m\)\\geq 2\\,\\mathrm\{TV\}\(p,m\)^\{2\}\(natural logarithm\), applied to each half of Eq\.[1](https://arxiv.org/html/2609.11061#S3.E1)with Step 2:
JS\(p∥q\)=12DKL\(p∥m\)\+12DKL\(q∥m\)≥12⋅2\(TV\(p,q\)2\)2\+12⋅2\(TV\(p,q\)2\)2=12TV\(p,q\)2,\\mathrm\{JS\}\(p\\\|q\)=\\tfrac\{1\}\{2\}D\_\{\\mathrm\{KL\}\}\(p\\\|m\)\+\\tfrac\{1\}\{2\}D\_\{\\mathrm\{KL\}\}\(q\\\|m\)\\;\\geq\\;\\tfrac\{1\}\{2\}\\cdot 2\\Bigl\(\\tfrac\{\\mathrm\{TV\}\(p,q\)\}\{2\}\\Bigr\)^\{2\}\+\\tfrac\{1\}\{2\}\\cdot 2\\Bigl\(\\tfrac\{\\mathrm\{TV\}\(p,q\)\}\{2\}\\Bigr\)^\{2\}=\\tfrac\{1\}\{2\}\\,\\mathrm\{TV\}\(p,q\)^\{2\},i\.e\.TV\(p,q\)≤2JS\(p∥q\)\\mathrm\{TV\}\(p,q\)\\leq\\sqrt\{2\\,\\mathrm\{JS\}\(p\\\|q\)\}, the second inequality\. Chaining Steps 1–3 gives Eq\.[9](https://arxiv.org/html/2609.11061#S3.E9)\.
*What the bound does and does not cover\.*The inequality is exact for the probe valueVeV\_\{e\}built from the scored beliefs\. Two approximations sit outside it and are assessed empirically: \(i\) top\-KKtruncation, which changes how faithfullyp~t\(a∗\)\\tilde\{p\}\_\{t\}\(a^\{\\ast\}\)tracks the untruncated elicited mass \(vanishing asKKgrows\); and \(ii\) elicitation itself, which makesVeV\_\{e\}a proxy for the true valueV∗V^\{\\ast\}\(forcing an answer atttversus letting the policy keep reasoning\)\. The pre\-RL reconstruction of Monte\-CarloV∗V^\{\\ast\}in Section[4](https://arxiv.org/html/2609.11061#S4)is the measurement of that gap\.
## Appendix DFork\-contrast diagnostic
The diagnostic quoted in Section[6](https://arxiv.org/html/2609.11061#S6)\(“Why the gains differ across models”\) is an offline measurement with HF transformers on one GPU, outside the RL loop, over three policies: Qwen3\-4B\-Base, the step\-600600checkpoint of an earlier QwenProbe\-JSarm trained with the same tree, data, and optimizer, and OLMo\-3\-7B\-Instruct\-SFT\. The three policies see the problems in the same order\.
#### Sibling agreement and JS terciles\.
For each policy we take the first six OlympiadBench problems whose chain yields a parsable answer and generate one chain per problem at temperature1\.01\.0\(at most15361536new tokens\)\. Newline boundaries are evenly subsampled to at most eight; at each boundary the belief is read with the boxed suffix of Section[3\.2](https://arxiv.org/html/2609.11061#S3.SS2)\(top\-2020log\-probabilities\) and theProbe\-JSscore is computed with the training code path\. Eight siblings are then forked at*every*boundary at temperature1\.01\.0and their final answers are compared with the root chain’s\. “Same answer” is the fraction of siblings that reproduce the root answer:62\.7%62\.7\\%for trained Qwen \(102102siblings at3434boundaries\),4\.8%4\.8\\%for the base model \(6363at2121\), and0%0\\%for OLMo \(3939at1313\)\. Sorting the trained\-Qwen boundaries byProbe\-JSscore, the top tercile gives55%55\\%sibling disagreement and the bottom tercile24%24\\%\.
#### Peak belief shift\.
For each policy,3030chains \(AIME 2026 problems\) at temperature0\.80\.8,1616evenly spaced newline boundaries per chain, and the per\-chain maximumProbe\-JSscore averaged over chains: OLMo0\.4100\.410, trained Qwen0\.2890\.289, base Qwen0\.2640\.264, hence the1\.4×1\.4\\timesof the main text\. The agreement measurement is criterion\-agnostic \(siblings are forked at every boundary\), so the contrast collapse it documents applies equally toMidpointandEntropyforks; the tercile split is the only part specific toProbe\-JS\.
## Appendix EAdditional results
Table 7:Pre\-RL value\-reconstruction errorEE\(lower is better\): the exact values behind Figure[3](https://arxiv.org/html/2609.11061#S4.F3), mean over non\-flat problems\.1/2= best / second best per column; belief\-shift signals above the rule in each block\.*AIME 2025–2026*\(nn: Gemma\-4\-E4B 51, OLMo\-3\-7B\-Think 34, Gemma\-4\-31B 29, OLMo\-3\.1\-32B\-Think 35\): BSV vectors fit on held\-out HMMT problems\.*GPQA\-Diamond transfer*\(nn: 180/181/83/152\): BSV vectors fit either in\-domain \(SuperGPQA\[[7](https://arxiv.org/html/2609.11061#bib.bib36)\]\) or fully out\-of\-domain \(math\); thecos\\cos\-selected OOD variant beats every baseline on average and edges the in\-domain fit, i\.e\., the belief direction is not benchmark\-specific\.Figure 6:Label\-free layer selection for BSV\.Per\-layercos\(u\+\(L\),u−\(L\)\)\\cos\(u\_\{\+\}^\{\(L\)\},u\_\{\-\}^\{\(L\)\}\)\(dashed, right axis\) against reconstruction error \(solid, left axis\) across the four probe models: correlation\+0\.26\+0\.26to\+0\.54\+0\.54\(“corr bsv” in each panel title\), soL∗=argminLcos\(u\+,u−\)L^\{\\ast\}=\\arg\\min\_\{L\}\\cos\(u\_\{\+\},u\_\{\-\}\), computable without any labels, selects a near\-optimal evaluation layer \(exact\-best on Gemma\-4\-31B\)\. The orange curves and the “val” correlation are a value\-change variant of the direction, fit on the rollout\-value flip instead of the belief move; it is shown for reference only and is not used in the paper\. Panel titles use short names: gemma\-4B = Gemma\-4\-E4B, Olmo\-7B = OLMo\-3\-7B\-Think, gemma\-31B = Gemma\-4\-31B, Olmo\-32B = OLMo\-3\.1\-32B\-Think\.The per\-benchmark validation trajectories behind the OLMo and Qwen rows of Table[3](https://arxiv.org/html/2609.11061#S6.T3)are in Figure[4](https://arxiv.org/html/2609.11061#S6.F4)\.
On BSV in RL: BSV leads the pre\-RL rankings \(Table[7](https://arxiv.org/html/2609.11061#A5.T7)\) but requires a per\-model offline vector fit, and its directions are not guaranteed to transfer across checkpoints of the same family; our RL arms therefore train with the two fit\-free read\-outs \(Section[3\.4](https://arxiv.org/html/2609.11061#S3.SS4)\)\.
## Appendix FThe Figure[1](https://arxiv.org/html/2609.11061#S1.F1)case in full
This appendix reproduces the complete case behind Figure[1](https://arxiv.org/html/2609.11061#S1.F1): the verbatim prompt, the twenty fork picks with their scores and ground\-truth value movement, and the full reasoning chain\. The chain is an OLMo\-3\.1\-32B\-Think sample \(67,117 characters, 128 newline boundaries\) that ends in the correct answer588\\boxed\{588\};V∗V^\{\\ast\}at each boundary is estimated with 16 Monte\-Carlo completions \(Section[4](https://arxiv.org/html/2609.11061#S4)\)\. In the transcript,▼\\blacktriangledownPkmarks thekk\-thProbe\-JSfork and▼\\blacktriangledownEkthekk\-th entropy fork, both numbered in chain order; each marker sits at the exact boundary where that fork’s siblings would branch\. All tenProbe\-JSforks fall in the first26%26\\%of the chain, where the shoelace computation is repeatedly set up, botched, and repaired andV∗V^\{\\ast\}swings between0\.620\.62and1\.001\.00; entropy places six of its ten \(E5–E10\) in the settled tail \(V∗≡1V^\{\\ast\}\\\!\\equiv\\\!1,ΔV∗=0\\Delta V^\{\\ast\}\\\!=\\\!0\), where high token uncertainty reflects interchangeable wording, not an undecided outcome\.
### F\.1Prompt
The problem is from AIME 2025 \(area of heptagonAFNBCEMAFNBCEM; answer 588\)\. The model received the statement below together with the Ti*k*Z source of the diagram, which is rendered here from that source\.
> On△ABC\\triangle ABCpointsA,D,EA,D,E, andBBlie in that order on sideAB¯\\overline\{AB\}withAD=4AD=4,DE=16DE=16,EB=8EB=8\. PointsA,F,GA,F,GandCClie in that order on sideAC¯\\overline\{AC\}withAF=13AF=13,FG=52FG=52, andGC=26GC=26\. LetMMbe the reflection ofDDthroughFF, and letNNbe the reflection ofGGthroughEE\. QuadrilateralDEGFDEGFhas area288288\. Find the area of heptagonAFNBCEMAFNBCEM\.
### F\.2The twenty fork picks
Table[8](https://arxiv.org/html/2609.11061#A6.T8)lists both selectors’ picks: character offset and fraction of the chain, the selector’s own score at the chosen boundary, and the ground\-truth value on the two sides of the boundary\.
Table 8:Fork picks for the Figure[1](https://arxiv.org/html/2609.11061#S1.F1)chain, in chain order\. “score” is the selector’s own criterion \(Probe\-JS: JS divergence of consecutive answer beliefs; entropy: boundary entropy\);V∗V^\{\\ast\}is the 16\-sample Monte\-Carlo value before and after the chosen boundary\.
### F\.3Reasoning chain
The complete chain, verbatim \(LaTeXin the model’s output is shown as source\)\. Fork markers are typeset in place; the surrounding text is otherwise unmodified\.
Okay, so I need to find the area of this heptagon AFNBCEM given some information about the triangle ABC and the points on its sides\. Let me start by trying to understand the problem and visualize it based on the description and the TikZ diagram\.
First, the problem says that on side AB, the points are in the order A, D, E, B with lengths AD=4, DE=16, and EB=8\. So the entire length of AB must be AD \+ DE \+ EB = 4 \+ 16 \+ 8 = 28 units\. Similarly, on side AC, the points are A, F, G, C with AF=13, FG=52, and GC=26\. Therefore, the length of AC is AF \+ FG \+ GC = 13 \+ 52 \+ 26 = 91 units\.
Now, they mention points M and N: M is the reflection of D through F, and N is the reflection of G through E\. Let me recall that reflecting a point through another point means that the second point is the midpoint of the original point and its reflection\. So, for example, reflecting D over F means that F is the midpoint of D and M\. So, the coordinates \(if we assign coordinates\) of M can be determined if we know D and F\. Similarly, N is the reflection of G over E, so E is the midpoint of G and N\.
The area of quadrilateral DEGF is given as 288\. The heptagon AFNBCEM is the area we need to find\. Let me try to break down the figure\.
First, perhaps assigning coordinates would help here\. Since we have points on sides AB and AC, maybe it’s a good idea to set coordinates with point A at the origin, but actually, looking at the TikZ code, they have coordinates for A at \(100,100\), but maybe we can set up a coordinate system ourselves for simplicity\.
Alternatively, maybe we can use coordinate geometry here\. Let’s try to set coordinate system with point A at \(0,0\)\. Let me see\.
Let me denote coordinates:
Let’s set point A at \(0,0\)\. Let’s let side AB lie along the x\-axis, so point B would be at \(28,0\), since AB is 28 units\. Then, points D, E are on AB\. Since AD=4, D is at \(4,0\)\. Then DE=16, so E is at \(4 \+16\)=20 units from A, so E is at \(20,0\)\. Then EB is 8, so from 20 to 28 is 8, which checks out\.
Now, side AC is another side\. Let’s let AC be along another axis, but since in the problem, points on AC are A, F, G, C with AF=13, FG=52, GC=26\. So total length AC is 13 \+52 \+26 =91 as I said\. Let me set point A at \(0,0\), and let’s set AC along the y\-axis for simplicity\. Wait, but then AB and AC would be perpendicular? The problem doesn’t specify that the triangle is right\-angled, so maybe that’s an incorrect assumption\. Hmm, that complicates things\.
Alternatively, perhaps we can assign coordinates more flexibly\. Let me suppose coordinates with A at \(0,0\), AB along the x\-axis, so B is at \(28,0\)\. Then point C is somewhere in the plane, say at \(p, q\)\. Then AC is from \(0,0\) to \(p, q\), and points F and G divide AC into segments of 13, 52, 26\. So we can parametrize the coordinates of F and G accordingly\.
Let me try this approach\. Let me formalize this:
Let’s set coordinate system with A at \(0,0\), and AB along x\-axis, so coordinates:
\- A: \(0,0\) \- B: \(28, 0\)
Point C is \(c\_x, c\_y\), which we don’t know yet\. Then AC is from \(0,0\) to \(c\_x, c\_y\)\. The points F and G divide AC into three parts: AF=13, FG=52, GC=26\. So the total length AC is 91, but the actual coordinates depend on the direction of AC\. Since the distances along the line AC are given, we can parametrize the points\.
The coordinates of a point dividing a segment in a ratio can be found using section formula\. For example, point F is 13 units from A, and since the total length AC is 91, the ratio AF : AC is 13:91 = 1:7\. So F divides AC in the ratio AF:FC = 13 : \(52 \+26\) =13:78 =1:6\. Wait, actually AF is 13, FG is 52, so from A to F is 13, F to G is 52, so from A to G is 13\+52=65, and then G to C is 26, so AC is 91\. Therefore, the ratio AF : AC is 13:91 = 1:7, so F is 1/7 of the way from A to C\. Similarly, G is \(13\+52\)/91 =65/91 =5/7 of the way from A to C\. So coordinates:
Let me denote vector from A to C as \(c\_x, c\_y\)\. Then coordinates:
\- F is at \( \(1/7\)c\_x, \(1/7\)c\_y \) \- G is at \( \(5/7\)c\_x, \(5/7\)c\_y \)
Similarly, on AB, the points D and E are at:
\- D is at AD=4 from A, so since AB is along x\-axis from \(0,0\) to \(28,0\), D is at \(4, 0\) \- E is at AD \+ DE =4 \+16=20 from A, so E is at \(20,0\)
Now, the reflections:
M is the reflection of D through F\. So reflection through point F means that F is the midpoint of D and M\. So if D is \(4,0\), then M must satisfy:
F\_x = \(D\_x \+ M\_x\)/2 =\> M\_x = 2F\_x \- D\_x
Similarly, M\_y = 2F\_y \- D\_y
Since D is \(4,0\), and F is \( \(1/7\)c\_x, \(1/7\)c\_y \), so:
M\_x = 2\*\( \(1/7\)c\_x \) \-4 = \(2/7\)c\_x \-4
M\_y = 2\*\(1/7 c\_y \) \-0 = \(2/7\)c\_y
Similarly, N is the reflection of G through E\. So E is the midpoint of G and N\. So coordinates of N:
E\_x = \(G\_x \+ N\_x\)/2 =\> N\_x = 2E\_x \- G\_x
E\_y = \(G\_y \+ N\_y\)/2 =\> N\_y = 2E\_y \- G\_y
Given E is at \(20,0\), and G is at \(5/7 c\_x, 5/7 c\_y\), so:
N\_x = 2\*20 \- \(5/7\)c\_x =40 \- \(5/7\)c\_x
N\_y = 2\*0 \- \(5/7\)c\_y = \- \(5/7\)c\_y
Now, quadrilateral DEGF has area 288\. Let me figure out the coordinates of D, E, G, F to compute its area\.
Points of quadrilateral DEGF are D, E, G, F\. Wait, need to confirm the order\. The problem says "quadrilateral DEGF", so the order is D\-E\-G\-F\-D? Let me confirm with the TikZ code given\.
Looking at the TikZ code:
The code draws the quadrilateral DEGF? Wait, actually in the TikZ code, they have:
\\fill\[draw=black, fill=gray\!20\] \(N\) \-\- \(E\) \-\- \(M\) \-\- \(F\) \-\- cycle;
\\fill\[draw=black, fill=gray\!20\] \(N\) \-\- \(E\) \-\- \(C\) \-\- \(B\) \-\- cycle;
\\fill\[draw=black, fill=gray\!20\] \(A\) \-\- \(F\) \-\- \(M\) \-\- cycle;
And the heptagon is AFNBCEM\. Let me see the coordinates in the TikZ code for better understanding\. The coordinates in the TikZ code are given as:
A at \(100,100\), D at \(95,80\), F at \(130,80\), M at \(165,80\); N at \(0,50\), E at \(87\.5,50\), G at \(175,50\); B is defined as \(D\)\!2\!\(E\), which is a midpoint? Wait, in TikZ, the syntax \(D\)\!2\!\(E\) means the point dividing the segment DE externally in the ratio 2:1? Wait, actually in TikZ, the coordinate \(A\)\!t\!\(B\) gives a point along AB such that the ratio is t\. If t\>1, it’s beyond B\. So \(D\)\!2\!\(E\) would be a point such that D\-\-\-E\-\-\-\-B with DB = 2 DE? Wait, maybe I need not get bogged down here, but perhaps the figure is constructed with AB and AC not perpendicular, but in the code, coordinates are given with A at \(100,100\), D at \(95,80\), so moving left and down from A\. F is at \(130,80\), which is to the right of A\. So in this coordinate system, AB is not a straight horizontal line? Hmm, perhaps the TikZ code is just a rough drawing, and the exact coordinates there might not align with the problem’s numerical values, since in the problem AD=4, but in the code, the distance from A\(100,100\) to D\(95,80\) would be sqrt\(52\+202\)=sqrt\(425\) which is not 4\. So probably the TikZ coordinates are just for drawing, not to scale\.
Therefore, better to ignore the TikZ coordinates and just go by the problem’s description\.
Back to quadrilateral DEGF\. Let’s confirm the order of the points\. The name DEGF suggests the quadrilateral is D\-E\-G\-F\-D\. To compute its area, perhaps using coordinates\.
We have coordinates for D, E, G, F in terms of c\_x and c\_y\.
D is \(4,0\)
E is \(20,0\)
G is \( \(5/7\)c\_x, \(5/7\)c\_y \)
F is \( \(1/7\)c\_x, \(1/7\)c\_y \)
So quadrilateral DEGF has coordinates:
D: \(4,0\)
E: \(20,0\)
G: \(5c\_x/7, 5c\_y/7\)
F: \(c\_x/7, c\_y/7\)
We can compute its area using the shoelace formula\.
First, list the coordinates in order D\-E\-G\-F\-D\.
So the coordinates are:
\(4,0\), \(20,0\), \(5c\_x/7,5c\_y/7\), \(c\_x/7, c\_y/7\), back to \(4,0\)\.
Compute the area using shoelace:
Area = 1/2 \| \(x1y2 \- x2y1\) \+ \(x2y3 \- x3y2\) \+ \(x3y4 \- x4y3\) \+ \(x4y1 \- x1y4\) \| ?
Wait, let me recall the shoelace formula properly\. For coordinates \(x1,y1\), \(x2,y2\), \.\.\., \(xn,yn\), area is 1/2 \|sum\_\{i=1 to n\} \(xi yi\+1 \- xi\+1 yi\)\| where xn\+1=x1, yn\+1=y1\.
So let’s apply that\.
List the coordinates in order D, E, G, F, D\.
So:
Point 1: D: \(4,0\)
Point 2: E: \(20,0\)
Point 3: G: \(5c\_x/7, 5c\_y/7\)
▼\\blacktriangledownP1Point 4: F: \(c\_x/7, c\_y/7\)
Point 5: D: \(4,0\)
Compute the sum:
x1 y2 \- x2 y1: 4\*0 \- 20\*0 = 0
x2 y3 \- x3 y2: 20\*\(5c\_y/7\) \- \(5c\_x/7\)\*0 = 100 c\_y /7 \-0 = 100 c\_y /7
x3 y4 \- x4 y3: \(5c\_x/7\)\(c\_y/7\) \- \(c\_x/7\)\(5c\_y/7\) = \(5 c\_x c\_y\)/49 \- \(5 c\_x c\_y\)/49 = 0
x4 y5 \- x5 y4: \(c\_x/7\)\(0\) \- 4\*\(c\_y/7\) = 0 \- \(4 c\_y\)/7 = \-4 c\_y /7
▼\\blacktriangledownP2Then sum all these terms:
0 \+ 100 c\_y /7 \+ 0 \-4 c\_y /7 = \(96 c\_y\)/7
Then take absolute value and multiply by 1/2:
Area = \(1/2\) \|96 c\_y /7\| = 48 \|c\_y\| /7
But the area is given as 288, so:
48 \|c\_y\| /7 = 288
Thus,
\|c\_y\| = \(288 \*7\)/48 = \(288/48\)\*7 =6\*7=42
So \|c\_y\|=42, so c\_y=42 or \-42, but since area is positive, we can take c\_y=42 \(assuming the triangle is above the x\-axis\)\. So c\_y=42\.
So we now know the y\-coordinate of point C is 42\. Now we need to find c\_x? Wait, but perhaps we don’t need c\_x? Because in the area of the heptagon, maybe it cancels out?
Wait, let’s see\. The problem is to find the area of heptagon AFNBCEM\. Let’s try to figure out what this heptagon consists of\. The name AFNBCEM suggests the vertices are A, F, N, B, C, E, M, and back to A? Let me confirm the order\. The name is AFNBCEM, so the vertices in order are A, F, N, B, C, E, M, and then back to A? So it’s a seven\-sided figure\.
Alternatively, perhaps I need to figure out the structure by looking at the TikZ code’s fill commands\. The TikZ code fills three regions:
1\. \(N\) \-\- \(E\) \-\- \(M\) \-\- \(F\) \-\- cycle: quadrilateral NEMF
2\. \(N\) \-\- \(E\) \-\- \(C\) \-\- \(B\) \-\- cycle: quadrilateral NEBC
3\. \(A\) \-\- \(F\) \-\- \(M\) \-\- cycle: triangle AFM
But the heptagon is AFNBCEM, so maybe the heptagon is the area of the triangle ABC minus the areas of these three filled regions? Or perhaps it’s combining some areas\.
Alternatively, let’s try to list the vertices of the heptagon AFNBCEM in order\. The letters are A, F, N, B, C, E, M\. So starting at A, going to F, then to N, then to B, then to C, then to E, then to M, and back to A? Let me see if that makes a closed shape\. Let me try to visualize:
\- From A to F: along AC, since F is on AC\.
\- From F to N: Hmm, N is the reflection of G over E\. Since G is on AC, and E is on AB, so N is somewhere\.\.\. Let’s see coordinates\.
Wait, perhaps I need coordinates for all the points to figure this out\. Let me proceed step by step\.
We have coordinates in terms of c\_x and c\_y, but we know c\_y=42\. We still need to find c\_x? Wait, maybe the area of the triangle ABC can be expressed in terms of c\_x and c\_y, but since c\_y is known, but c\_x is still unknown\. Wait, but perhaps the area of ABC is not needed directly, but since the heptagon is part of ABC, maybe the total area can be found via relations\.
Alternatively, maybe c\_x can be arbitrary? But that can’t be, since the area of DEGF was given, but in our calculation for DEGF’s area, the c\_x didn’t matter\. The area only depended on c\_y\. Because in the shoelace formula for DEGF, the terms involving c\_x canceled out\. Let me check that again\.
Looking back at the shoelace computation for DEGF:
The coordinates were D\(4,0\); E\(20,0\); G\(5c\_x/7,5c\_y/7\); F\(c\_x/7, c\_y/7\)\.
The shoelace terms:
First term: x1y2 \- x2y1: 4\*0 \-20\*0 =0
Second term: x2y3 \-x3y2: 20\*\(5c\_y/7\) \- \(5c\_x/7\)\*0 = 100 c\_y /7
▼\\blacktriangledownP3Third term: x3y4 \-x4y3: \(5c\_x/7\)\(c\_y/7\) \- \(c\_x/7\)\(5c\_y/7\) = \(5c\_x c\_y\)/49 \- \(5c\_x c\_y\)/49 =0
Fourth term: x4y1 \-x1y4: \(c\_x/7\)\(0\) \-4\*\(c\_y/7\) = \-4 c\_y /7
So total sum: 100 c\_y /7 \-4 c\_y /7 =96 c\_y /7, so indeed only c\_y matters, so c\_x can be any value? But that seems odd\. How come the area of DEGF does not depend on c\_x?
Wait, perhaps because points D and E are on AB \(the x\-axis\), and points F and G are on AC\. The quadrilateral DEGF is a four\-sided figure with two points on AB and two on AC? Wait, actually D and E are on AB, but F and G are on AC, so DEGF is a quadrilateral with two vertices on AB and two on AC, connected how? Let me see the order D\-E\-G\-F\-D\. So from D\(4,0\) to E\(20,0\) along AB, then to G which is on AC, then to F which is also on AC, then back to D? That seems a bit irregular, but the shoelace formula still works regardless\.
But the area not depending on c\_x is interesting\. It must be that the area is determined purely by the vertical component, since moving along AC affects the y\-coordinate but since both F and G are along AC, their x\-coordinates depend on c\_x, but in the shoelace terms, the cross terms canceled out the c\_x\. So perhaps that’s correct\. So as a result, we only needed c\_y to compute that area, and c\_x remains free? But then how do we compute other areas?
Wait, but the area of triangle ABC is \(base \* height\)/2\. If AB is the base, length 28, then the height would be the altitude from C to AB, which is the y\-coordinate of C if AB is on x\-axis\. Since in our coordinate system, AB is on x\-axis from \(0,0\) to \(28,0\), and point C is at \(c\_x,42\), so the height is indeed 42, regardless of c\_x\. Therefore area of triangle ABC is \(28 \*42\)/2 = 28\*21=588\. Wait, is this correct? Wait, actually the area would be 1/2 \* base AB \* height, where height is the vertical distance from C to AB \(which is along x\-axis\)\. Since AB is on x\-axis, the y\-coordinate of C is indeed the height, so yes, area is \(28 \*42\)/2 = 588\. So the area of ABC is fixed at 588, regardless of c\_x? That’s interesting, because c\_x can be any value, but as long as the y\-coordinate is 42, the area remains the same\. So in this problem, even though we don’t know c\_x, the total area of ABC is fixed\. That might be key here\.
Therefore, perhaps the area of the heptagon can be found by subtracting certain areas from the total area of ABC\. Let me see\.
The heptagon is AFNBCEM\. Let me try to understand which regions are inside or outside\.
Alternatively, looking back at the TikZ code’s filled regions:
The code fills three regions:
1\. N\-E\-M\-F: quadrilateral
2\. N\-E\-C\-B: quadrilateral
3\. A\-F\-M: triangle
So perhaps the heptagon is the area of ABC minus these three filled regions? Let me check:
The heptagon AFNBCEM would consist of the following edges:
From A to F: along AC\.
F to N: from F to N, which is a point related to G and E\.
N to B: from N to B\.
B to C: along BC\.
C to E: but E is on AB, so from C to E? Wait, but E is on AB, so CE is a line from C to E on AB\.
Then E to M: from E to M, which is the reflection of D over F\.
Then M back to A? Wait, no, the name is AFNBCEM, so after M it should close back to A? Wait the vertices are A, F, N, B, C, E, M, and then back to A? So the edges are AF, FN, NB, BC, CE, EM, MA? Hmm, not sure\. Alternatively, maybe the heptagon is formed by connecting those points in order, so the area would be the polygon A\-F\-N\-B\-C\-E\-M\-A\.
To compute its area, perhaps it’s better to use coordinates\. Since we can express all points in terms of c\_x and c\_y, but we know c\_y=42, and c\_x is still unknown, but maybe in the end it will cancel out\.
Wait, but in the area of ABC is fixed at 588, so if I can express the heptagon’s area in terms of ABC’s area minus the other regions, maybe the unknown c\_x will not matter\.
Alternatively, let’s try to find coordinates for all the points involved\.
We have:
Point A: \(0,0\)
Point B: \(28,0\)
Point C: \(c\_x,42\)
Point D: \(4,0\)
Point E: \(20,0\)
Point F: \( \(1/7\)c\_x, \(1/7\)\(42\) \) = \(c\_x/7, 6 \)
Because c\_y=42, so F’s y\-coordinate is 42/7=6\.
Similarly, G is \(5/7 c\_x, 5/7 \*42\)= \(5c\_x/7, 30 \)
Then M is the reflection of D over F:
M\_x = 2\*\(c\_x/7\) \-4 = \(2c\_x\)/7 \-4
M\_y = 2\*6 \-0 =12
So M is at \( \(2c\_x/7 \-4\), 12 \)
Similarly, N is reflection of G over E:
E is at \(20,0\)\. So:
N\_x = 2\*20 \- \(5c\_x/7\) =40 \- \(5c\_x\)/7
N\_y = 2\*0 \-30 = \-30
Thus N is at \(40 \- \(5c\_x\)/7, \-30 \)
Now we have coordinates for all points:
A: \(0,0\)
F: \(c\_x/7, 6\)
N: \(40 \-5c\_x/7, \-30\)
B: \(28,0\)
C: \(c\_x,42\)
E: \(20,0\)
M: \(2c\_x/7 \-4, 12\)
Now, the heptagon is AFNBCEM, so the vertices in order are:
A \(0,0\),
F \(c\_x/7,6\),
▼\\blacktriangledownP4N \(40 \-5c\_x/7, \-30\),
B \(28,0\),
C \(c\_x,42\),
E \(20,0\),
M \(2c\_x/7 \-4, 12\),
and back to A \(0,0\)\.
To compute the area of this polygon, we can use the shoelace formula\. But since this involves variables c\_x, but the total area of ABC is fixed, perhaps when we compute the shoelace formula, the terms with c\_x will cancel out?
▼\\blacktriangledownP5Let me attempt to compute the area step by step\.
First, list all the vertices in order:
1\. A: \(0,0\)
2\. F: \(c\_x/7, 6\)
3\. N: \(40 \- \(5c\_x\)/7, \-30\)
4\. B: \(28,0\)
5\. C: \(c\_x,42\)
6\. E: \(20,0\)
7\. M: \( \(2c\_x\)/7 \-4, 12 \)
8\. Back to A: \(0,0\)
Now, apply shoelace formula\. Let’s denote each point as \(x\_i, y\_i\) for i from 1 to 7, then back to \(x\_8,y\_8\)=\(0,0\)\.
Compute the sum over i=1 to 7 of \(x\_i y\_\{i\+1\} \- x\_\{i\+1\} y\_i \)
Let me compute term by term\.
▼\\blacktriangledownP6Term 1: i=1 \(A to F\)
x1 y2 \- x2 y1 = 0 \*6 \- \(c\_x/7\)\*0 =0 \-0=0
Term 2: i=2 \(F to N\)
x2 y3 \- x3 y2 = \(c\_x/7\)\(\-30\) \- \[40 \-5c\_x/7\]\(6\)
Compute:
= \(\-30 c\_x\)/7 \- \[40\*6 \- \(5c\_x/7\)\*6\]
= \(\-30 c\_x\)/7 \- \[240 \- 30 c\_x /7 \]
▼\\blacktriangledownP7= \(\-30 c\_x\)/7 \-240 \+30 c\_x /7
= \(\-30 c\_x \+30 c\_x\)/7 \-240 = 0 \-240 = \-240
Term3: i=3 \(N to B\)
x3 y4 \- x4 y3 = \[40 \-5c\_x/7\]\(0\) \-28\*\(\-30\)
= 0 \-28\*\(\-30\) = \+840
Term4: i=4 \(B to C\)
x4 y5 \- x5 y4 =28\*42 \- c\_x \*0 = 1176 \-0 =1176
▼\\blacktriangledownP8Term5: i=5 \(C to E\)
x5 y6 \-x6 y5 = c\_x \*0 \-20\*42 =0 \-840= \-840
Term6: i=6 \(E to M\)
x6 y7 \-x7 y6 =20\*12 \- \[ \(2c\_x/7 \-4\) \]\*0 =240 \-0=240
Term7: i=7 \(M to A\)
x7 y8 \-x8 y7 = \[ \(2c\_x/7 \-4\) \]\*0 \-0\*12 =0 \-0=0
Now sum all these terms:
▼\\blacktriangledownP9Term1:0
Term2: \-240
Term3:\+840
Term4:\+1176
Term5:\-840
Term6:\+240
▼\\blacktriangledownP10Term7:0
Total sum: \(\-240\) \+840 =600; 600 \+1176=1776; 1776\-840=936; 936 \+240=1176\.
So total sum is 1176\. Then area is \(1/2\)\|1176\| =588\.
Wait, that’s exactly the area of triangle ABC, which was 588\. But that can’t be, because the heptagon is supposed to be inside the triangle, but according to this, its area is equal to the whole triangle? That must be wrong\. Probably I messed up the order of the points\. Because depending on the order \(clockwise or counter\-clockwise\), the shoelace formula can give negative, but absolute value\. But if the points are ordered in a way that the polygon winds around, it might subtract areas\.
Wait, let’s check the order of the points\. The heptagon is AFNBCEM\. Let me see the order:
Starting at A \(0,0\), going to F which is on AC, then to N\. Where is N? From coordinates, N has a y\-coordinate of \-30, which is below the x\-axis, since in our coordinate system AB is on x\-axis from \(0,0\) to \(28,0\), and C is at \(c\_x,42\)\. So point N is at \(40 \-5c\_x/7, \-30\), so it’s below AB\. Then from N to B \(28,0\), which is on AB\. Then to C \(c\_x,42\), up to the triangle, then to E \(20,0\) on AB, then to M, which is at \(2c\_x/7 \-4,12\), so y=12, which is above x\-axis, and then back to A\.
This path probably crosses over itself or encloses areas outside the triangle? For instance, going from N below AB to B on AB, then to C, then to E on AB, then to M, which is above\. This might create a non\-simple polygon or overlapping areas\. Hence, the shoelace formula might not work correctly if the points are not ordered properly \(either clockwise or counter\-clockwise without crossing\)\.
Alternatively, perhaps I ordered the points incorrectly\. The name AFNBCEM could be in a different order? Let me check the TikZ code again\.
In the TikZ code, the heptagon is not directly drawn, but the filled regions are three parts\. The code draws the triangle ABC, and some internal lines\. The heptagon is probably the area that’s left after subtracting the gray regions\. The gray regions are:
1\. N\-E\-M\-F: quadrilateral
2\. N\-E\-C\-B: quadrilateral? Wait, N\-E\-C\-B: connecting N to E to C to B\. Since N is below AB, and E and B are on AB, and C is the top vertex\. So this would be a quadrilateral that includes the area below AB? But the triangle ABC is above AB\. Hmm, perhaps in the diagram, the figure is drawn with AB and AC not along axes, so N might be inside?
Alternatively, given that in the shoelace calculation, I got the total area equal to ABC, which is suspicious, so likely the order of the points is incorrect\.
Alternatively, perhaps the correct order is different\. Let me try to think of the heptagon AFNBCEM\.
Starting at A, then F \(on AC\), then N\.\.\. where is N? Since N is reflection of G over E\. G is on AC, so reflecting over E \(on AB\) would place N somewhere\. Similarly, M is reflection of D over F, so reflecting D \(on AB\) over F \(on AC\) would place M somewhere\.
Alternatively, maybe the heptagon is formed by the outer edges and some internal lines\. Alternatively, perhaps it’s better to compute the area by breaking it into parts\.
Alternatively, note that in my shoelace calculation, I got 588, which is the total area, but that can’t be the heptagon\. So perhaps I have the wrong order\. Let me try reversing the order or adjusting the point sequence\.
Alternatively, let’s list the coordinates numerically with a sample c\_x\. Since c\_x is arbitrary except that it defines point C\. Since in the area calculations, the final area didn’t depend on c\_x, but in coordinates of the heptagon, it does\. Wait in the shoelace sum above, all the terms with c\_x canceled out, leading to total 1176, so area 588\. That is interesting\. Because when I computed term by term:
Looking back at the terms:
Term2 had \-240, but also terms with c\_x canceled\.
Term3: \+840 \(no c\_x\)
▼\\blacktriangledownE1Term4: \+1176 \(no c\_x\)
Term5: \-840 \(no c\_x\)
Term6:\+240 \(no c\_x\)
So all c\_x terms canceled out in the total sum, resulting in 1176, so area 588, same as ABC\. That suggests that regardless of c\_x, the heptagon AFNBCEM has the same area as ABC? But that contradicts the problem statement, since they mention the area of DEGF is 288, so probably the heptagon is smaller\.
This inconsistency suggests I have the wrong order of the points in the heptagon\. Let me check the TikZ code again for the structure\.
In the TikZ code:
\- It draws the triangle ABC\.
\- Draws lines D\-M and G\-N\.
\- Fills three regions:
1\. N\-E\-M\-F: a quadrilateral\.
2\. N\-E\-C\-B: another quadrilateral\.
3\. A\-F\-M: a triangle\.
So the gray areas are these three regions\. Then the heptagon AFNBCEM is probably the remaining area of ABC minus these three gray regions\. So total area would be Area ABC \- \(Area NEMF \+ Area NECB \+ Area AFM\)\.
If that is the case, then compute each of these areas and subtract from 588\.
Given that, let’s try this approach\.
First, Area ABC is 588\.
Now compute the areas of the three gray regions:
1\. Quadrilateral NEMF: points N, E, M, F\.
2\. Quadrilateral NEBC: points N, E, B, C? Wait the code says N\-E\-C\-B, so N to E to C to B to N? But B to N? Wait the coordinates: N is at \(0,50\) in the TikZ code, but in our coordinates it’s \(40 \-5c\_x/7, \-30\)\. Hmm, perhaps in actual coordinates it’s a quadrilateral N\-E\-C\-B, but depending on positions, this could be a non\-convex or crossing polygon\.
Alternatively, better to use coordinates\.
First, compute Area of NEMF:
Points N, E, M, F\.
Coordinates:
N: \(40 \-5c\_x/7, \-30\)
E: \(20,0\)
M: \(2c\_x/7 \-4, 12\)
F: \(c\_x/7,6\)
We can apply shoelace formula here\.
Order of points: N\-E\-M\-F\-N\.
List the coordinates:
1\. N: \(40 \-5c\_x/7, \-30\)
2\. E: \(20,0\)
3\. M: \(2c\_x/7 \-4, 12\)
4\. F: \(c\_x/7,6\)
Back to N\.
Compute shoelace sum:
Term1: x1 y2 \- x2 y1:
x1=40 \-5c\_x/7, y2=0; x2=20, y1=\-30
So term1: \(40 \-5c\_x/7\)\(0\) \-20\*\(\-30\) =0 \+600=600
Term2: x2 y3 \-x3 y2:
x2=20, y3=12; x3=2c\_x/7 \-4, y2=0
So 20\*12 \- \(2c\_x/7 \-4\)\*0 =240 \-0=240
Term3: x3 y4 \-x4 y3:
x3=2c\_x/7 \-4, y4=6; x4= c\_x/7, y3=12
So \(2c\_x/7 \-4\)\*6 \- \(c\_x/7\)\*12 = \(12 c\_x /7 \-24\) \-12 c\_x /7 = \-24
Term4: x4 y1 \-x1 y4:
x4= c\_x/7, y1=\-30; x1=40 \-5c\_x/7, y4=6
So \(c\_x/7\)\(\-30\) \- \(40 \-5c\_x/7\)\(6\) = \-30 c\_x /7 \-240 \+30 c\_x /7 = \-240
Total sum: 600 \+240 \+ \(\-24\) \+ \(\-240\) = 600 \+240=840; 840\-24=816; 816\-240=576
Area is 1/2 \|576\| =288\. Wait, that’s interesting\! The area of NEMF is 288, which is exactly the given area of DEGF\. But in the problem statement, it’s quadrilateral DEGF that has area 288, not NEMF\. Hmm\. So perhaps there is a confusion here\. Wait in the problem statement, it says "Quadrilateral DEGF has area 288", but according to my calculation here, quadrilateral NEMF has area 288, which matches\. So perhaps in the problem, DEGF is actually referring to NEMF? Or maybe a mislabeling?
Wait let me check coordinates of DEGF:
D is \(4,0\), E is \(20,0\), G is \(5c\_x/7,30\), F is \(c\_x/7,6\)\. So quadrilateral DEGF is D\-E\-G\-F\-D\.
Compute its area with shoelace:
Points D\(4,0\); E\(20,0\); G\(5c\_x/7,30\); F\(c\_x/7,6\); back to D\.
Compute terms:
Term1: x1 y2 \-x2 y1: 4\*0 \-20\*0=0
Term2: x2 y3 \-x3 y2:20\*30 \- \(5c\_x/7\)\*0=600 \-0=600
Term3:x3 y4 \-x4 y3: \(5c\_x/7\)\(6\) \- \(c\_x/7\)\(30\)= 30 c\_x /7 \-30 c\_x /7=0
Term4:x4 y1 \-x1 y4: \(c\_x/7\)\(0\) \-4\*6= 0 \-24= \-24
Total sum:0 \+600 \+0 \-24=576, so area 288\. Yes\! So DEGF has area 288, as given, which matches\. So in this case, the area of DEGF is indeed 288, as given, so that checks out\. So my initial calculation was correct, and that’s how we found c\_y=42\.
But then the area of NEMF is also 288? Wait no, in my calculation above for NEMF, I had coordinates N, E, M, F and got area 288 as well\. Wait but according to coordinates:
Wait DEGF area is 288, NEMF area is also 288? That can’t be a coincidence\. Wait let me recalculate NEMF area to confirm\.
NEMF points: N, E, M, F\.
Coordinates:
N: \(40 \-5c\_x/7, \-30\)
E: \(20,0\)
M: \(2c\_x/7 \-4, 12\)
F: \(c\_x/7,6\)
Applying shoelace:
Order N\-E\-M\-F\-N\.
Compute terms step by step:
List the coordinates:
1\. N: \(x1,y1\)= \(40 \-5c\_x/7, \-30\)
2\. E: \(x2,y2\)= \(20,0\)
3\. M: \(x3,y3\)= \(2c\_x/7 \-4,12\)
4\. F: \(x4,y4\)= \(c\_x/7,6\)
Back to N: \(x5,y5\)= \(x1,y1\)
Compute sum of x\_i y\_\{i\+1\} \- x\_\{i\+1\}y\_i:
Term1: x1 y2 \- x2 y1 = \(40 \-5c\_x/7\)\(0\) \-20\*\(\-30\) = 0 \+600=600
Term2: x2 y3 \-x3 y2 =20\*12 \- \(2c\_x/7 \-4\)\(0\) =240 \-0=240
Term3:x3 y4 \-x4 y3 = \(2c\_x/7 \-4\)\(6\) \- \(c\_x/7\)\(12\)
Let’s compute this:
First part: \(2c\_x/7 \*6\) \-4\*6 = \(12 c\_x\)/7 \-24
Second part: \- \(c\_x/7 \*12\) = \-12 c\_x /7
Total term3: \(12 c\_x /7 \-24\) \-12 c\_x /7 = \-24
Term4: x4 y1 \-x1 y4 = \(c\_x/7\)\(\-30\) \- \(40 \-5c\_x/7\)\(6\)
Compute:
First part: \-30 c\_x /7
Second part: \- \[40\*6 \- \(5c\_x/7\)\*6\] = \-\[240 \-30 c\_x /7\] = \-240 \+30 c\_x /7
Total term4: \(\-30 c\_x /7\) \+ \(\-240 \+30 c\_x /7\) = \-240
Term5: x1 y5 \-x5 y4? Wait no, in shoelace, it’s up to term4 because we have 4 points, so actually for quadrilateral, the terms are 1 to 4, and the total sum is terms1\-4\.
Wait in my earlier calculation I included term4 as x4 y1 \-x1 y4, but actually for a quadrilateral, the shoelace formula is sum over i=1 to 4 of \(xi yi\+1 \- xi\+1 yi\), with \(x5,y5\)=\(x1,y1\)\. So:
Total sum = Term1 \+ Term2 \+ Term3 \+ Term4
Where Term4 is x4 y1 \- x1 y4? Wait no:
Wait Term4 should be x4 y5 \- x5 y4, but y5 is y1, x5 is x1\. So:
Term4: x4 y1 \- x1 y4
Yes, that’s correct\. So the total sum is 600 \+240 \-24 \-240 = 600\+240=840; 840\-24=816; 816\-240=576\. So absolute value 576, area 288\. So yes, area of NEMF is also 288\. Interesting\.
But in the TikZ code, the first filled region is N\-E\-M\-F, so that is 288, same as DEGF\. So the problem mentions DEGF area 288, which is correct, and the gray area NEMF is also 288\. Then the other gray areas are NECB and AFM\.
So total gray area is 288 \(NEMF\) \+ area NEBC \+ area AFM\.
Then the heptagon’s area would be ABC area \(588\) minus these gray areas\.
So we need to compute area of NEBC and AFM\.
First, compute area of NEBC: points N, E, C, B\.
Points:
N: \(40 \-5c\_x/7, \-30\)
E: \(20,0\)
C: \(c\_x,42\)
B: \(28,0\)
Back to N\.
Apply shoelace formula\.
Order N\-E\-C\-B\-N\.
Coordinates:
1\. N: \(40 \-5c\_x/7, \-30\)
2\. E: \(20,0\)
3\. C: \(c\_x,42\)
4\. B: \(28,0\)
Back to N\.
Compute terms:
Term1: x1 y2 \-x2 y1 = \(40 \-5c\_x/7\)\(0\) \-20\*\(\-30\) =0 \+600=600
Term2: x2 y3 \-x3 y2 =20\*42 \- c\_x \*0=840 \-0=840
Term3: x3 y4 \-x4 y3 =c\_x \*0 \-28\*42=0 \-1176= \-1176
Term4: x4 y1 \-x1 y4 =28\*\(\-30\) \- \(40 \-5c\_x/7\)\(0\) = \-840 \-0= \-840
Total sum:600 \+840=1440; 1440 \-1176=264; 264 \-840= \-576
Area is \(1/2\)\|\-576\|=288\.
So area of NEBC is also 288\.
Then the third gray area is AFM: triangle A\-F\-M\.
Points A\(0,0\), F\(c\_x/7,6\), M\(2c\_x/7 \-4,12\)\.
Compute area via shoelace:
Coordinates:
A\(0,0\); F\(c\_x/7,6\); M\(2c\_x/7 \-4,12\); back to A\.
Terms:
Term1: x1 y2 \-x2 y1 =0\*6 \- \(c\_x/7\)\*0=0
Term2: x2 y3 \-x3 y2 = \(c\_x/7\)\(12\) \- \(2c\_x/7 \-4\)\(6\)
= \(12 c\_x\)/7 \- \[ \(2c\_x/7 \*6\) \-4\*6 \]
= \(12 c\_x\)/7 \- \(12 c\_x /7 \-24\)
= \(12 c\_x\)/7 \-12 c\_x /7 \+24 =24
Term3: x3 y1 \-x1 y3 = \(2c\_x/7 \-4\)\(0\) \-0\*12=0
Total sum:0 \+24 \+0=24\. Area=1/2 \|24\|=12\.
Thus area of AFM is 12\.
Therefore total gray area is 288 \+288 \+12= 588? Wait 288\+288=576 \+12=588\. So total gray area equals the area of ABC? That can’t be, since the heptagon is supposed to be the non\-gray area? But according to this, the gray areas sum up to the entire triangle, which would imply the heptagon has area zero, which is impossible\.
This inconsistency suggests that my assumption is wrong: that the heptagon is ABC minus the gray areas\. But according to TikZ code, the heptagon AFNBCEM is drawn, but the filled regions are three parts\. Maybe the heptagon is the union of some regions and excludes others? Alternatively, perhaps the heptagon includes some gray areas and excludes others?
Looking back at the TikZ code:
The code fills three regions:
1\. N\-E\-M\-F \(gray\)
2\. N\-E\-C\-B \(gray\)
3\. A\-F\-M \(gray\)
Then it draws the main triangle ABC, and draws lines D\-M and G\-N\.
The heptagon is named AFNBCEM\. Let’s see the vertices:
A, then F \(which is part of the gray triangle AFM\), so if AFM is gray, then the heptagon might exclude it? But then it goes to N, which is part of the gray quadrilaterals\. This is getting confusing\.
Alternatively, perhaps the heptagon is the intersection of non\-gray regions? But this is complicated\.
Alternatively, look at the Asymptote code’s draw commands:
It draws the triangle ABC, draws D\-M and G\-N\. Then it fills three regions\. The heptagon is probably the main figure minus those filled regions, but according to areas, that would be zero, which is wrong\. Hence, likely my approach is incorrect\.
Wait, let’s think differently\. Since in my first shoelace calculation of the heptagon AFNBCEM, I got area 588 equal to ABC, but the order of the points might have been incorrect, causing the polygon to overlap itself, thus canceling areas\. For example, going from N below AB to B, then to C, then to E on AB, then to M above, then back to A\-\-\-this path likely overlaps regions, leading the shoelace formula to subtract areas\.
Alternatively, arrange the points in correct order without crossing\.
Let me try to determine the correct order of the heptagon AFNBCEM\.
The name is A\-F\-N\-B\-C\-E\-M\-A\.
Let me plot approximate positions with a sample c\_x\. Let’s choose c\_x such that calculations are easy\. Since c\_y=42, let’s set c\_x=0 for simplicity? Wait if c\_x=0, then point C is at \(0,42\), so AC is vertical line, but then points F and G would be along y\-axis\. Let’s see:
If c\_x=0,
Then:
C is \(0,42\)
F is \(0/7,6\)=\(0,6\)
G is \(0,30\)
M is reflection of D\(4,0\) over F\(0,6\):
M\_x=2\*0 \-4= \-4, M\_y=2\*6 \-0=12→\\rightarrowM\(\-4,12\)
N is reflection of G\(0,30\) over E\(20,0\):
N\_x=2\*20 \-0=40, N\_y=2\*0 \-30= \-30→\\rightarrowN\(40, \-30\)
So coordinates would be:
A\(0,0\); F\(0,6\); N\(40,\-30\); B\(28,0\); C\(0,42\); E\(20,0\); M\(\-4,12\)
Now, let’s see the polygon AFNBCEM:
A\(0,0\) to F\(0,6\): up along y\-axis\.
F\(0,6\) to N\(40,\-30\): line going to the right and down sharply\.
N\(40,\-30\) to B\(28,0\): line to the left and up\.
B\(28,0\) to C\(0,42\): line to the left and up to \(0,42\)\.
C\(0,42\) to E\(20,0\): line to the right and down\.
E\(20,0\) to M\(\-4,12\): line to the left and up\.
M\(\-4,12\) back to A\(0,0\): line to the right and down\.
This seems a non\-intersecting polygon? Maybe\. Let’s compute its area with shoelace\.
List the coordinates in order:
1\. A: \(0,0\)
2\. F: \(0,6\)
3\. N: \(40, \-30\)
4\. B: \(28,0\)
5\. C: \(0,42\)
6\. E: \(20,0\)
7\. M: \(\-4,12\)
Back to A: \(0,0\)
Apply shoelace formula:
Compute terms:
Term1: x1 y2 \-x2 y1 =0\*6 \-0\*0=0
Term2:x2 y3 \-x3 y2 =0\*\(\-30\) \-40\*6= 0 \-240= \-240
Term3:x3 y4 \-x4 y3 =40\*0 \-28\*\(\-30\)=0 \+840=840
Term4:x4 y5 \-x5 y4 =28\*42 \-0\*0=1176 \-0=1176
Term5:x5 y6 \-x6 y5 =0\*0 \-20\*42=0 \-840= \-840
Term6:x6 y7 \-x7 y6 =20\*12 \- \(\-4\)\*0=240 \-0=240
Term7:x7 y1 \-x1 y7 =\-4\*0 \-0\*12=0 \-0=0
Total sum: 0 \-240 \+840 \+1176 \-840 \+240 \+0
Compute step by step:
Start at 0\.
\-240→\\rightarrow\-240
\+840→\\rightarrow600
\+1176→\\rightarrow1776
\-840→\\rightarrow936
\+240→\\rightarrow1176
Total sum 1176\. Area=1/2 \|1176\|=588, same as ABC\. So even with c\_x=0, the area is 588, which is the whole triangle, but visually, in this case, the polygon seems to cover the whole triangle? Let’s see:
With c\_x=0, triangle ABC is from \(0,0\) to \(28,0\) to \(0,42\)\. So it’s a right triangle with legs 28 and 42, area \(28\*42\)/2=588, correct\.
The heptagon AFNBCEM with coordinates as above: does it cover the entire triangle?
Looking at the path:
From A\(0,0\) to F\(0,6\): along left edge\.
Then to N\(40,\-30\): way outside the triangle \(since triangle’s x only up to 28, and y down to 0\)\. So this point is outside the triangle, so the polygon goes outside, hence covering area outside the triangle, leading the total area to still be 588 but including external areas\. Thus, my initial approach is wrong\.
This indicates that the order of the points in the heptagon is incorrect, or perhaps the heptagon is defined differently\.
Alternatively, perhaps the correct heptagon is ABC minus the three gray areas, but since the gray areas sum to 588, that can’t be\. Wait in this c\_x=0 case, the gray areas:
Compute gray areas:
1\. NEMF: N\(40,\-30\), E\(20,0\), M\(\-4,12\), F\(0,6\)\. This is a quadrilateral partly outside the triangle\.
2\. NEBC: N\(40,\-30\), E\(20,0\), C\(0,42\), B\(28,0\)\. This also extends outside\.
3\. AFM: A\(0,0\), F\(0,6\), M\(\-4,12\)\. This is a triangle outside the main triangle ABC \(since M is at \-4,12\)\.
Thus the gray areas include regions outside ABC, so subtracting them from ABC doesn’t make sense\.
This is getting too confusing\. Maybe another approach\.
Let me recall that in the problem, the area of DEGF is 288, which we used to find c\_y=42, and total area ABC is 28\*42/2=588\. The question is to find the area of heptagon AFNBCEM\.
Perhaps the heptagon can be expressed as ABC minus the areas of certain triangles or quadrilaterals, whose areas we can compute\.
Looking at the figure description:
The heptagon is AFNBCEM\. The points involved are A, F, N, B, C, E, M\.
Looking at the original triangle ABC, the heptagon is formed by connecting these points, possibly cutting off some corners\.
The points D and G are on AB and AC, but the heptagon doesn’t include D or G, but includes their reflections M and N\.
Reflections: M is reflection of D over F, so FM = FD, and similarly EN = EG\.
Given AD=4, DE=16, so AE=AD\+DE=20, and EB=8, so AB=28\.
On AC: AF=13, FG=52, so AG=65, and GC=26, AC=91\.
Since M is reflection of D over F, then F is midpoint of D and M, so DM=2 FD\. Since FD is the distance from F to D\.
Wait, but we need to find coordinates or vectors?
Alternatively, use mass point or area ratios\.
Alternatively, note that reflections might create parallelograms\. For example, reflecting D over F to get M implies that DF = FM, and similarly for N\.
Thus, quadrilateral DFM is a parallelogram? Wait, reflection over a point makes it so that F is midpoint, so vector from D to F is same as from F to M, so DM is twice DF, and DF and FM are equal and opposite\.
Similarly for N: reflection of G over E, so E is midpoint of G and N, so EN = EG\.
Given that, perhaps areas can be related\.
But I need to relate this to the heptagon\.
Alternatively, think of the heptagon as ABC with certain triangles removed or added\.
The heptagon AFNBCEM consists of the following edges:
From A to F: along AC\.
F to N: connecting to N, which is reflection of G over E\.
N to B: connecting to B\.
B to C: edge of triangle\.
C to E: connecting to E on AB\.
E to M: connecting to M, reflection of D over F\.
M to A: connecting back\.
This path likely excludes the areas of triangles AFD, some quadrilaterals, but it’s complicated\.
Alternatively, use coordinate geometry with c\_x arbitrary but c\_y=42, and notice that in the shoelace formula for the heptagon, the area was always 588 regardless of c\_x, but that includes areas outside ABC\. But the problem likely wants the area within the triangle ABC\.
Wait in the shoelace calculation with c\_x=0, the heptagon’s area equaled ABC’s area, but part of the heptagon was outside ABC, so the actual area inside ABC would be less\. But without knowing c\_x, can’t compute\.
Wait but in our first calculation with general c\_x, the area came out exactly as ABC’s area, but perhaps because the parts outside and inside cancel out? For example, when c\_x is such that point C is to the right, then N might be to the left, and vice versa, but the total area remains same\.
Alternatively, perhaps the correct area is 588 \- 288 \-288 \-12 =0, which is impossible, so definitely wrong\.
Wait, problem says "Find the area of heptagon AFNBCEM"\. Given that in my first shoelace with coordinates in the order given, it’s equal to ABC, but perhaps the correct area is 588 \- \(288 \+ 288 \+12\) \+ some overlaps? This is too confusing\.
Alternative approach: use vectors\.
Let me denote vectors with position vectors\.
▼\\blacktriangledownE2Let me set coordinate system with A as origin\.
Let vector AB = \*\*b\*\*, and vector AC = \*\*c\*\*\. So coordinates:
\- Point A: \*\*0\*\*
\- Point B: \*\*b\*\*
\- Point C: \*\*c\*\*
\- Point D is on AB with AD=4, AB=28, so AD/AB=4/28=1/7, so position vector of D is \(1/7\)\*\*b\*\*
Similarly, AE = AD \+ DE =4\+16=20, so AE/AB=20/28=5/7, so E is at \(5/7\)\*\*b\*\*
On AC:
AF=13, AC=91, so AF/AC=13/91=1/7, so F is at \(1/7\)\*\*c\*\*
FG=52, so from A to G is 13\+52=65, so AG/AC=65/91=5/7, so G is at \(5/7\)\*\*c\*\*
Now, M is reflection of D over F\. So:
The reflection of D over F means that F is the midpoint of D and M\. So:
\*\*F\*\* = \( \*\*D\*\* \+ \*\*M\*\* \) /2→\\rightarrow\*\*M\*\* = 2\*\*F\*\* \- \*\*D\*\*
Similarly, N is reflection of G over E:
\*\*N\*\* = 2\*\*E\*\* \- \*\*G\*\*
Compute \*\*M\*\* and \*\*N\*\*:
\*\*M\*\* = 2\*\(1/7 \*\*c\*\*\) \- \(1/7 \*\*b\*\*\) = \(2/7\)\*\*c\*\* \- \(1/7\)\*\*b\*\*
\*\*N\*\* = 2\*\(5/7 \*\*b\*\*\) \- \(5/7 \*\*c\*\*\) = \(10/7\)\*\*b\*\* \- \(5/7\)\*\*c\*\*
Now, we need to find the area of heptagon AFNBCEM\. Let’s list the points in order:
A, F, N, B, C, E, M, A\.
Expressed as vectors:
A: \*\*0\*\*
F: \(1/7\)\*\*c\*\*
N: \(10/7 \*\*b\*\* \-5/7 \*\*c\*\*\)
B: \*\*b\*\*
C: \*\*c\*\*
E: \(5/7\)\*\*b\*\*
M: \(\-1/7 \*\*b\*\* \+ 2/7 \*\*c\*\*\)
Back to A: \*\*0\*\*
To compute the area of polygon with vertices in order, we can use the shoelace formula in vectors, but it’s complex\. Alternatively, divide the heptagon into triangles or other shapes\.
Alternatively, use the shoelace formula in terms of \*\*b\*\* and \*\*c\*\*\.
But since area is a scalar, we need to express it in terms of the area of the triangle ABC, which is \(1/2\)\|\*\*b\*\*×\\times\*\*c\*\*\| =588\. Since we know \|\*\*b\*\*×\\times\*\*c\*\*\| = 2\*588=1176\.
Let me denote the cross product \*\*b\*\*×\\times\*\*c\*\* = 1176 \(taking magnitude, assuming orientation\)\.
Now, express all the coordinates in terms of \*\*b\*\* and \*\*c\*\*, but this might get involved\. Let’s assign coordinates symbolically\.
Let me set coordinate system with point A at \(0,0\), vector AB along x\-axis, so \*\*b\*\* = \(28,0\), since AB length is28\. Then point B is at \(28,0\)\. Vector \*\*c\*\* is \(p,42\), since the y\-coordinate is 42 \(as found earlier\), and p is the x\-coordinate of point C, which we don’t know\. So \*\*c\*\* = \(p,42\)\. Then the cross product \*\*b\*\*×\\times\*\*c\*\* is 28\*42 \-0\*p = 28\*42=1176, which matches, since area ABC is 588=1176/2\.
Now, express all points in coordinates:
A: \(0,0\)
F: \(1/7 \*\*c\*\*\) = \(p/7, 6\) since 42/7=6
N: \(10/7 \*\*b\*\* \-5/7 \*\*c\*\*\) = \(10/7 \*28, 0\) \-5/7\*\(p,42\) = \(40,0\) \- \(5p/7, 30\) = \(40 \-5p/7, \-30\)
B: \(28,0\)
C: \(p,42\)
E: \(5/7 \*\*b\*\*\) = \(5/7 \*28, 0\) = \(20,0\)
M: \(\-1/7 \*\*b\*\* \+ 2/7 \*\*c\*\*\) = \(\-4,0\) \+ \(2p/7, 12\) = \(2p/7 \-4, 12\)
So these are the same coordinates as before\. So now, with \*\*b\*\*=\(28,0\), \*\*c\*\*=\(p,42\)\.
Now, to compute the area of heptagon AFNBCEM with coordinates:
1\. A: \(0,0\)
2\. F: \(p/7,6\)
3\. N: \(40 \-5p/7, \-30\)
4\. B: \(28,0\)
5\. C: \(p,42\)
6\. E: \(20,0\)
7\. M: \(2p/7 \-4,12\)
Back to A\.
We can apply shoelace formula here\. Let’s denote each point as \(xi,yi\):
x0=0, y0=0
x1=p/7, y1=6
x2=40 \-5p/7, y2=\-30
x3=28, y3=0
x4=p, y4=42
x5=20, y5=0
x6=2p/7 \-4, y6=12
x7=0, y7=0
Compute the shoelace sum S = sum\_\{i=1 to 6\} \(xi yi\+1 \- xi\+1 yi\)
Compute term by term:
Term1: x1 y2 \- x2 y1 = \(p/7\)\(\-30\) \- \(40 \-5p/7\)\(6\)
= \-30p/7 \-240 \+ 30p/7 = \-240
Term2: x2 y3 \-x3 y2 = \(40 \-5p/7\)\(0\) \-28\*\(\-30\) =0 \+840=840
Term3:x3 y4 \-x4 y3=28\*42 \- p\*0=1176 \-0=1176
Term4:x4 y5 \-x5 y4= p\*0 \-20\*42=0 \-840= \-840
Term5:x5 y6 \-x6 y5=20\*12 \- \(2p/7 \-4\)\*0=240 \-0=240
Term6:x6 y7 \-x7 y6= \(2p/7 \-4\)\(0\) \-0\*12=0 \-0=0
Total sum S= Term1 \+ Term2 \+ Term3 \+ Term4 \+ Term5 \+ Term6
= \(\-240\) \+840 \+1176 \-840 \+240 \+0
Calculate step by step:
Start at 0:
\-240 \+840 =600
600 \+1176=1776
1776\-840=936
936\+240=1176
So S=1176, area=1/2 \|1176\|=588, same as before\.
This confirms that no matter the value of p \(i\.e\., the x\-coordinate of C\), the area of the heptagon is always 588, equal to the area of triangle ABC\. But this seems contradictory, unless the heptagon coincides with ABC, but looking at the coordinates, when p is chosen such that point C is to the right, the heptagon may include areas outside ABC, but mathematically the shoelace formula accounts for net area, considering orientation\.
But the problem is from an Olympiad, so likely the answer is 588 \- 288 \(area of DEGF\) but no, DEGF is given as 288\. Alternatively, the heptagon’s area is ABC minus the two quadrilaterals and the triangle, but they summed to 588, so 588\-588=0, which is impossible\.
Wait, but in the problem statement, it’s possible that the heptagon is the intersection or something else, but according to calculations, with coordinates, it’s equal to ABC\. But in the Asymptote code, the heptagon is drawn with those points, and the filled areas are separate, so perhaps the answer is 588 \- \(288 \+ 288 \+12\) = 588 \- 588 =0? No\.
Alternatively, maybe I miscounted the gray areas\. The problem mentions quadrilateral DEGF has area 288, but in my calculation DEGF area is 288, and NEMF also 288, but maybe the problem considers only DEGF, and the heptagon excludes it? But the heptagon’s area came as 588 regardless\.
Alternatively, perhaps the question had a typo and the heptagon is actually the quadrilateral DEGF, but no, it specifies heptagon\.
Wait, going back to the original problem statement:
"Quadrilateral DEGF has area 288\. Find the area of heptagon AFNBCEM\."
Possibly the heptagon is ABC minus DEGF and some other regions? But in my calculation, ABC is 588, DEGF is 288, but then need more info\.
▼\\blacktriangledownE3Alternatively, perhaps my coordinate approach is missing that the reflections M and N create figures whose areas relate\.
Note that M is reflection of D over F, so the area related to D and M might be twice the area of D to F or something\.
Alternatively, the area of the heptagon can be computed as the area of ABC plus the areas of the reflections? Not sure\.
Wait, let’s think of transformations\. Reflecting D over F to get M: this is a translation? No, reflection through a point is equivalent to a 180 degree rotation about F\. So the segment DM is such that F is midpoint, and this transformation maps D to M\. Similarly for G to N over E\.
The area of figure AFNBCEM might be related to ABC with some added or subtracted areas due to the reflections\.
Alternatively, consider that the heptagon can be divided into parts whose areas can be expressed via the given 288\.
Wait, let’s see what’s inside the heptagon\.
Looking at the vertices:
A, then F \(on AC\), then N \(reflection of G over E\), then B, then C, then E \(on AB\), then M \(reflection of D over F\), then back to A\.
This path likely encloses the following areas:
\- The main triangle ABC, but with certain internal areas subtracted or added\.
Alternatively, the heptagon comprises:
\- The quadrilateral AFM \(area 12\), but in the heptagon, from A\-F\-N\.\.\., so it might exclude AFM?
▼\\blacktriangledownE4Wait this is too vague\. Alternatively, notice that in my coordinate calculation, despite the path going outside the triangle, the net area equals ABC\. But the problem might be set in such a way that regardless of the position of C \(as long as height is 42\), the heptagon’s area equals ABC’s area, so the answer is 588\. But 588 is more than 999? No, 588 is less than 999, but the problem says "an integer between 0 and 999", so possible, but I need to confirm\.
Wait, but in the problem statement from competitions, often the area would be an integer, 588 is possible, but let me see with c\_x arbitrary\. Wait in our first calculation with c\_x=0, the area was 588, but part of the heptagon was outside ABC\. But perhaps in the actual figure, due to the reflection points, the heptagon is exactly the triangle ABC? But the names of the vertices suggest otherwise\.
Wait in the Asymptote code, the heptagon is drawn as AFNBCEM, and in the code’s coordinates:
A is at \(100,100\), F at \(130,80\), which is to the right and down from A, then N is at \(0,50\), which is far to the left, then to B at \(D\)\!2\!\(E\)\. Given the code’s coordinates, it’s likely a non\-convex heptagon covering the same area as ABC through some symmetry\.
Alternatively, the area is indeed 588, so the answer is 588\. But I recall that in some olympiad problems, when reflections are involved and areas are preserved, the total area remains same\.
Wait, reflecting points over other points doesn’t change the overall area? For instance, reflecting D over F to get M: the figure might be transformed but area preserved\. However, the heptagon includes both original and reflected points, so it’s complicated\.
Alternatively, note that the mapping from D to M and G to N involves translations? Wait reflection through a point is equivalent to a 180 rotation about that point, which is an isometry, preserving areas\. But combining these might not directly help\.
Alternatively, the heptagon AFNBCEM can be seen as the original triangle ABC plus the quadrilaterals related to the reflections, but areas might cancel\.
But according to coordinate calculations, it’s exactly ABC’s area, so maybe the answer is 588\. But I need to verify once more\.
Wait in the problem statement from the user, they mention "heptagon AFNBCEM", and in their Asymptote code, they have specific coordinates where the area might be different, but since in the code the coordinates are not to scale, as I saw distances don’t match, so we must go by the mathematical approach\.
Given that through coordinate analysis with general c\_x \(as long as c\_y=42\), the heptagon’s area is always 588, I think this is the answer\. But let me check with c\_x=28 for example\.
Let me choose c\_x=28, so point C is at \(28,42\), making triangle ABC a right triangle with legs 28 and 42, same as before\. Then:
F is at \(28/7,6\)=\(4,6\)
G is at \(5\*28/7,30\)=\(20,30\)
M is reflection of D\(4,0\) over F\(4,6\): midpoint is F, so M must be \(4,12\), since from D\(4,0\) to F\(4,6\) is \+6 in y, so M is \+6 more: \(4,12\)
N is reflection of G\(20,30\) over E\(20,0\): midpoint is E, so N is \(20, \-30\)
Now coordinates of heptagon AFNBCEM:
A\(0,0\); F\(4,6\); N\(20,\-30\); B\(28,0\); C\(28,42\); E\(20,0\); M\(4,12\); back to A\.
Applying shoelace:
List the coordinates:
1\. \(0,0\)
2\. \(4,6\)
3\. \(20,\-30\)
4\. \(28,0\)
5\. \(28,42\)
6\. \(20,0\)
7\. \(4,12\)
Back to \(0,0\)
Compute terms:
Term1: x1y2 \-x2y1 =0\*6 \-4\*0=0
Term2:x2y3 \-x3y2=4\*\(\-30\) \-20\*6= \-120 \-120= \-240
Term3:x3y4 \-x4y3=20\*0 \-28\*\(\-30\)=0 \+840=840
Term4:x4y5 \-x5y4=28\*42 \-28\*0=1176 \-0=1176
Term5:x5y6 \-x6y5=28\*0 \-20\*42=0 \-840= \-840
Term6:x6y7 \-x7y6=20\*12 \-4\*0=240 \-0=240
Term7:x7y1 \-x1y7=4\*0 \-0\*12=0\-0=0
Total sum: 0 \-240 \+840 \+1176 \-840 \+240 \+0
Calculations:
\-240 \+840 =600
600 \+1176 =1776
1776 \-840 =936
936 \+240 =1176
Area=1176/2=588\. Same result\.
But visually, in this case, triangle ABC is right triangle with vertices \(0,0\),\(28,0\),\(28,42\)\. The heptagon goes from A\(0,0\) to F\(4,6\), then to N\(20,\-30\) which is outside the triangle, then to B\(28,0\), up to C\(28,42\), down to E\(20,0\), to M\(4,12\), and back\. This clearly includes areas outside the original triangle, but the total signed area still computes to 588, same as ABC\. But the actual area "inside" might be different, but the problem doesn’t specify; it just says "area of heptagon AFNBCEM", which is a polygon, and by shoelace it’s 588\.
But the problem statement is from an Olympiad, likely answer is different\. Wait, the area of DEGF is 288, and total area ABC is 588, so 588\-288=300, but not sure\.
Alternatively, perhaps I misread the heptagon’s name\. The user wrote "heptagon AFNBCEM"\. Let me count the letters: A,F,N,B,C,E,M→\\rightarrow7 points, so heptagon\. But in my calculation it’s 588\. But let me see the Asymptote code’s filled regions: they filled three regions with total area 288\+288\+12=588, so if the heptagon is the remaining, it would be zero, but that’s impossible, so likely my misinterpretation\.
Wait in the Asymptote code, the filled regions are:
1\. N\-E\-M\-F: gray
2\. N\-E\-C\-B: gray
3\. A\-F\-M: gray
Then the heptagon AFNBCEM is drawn on top? Or perhaps the heptagon is the union of non\-gray regions? But the non\-gray region would be ABC minus the three gray areas, which is zero\. So this is conflicting\.
Alternatively, the heptagon comprises the three gray areas? But three regions totaling 588, but it’s called a heptagon\.
Alternatively, my only mathematical conclusion is that via coordinates, the heptagon has area 588, so I think that is the answer, but I need to verify problem constraints\.
Wait the problem says "points A, D, E, and B lie in that order on side AB", so AD=4, DE=16, EB=8\. Similarly on AC: A,F,G,C with AF=13, FG=52, GC=26\. So all points are in order, no overlaps\.
The reflections: M is reflection of D over F\. Since F is on AC, and D is on AB, M would be somewhere inside or outside the triangle? Similarly N is reflection of G over E, with G on AC and E on AB, so N would be outside\.
But according to calculations, the area is fixed at 588, which is ABC’s area, so I think this is the answer\. Maybe the trick is that the heptagon’s area equals the whole triangle\.
But to confirm, in my c\_x=28 example, even with the heptagon extending outside, but mathematically the area is same\. Perhaps in the problem’s configuration, due to the reflection, the total area cancels out to the original triangle\.
Thus I think the answer is 588, but let me see if 588 is between 0\-999, yes, so \\boxed\{588\}
Wait, but I recall that in some similar problems, the area might be double or something\. Wait let’s see the area of DEGF is 288, and in my calculation DEGF area is 288, which came from 48\|c\_y\|/7 =288, so \|c\_y\|=42, leading to ABC area 28\*42/2=588\. So it’s consistent\.
Alternatively, maybe the heptagon is ABC minus DEGF, so 588\-288=300? But why?
Alternatively, the area of AFNBCEM is equal to area of ABC plus area of DEGF? 588\+288=876? But no basis\.
Alternatively, note that the three gray areas in the Asymptote code sum to 588, so if the heptagon is what’s left, but it would be zero\. But since the code draws the heptagon, perhaps it’s the combination of the gray areas? But three regions, not a heptagon\.
Wait the Asymptote code draws:
\- The main triangle ABC\.
\- Draws lines D\-M and G\-N\.
\- Fills three regions: NEMF, NEBC, and AFM\.
Then the heptagon AFNBCEM is probably the outer boundary, which would be the same as ABC, but according to the coordinates in the Asymptote code, it’s not\. For example, in the code, point N is at \(0,50\), which is outside triangle ABC \(if ABC has points at \(100,100\), etc\.\), but this is just drawing code\.
Given the mathematical analysis shows the area is 588 regardless of c\_x, and the problem gives enough data to find ABC area as 588 via DEGF area, I think the answer is 588\. But I recall that in some reflection problems, the area might be double, but here reflections are through points, not lines\.
▼\\blacktriangledownE5Alternatively, cross\-verifying with mass point or another method\.
The coordinates method is solid here, despite the heptagon extending outside, but the question just asks for the area of the polygon AFNBCEM, which is a mathematical polygon with those vertices, and by shoelace it’s 588\. So I think that is the answer\.
But I’m unsure because 588 seems straightforward, but maybe I missed a negative sign? In the shoelace formula, if the points are ordered clockwise, the area would be negative, but absolute value is taken, so still 588\.
Yes, I think the answer is 588, so I will go with \\boxed\{588\}
Wait, but hold on\! In the initial problem statement, it says "the area of heptagon AFNBCEM"\. But according to the Asymptote code’s filled regions, the gray areas include parts that might be inside the heptagon\. For example, the region AFM is filled, but AFM is part of the heptagon’s edge? From A\-F\-M\-A is a triangle, but in the heptagon, the path is A\-F\-N\-\.\.\.\-M\-A, so it doesn’t include the interior of AFM\. Thus the heptagon excludes the gray triangle AFM \(area12\) and includes other areas\. Similarly, the quadrilaterals NEMF and NEBC may be partially inside or outside\.
This is too ambiguous without seeing the figure, but according to the coordinate calculation with two different c\_x values giving 588, I think it’s safe\. But let me see an official solution approach if I recall\.
Wait, no, but I just realized in my general calculation, the area was always 588, so I think that’s the answer\. Maybe the test\-maker intended that, as the area remains constant due to the reflections, so \\boxed\{588\} is the answer\.
But wait, hang on, I just found an official problem? Wait no, but let me think again\. When I calculated the area of the heptagon as 588 using coordinates, but in the problem, they mention "quadrilateral DEGF has area 288", which we used to find that the height is 42, leading to ABC area 588, so if the heptagon is ABC, then answer is 588, but the name heptagon suggests it’s not the whole triangle\. But with seven vertices, it must be different\. For example, in my c\_x=28 case:
The heptagon has points:
A\(0,0\), F\(4,6\), N\(20,\-30\), B\(28,0\), C\(28,42\), E\(20,0\), M\(4,12\)
Plotting these, the polygon would have a "bite" taken out near the bottom, but extending below the triangle, but the total area computation includes that negative space as positive due to the order\. But mathematically the shoelace counts the net area, so if the polygon is simple \(non\-intersecting\), but in this case, with N at \(20,\-30\), the polygon would intersect itself?
Check if the polygon AFNBCEM is simple:
List the points in order:
▼\\blacktriangledownE61\. A\(0,0\)
2\. F\(4,6\)
3\. N\(20,\-30\)
4\. B\(28,0\)
5\. C\(28,42\)
6\. E\(20,0\)
7\. M\(4,12\)
8\. A\(0,0\)
Check for intersections between edges\.
Edge AF: from \(0,0\) to \(4,6\)
Edge FN: \(4,6\) to \(20,\-30\): this line goes down sharply\.
Edge NB: \(20,\-30\) to \(28,0\): up to B\.
Edge BC: to C\(28,42\)
Edge CE: down to E\(20,0\)
Edge EM: to M\(4,12\)
Edge MA: to A\.
Visually, this likely intersects itself between EM and AF or something, making it a complex polygon, but shoelace formula still applies if we traverse the boundary once, but the area might not be what’s expected\.
However, in all coordinate examples I tried, the area was 588, so unless there is a miscalculation, I think the answer is 588\. But I recall that in some cases, the area might need to subtract, but according to calculations it’s not\.
Wait let me do a small c\_x example, say c\_x=0, then:
Heptagon area 588, but the actual figure includes area below AB \(from N at \(40,\-30\)\), so total area would be ABC area \(588\) plus the area below AB\. But in shoelace, since the polygon winds around, the net area is still 588? Let’s see:
▼\\blacktriangledownE7In c\_x=0 case:
Coordinates:
A\(0,0\); F\(0,6\); N\(40,\-30\); B\(28,0\); C\(0,42\); E\(20,0\); M\(\-4,12\)
Compute area via shoelace as 588, but let’s compute the area manually:
The polygon can be divided into parts\.
From A\(0,0\) to F\(0,6\) to N\(40,\-30\) to B\(28,0\) to C\(0,42\) to E\(20,0\) to M\(\-4,12\) to A\.
This is complex, but let’s see:
The part from A\-F\-N\-B\-A is a quadrilateral\. From B\-C\-E\-M\-A is another part\.
But this is too time\-consuming\. Alternatively trust the shoelace formula\.
Given that in two different c\_x the area is 588, I think it’s safe\. So I think the answer is \\boxed\{588\}\. But I have a doubt because the problem says "heptagon", but according to this it’s the whole triangle’s area\. Maybe the correct answer is 588 \- 288 = 300? But why subtract DEGF?
Alternatively, the area of the heptagon is ABC minus DEGF, so 588\-288=300\. But in coordinate calculation it’s not\. Alternatively, twice?
Wait, in the Asymptote code, they filled three regions: NEMF \(288\), NEBC \(288\), and AFM\(12\)\. Total 588, so the remaining area is zero, but the heptagon might be their union? But three regions can’t form a heptagon\.
Alternatively, the heptagon is N\-E\-M\-F\-A\-F\-M? No\.
Given the time I’ve spent and calculations consistently give 588, I think I will proceed with that\.
▼\\blacktriangledownE8But wait, hold on\! I just realized the problem says "the area of heptagon AFNBCEM"\. In my coordinate examples, when c\_x=28, the heptagon includes the point N\(20,\-30\) which is outside triangle ABC\. If the problem considers only the area inside triangle ABC, then we need to calculate that\.
In the c\_x=28 case:
Triangle ABC is from \(0,0\), \(28,0\), \(28,42\)\. The heptagon AFNBCEM has parts inside and outside\.
To compute the area inside ABC, we need to clip the heptagon to the triangle\.
This complicates, but perhaps in the problem’s configuration, due to the reflection points, the heptagon lies entirely within ABC?
Check with c\_x=28:
Point N is \(20, \-30\), which is outside ABC \(which is above y=0 except on AB\)\. So part of the heptagon is outside\. Similarly, point M is \(4,12\) which is inside\.
Thus part of the heptagon is outside ABC, so the area inside ABC would be less than 588\.
But this complicates, and the problem didn’t specify to consider only the interior\. But likely in competitions, the area is considered as the absolute value from shoelace, so 588\. But I need to verify once more with the initial data\.
Wait, the key was that from DEGF area we found c\_y=42, leading to ABC area 588, and the heptagon area via coordinates is same, so answer is 588\. I think I’ll go with that\.
But let me see similar problems\. For example, if you reflect points and form a polygon, sometimes areas add up\. But here reflections are through points, so perhaps area relations\.
The area of AFNBCEM:
It can be seen as ABC plus the areas of the reflected parts? But M and N are reflections, so areas might cancel\.
▼\\blacktriangledownE9Alternatively, the heptagon consists of ABC minus the areas of DFM and ENG or something\.
Alternatively, note that:
The heptagon AFNBCEM can be partitioned into:
\- Triangle AFC: but F is on AC\.
Wait this is too vague\.
Alternatively, note that in the coordinate system, regardless of c\_x, area is 588, so I think the answer is 588\. So I write \\boxed\{588\}
Wait, but in the initial problem statement from the user, the Asymptote code might have a different configuration\. Looking at the Asymptote code coordinates:
▼\\blacktriangledownE10In the code:
A is at \(100,100\)
D is at \(95,80\) \-\-\- so left and down from A
F is at \(130,80\) \-\-\- right and down from A
M is at \(165,80\) \-\-\- further right
N is at \(0,50\) \-\-\- far left and down
E is at \(87\.5,50\) \-\-\- left of A?
G is at \(175,50\) \-\-\- right
B is defined as \(D\)\!2\!\(E\), which in Asymptote means extending from D through E twice the length? So from D to E is vector, then extend\.
Similarly C is \(F\)\!2\!\(G\)\.
This suggests B is located such that E divides DB in ratio 2:1, so BE:ED=2:1? Wait, if it’s \(D\)\!2\!\(E\), that would be a point such that D\-\-\-E\-\-\-\-B with DE:EB=1:2? Because the syntax \(A\)\!t\!\(B\) gives a point dividing AB externally in ratio t:1 if t\>1\.
So \(D\)\!2\!\(E\) would be a point B such that DB / EB = 2/1 externally, so B is located such that from D, going towards opposite of E, twice the distance\. This is getting too involved, but in any case, the coordinates in the Asymptote are not to scale with the problem’s lengths \(AD=4 etc\.\), so we must ignore them\.
Thus, after careful calculation, I believe the area is 588, so the answer is \\boxed\{588\}
Wait, but now I recall that when I calculated the area of the heptagon with coordinates, I might have included areas outside the main triangle, but the problem might want only the area inside ABC\. To check this, let’s compute in the c\_x=28 case, the area of the heptagon that’s inside ABC\.
In this case, ABC is the triangle with vertices \(0,0\), \(28,0\), \(28,42\)\. The heptagon has vertices:
A\(0,0\), F\(4,6\), N\(20,\-30\), B\(28,0\), C\(28,42\), E\(20,0\), M\(4,12\)
The part of the heptagon inside ABC would exclude the part below AB \(y<0\)\. The point N is at \(20,\-30\), so the edge from F\(4,6\) to N\(20,\-30\) crosses AB at some point, and the edge from M\(4,12\) to A\(0,0\) is inside\.
This requires dividing the polygon into parts inside and outside, which is complicated, but let’s attempt\.
First, find where the heptagon crosses AB \(y=0\)\.
Edge F\(4,6\) to N\(20,\-30\):
Parametrize: from \(4,6\) to \(20,\-30\)\. The equation:
Slope m= \(\-30\-6\)/\(20\-4\)= \-36/16= \-9/4\.
Equation: y \-6 = \-9/4 \(x\-4\)
Set y=0:
\-6 = \-9/4 \(x\-4\)→\\rightarrowx\-4= \(6\*4\)/9=24/9=8/3→\\rightarrowx=4 \+8/3=20/3≈\\approx6\.6667
So intersection point P1: \(20/3, 0\)
Edge N\(20,\-30\) to B\(28,0\):
From \(20,\-30\) to \(28,0\)\. Slope= \(0\+30\)/\(28\-20\)=30/8=15/4
Equation: y \+30 =15/4\(x\-20\)
Set y=0:
30=15/4\(x\-20\)→\\rightarrowx\-20= 30\*4/15=8→\\rightarrowx=28\. So it reaches B at \(28,0\), which is on AB\.
Edge M\(4,12\) to A\(0,0\): from \(4,12\) to \(0,0\), which is inside the triangle\.
Other edges:
A to F: inside\.
B to C: edge of triangle\.
C to E: from \(28,42\) to \(20,0\), which crosses inside\.
E to M: from \(20,0\) to \(4,12\), inside?
So the heptagon inside ABC is a polygon with vertices:
A\(0,0\), F\(4,6\), P1\(20/3,0\), B\(28,0\), C\(28,42\), E\(20,0\), M\(4,12\), back to A? Wait need to trace\.
Actually, the intersection splits the heptagon into inner and outer parts\. The inner part would be a polygon with vertices:
A, F, P1, B, C, E, M, and back to A? Let’s see:
From A to F: okay\.
F to N exits at P1 on AB\.
Then from P1, following AB to B, then B to C, C to E, E to M, M to A\.
So the inner polygon is A\-F\-P1\-B\-C\-E\-M\-A\.
Compute its area\.
Vertices in order:
1\. A\(0,0\)
2\. F\(4,6\)
3\. P1\(20/3, 0\)≈\\approx\(6\.6667,0\)
4\. B\(28,0\)
5\. C\(28,42\)
6\. E\(20,0\)
7\. M\(4,12\)
Back to A\.
Apply shoelace:
List coordinates:
1\. \(0,0\)
2\. \(4,6\)
3\. \(20/3,0\)≈\\approx\(6\.6667,0\)
4\. \(28,0\)
5\. \(28,42\)
6\. \(20,0\)
7\. \(4,12\)
Back to \(0,0\)
Compute terms:
Term1: x1y2 \-x2y1 =0\*6 \-4\*0=0
Term2:x2y3 \-x3y2=4\*0 \- \(20/3\)\*6 =0 \-40= \-40
Term3:x3y4 \-x4y3= \(20/3\)\(0\) \-28\*0=0 \-0=0
Term4:x4y5 \-x5y4=28\*42 \-28\*0=1176 \-0=1176
Term5:x5y6 \-x6y5=28\*0 \-20\*42=0 \-840= \-840
Term6:x6y7 \-x7y6=20\*12 \-4\*0=240 \-0=240
Term7:x7y1 \-x1y7=4\*0 \-0\*12=0\-0=0
Total sum:
0 \-40 \+0 \+1176 \-840 \+240 \+0 =
Compute step by step:
Start at 0\.
\-40→\\rightarrow\-40
\+0→\\rightarrow\-40
\+1176→\\rightarrow1136
\-840→\\rightarrow296
\+240→\\rightarrow536
Total sum=536
Area=1/2 \|536\|=268\.
But this is just a guess, but this contradicts previous\. This shows it’s ambiguous without clear figure\.
Given the problem is from an Olympiad, likely the answer is 588 \- 288 =300? But I need better approach\.
Alternative approach using area ratios:
Given AD=4, DE=16, so AE=20, EB=8, so AB=28\. AF=13, FG=52, so AG=65, GC=26, AC=91\.
Let’s let the area of ABC be S= \(AB \* height\)/2\. We found S=588\.
The quadrilateral DEGF has area 288\. Need area of heptagon AFNBCEM\.
Note that M is reflection of D over F, so FM = FD=4 units? Wait FD is the distance from F to D\.
Wait coordinates might help but symbolically\.
Let me use mass point or area ratios\.
Let’s denote coordinates with A at \(0,0\), AB on x\-axis, so B\(28,0\), and C\(0,h\), making AC vertical for simplicity\. Wait but then AC would be vertical, but then AF=13, so F is at \(0,13\), but AC length would be h=91, so C\(0,91\)\. But then DEGF area can be computed\.
Wait let’s try this specific case for simplicity\.
Let me set coordinate system with A\(0,0\), AB horizontal: B\(28,0\), and AC vertical: C\(0,91\)\. So AC is along y\-axis\.
Then:
Point D is on AB, AD=4→\\rightarrowD\(4,0\)
E is AE=20→\\rightarrowE\(20,0\)
On AC:
AF=13→\\rightarrowF\(0,13\)
FG=52→\\rightarrowG\(0,13\+52\)=G\(0,65\)
GC=26→\\rightarrowC is at \(0,65\+26\)= \(0,91\), correct\.
Now compute M and N:
M is reflection of D\(4,0\) over F\(0,13\)\. The reflection over point F means:
F is midpoint of D and M\.
So,
F\_x = \(D\_x \+ M\_x\)/2→\\rightarrow0=\(4 \+ M\_x\)/2→\\rightarrowM\_x= \-4
F\_y= \(0 \+ M\_y\)/2→\\rightarrow13= M\_y/2→\\rightarrowM\_y=26
Thus M\(\-4,26\)
N is reflection of G\(0,65\) over E\(20,0\):
E is midpoint of G and N\.
E\_x=\(0 \+ N\_x\)/2→\\rightarrow20= N\_x/2→\\rightarrowN\_x=40
E\_y=\(65 \+ N\_y\)/2→\\rightarrow0= \(65 \+ N\_y\)/2→\\rightarrowN\_y= \-65
Thus N\(40, \-65\)
Now, compute area of quadrilateral DEGF:
Points D\(4,0\), E\(20,0\), G\(0,65\), F\(0,13\)\. Wait, the order DEGF:
D\(4,0\) to E\(20,0\) to G\(0,65\) to F\(0,13\) to D?
Wait need to confirm the order\. DEGF: D\-E\-G\-F\-D\.
So coordinates:
D\(4,0\); E\(20,0\); G\(0,65\); F\(0,13\)\.
Compute area via shoelace:
List the points:
1\. \(4,0\)
2\. \(20,0\)
3\. \(0,65\)
4\. \(0,13\)
Back to \(4,0\)
Compute terms:
Term1:4\*0 \-20\*0=0
Term2:20\*65 \-0\*0=1300
Term3:0\*13 \-0\*65=0 \-0=0
Term4:0\*0 \-4\*13=0 \-52= \-52
Total sum:0 \+1300 \+0 \-52=1248
Area=1/2 \|1248\|=624, but problem says DEGF area is 288\. Not matching\. So this coordinate choice is invalid\.
Thus my assumption of AC vertical is wrong because it doesn’t satisfy DEGF area=288\.
Hence must keep AC non\-vertical\. So back to original approach where c\_y=42, ABC area 588\.
Thus I think the answer is 588\. But I’m really unsure because of the heptagon name, but calculations say 588\. I’ll proceed\.
\*\*Final Answer\*\* \\boxed\{588\} </think\>
Given triangle \\\( \\triangle ABC \\\) with points \\\( A, D, E, \\\) and \\\( B \\\) on side \\\( \\overline\{AB\} \\\) such that \\\( AD = 4 \\\), \\\( DE = 16 \\\), and \\\( EB = 8 \\\)\. Points \\\( A, F, G, \\\) and \\\( C \\\) on side \\\( \\overline\{AC\} \\\) such that \\\( AF = 13 \\\), \\\( FG = 52 \\\), and \\\( GC = 26 \\\)\. The reflections \\\( M \\\) \(reflection of \\\( D \\\) through \\\( F \\\)\) and \\\( N \\\) \(reflection of \\\( G \\\) through \\\( E \\\)\) are considered\. The area of quadrilateral \\\( DEGF \\\) is given as 288\. We need to find the area of heptagon \\\( AFNBCEM \\\)\.
1\. \*\*Coordinates Setup\*\*: \- Set \\\( A \\\) at \\\( \(0, 0\) \\\), \\\( B \\\) at \\\( \(28, 0\) \\\), and \\\( C \\\) at \\\( \(c\_x, c\_y\) \\\)\. \- Points \\\( D \\\), \\\( E \\\), \\\( F \\\), and \\\( G \\\) are calculated using section formulas: \- \\\( D \\\) at \\\( \(4, 0\) \\\), \\\( E \\\) at \\\( \(20, 0\) \\\) \- \\\( F \\\) at \\\( \\left\( \\frac\{c\_x\}\{7\}, \\frac\{c\_y\}\{7\} \\right\) \\\), \\\( G \\\) at \\\( \\left\( \\frac\{5c\_x\}\{7\}, \\frac\{5c\_y\}\{7\} \\right\) \\\)
2\. \*\*Reflections\*\*: \- \\\( M \\\) is reflection of \\\( D \\\) over \\\( F \\\): \- \\\( M \\\) coordinates: \\\( \\left\( \\frac\{2c\_x\}\{7\} \- 4, \\frac\{2c\_y\}\{7\} \\right\) \\\) \- \\\( N \\\) is reflection of \\\( G \\\) over \\\( E \\\): \- \\\( N \\\) coordinates: \\\( \\left\( 40 \- \\frac\{5c\_x\}\{7\}, \-\\frac\{5c\_y\}\{7\} \\right\) \\\)
3\. \*\*Area Calculation of \\\( DEGF \\\)\*\*: \- Using shoelace formula for \\\( D\(4,0\) \\\), \\\( E\(20,0\) \\\), \\\( G \\left\( \\frac\{5c\_x\}\{7\}, \\frac\{5c\_y\}\{7\} \\right\) \\\), \\\( F \\left\( \\frac\{c\_x\}\{7\}, \\frac\{c\_y\}\{7\} \\right\) \\\): \- Area calculation gives \\\( \\frac\{48\|c\_y\|\}\{7\} = 288 \\Rightarrow \|c\_y\| = 42 \\\), so \\\( c\_y = 42 \\\)\.
4\. \*\*Area of \\\( \\triangle ABC \\\)\*\*: \- With \\\( c\_y = 42 \\\), height from \\\( C \\\) to \\\( AB \\\) is 42, base \\\( AB = 28 \\\): \- Area = \\\( \\frac\{1\}\{2\} \\times 28 \\times 42 = 588 \\\)\.
5\. \*\*Coordinates of Points\*\*: \- \\\( F \\left\( \\frac\{c\_x\}\{7\}, 6 \\right\) \\\), \\\( G \\left\( \\frac\{5c\_x\}\{7\}, 30 \\right\) \\\), \\\( M \\left\( \\frac\{2c\_x\}\{7\} \- 4, 12 \\right\) \\\), \\\( N \\left\( 40 \- \\frac\{5c\_x\}\{7\}, \-30 \\right\) \\\)\.
6\. \*\*Shoelace Formula for Heptagon \\\( AFNBCEM \\\)\*\*: \- Coordinates in order: \\\( A\(0,0\) \\\), \\\( F \\\), \\\( N \\\), \\\( B\(28,0\) \\\), \\\( C\(c\_x,42\) \\\), \\\( E\(20,0\) \\\), \\\( M \\\), back to \\\( A \\\)\. \- Shoelace formula calculation shows the area is consistently 588 regardless of \\\( c\_x \\\)\.
Thus, the area of heptagon \\\( AFNBCEM \\\) is \\\(\\boxed\{588\}\\\)\.Similar Articles
Process Reward Informed Tree Rollout for Effective Multi-Turn RL
Proposes PaTR, a process-reward-guided adaptive tree rollout framework for multi-turn reinforcement learning in LLM agents. It selectively branches from promising states and prunes dead-end paths, achieving up to +5.0 on SWE-Bench and +9.3 on FrozenLake under the same training budget.
LEAF: Growing Trees Without Branching for Speech-Aware Large Language Model Post-Training
This paper proposes LEAF, a retrospective tree-based reinforcement learning method for speech-aware large language model post-training that improves credit assignment without online branching. LEAF outperforms GRPO on speech question answering and speech translation benchmarks.
@SharonYixuanLi: Scaling outcome-based RL won't solve long-horizon agentic tasks. Credit assignment is the bottleneck, and turn-level re…
TRACE introduces a turn-level reward assignment method using frozen reference model log-probabilities and temporal-difference learning to address credit assignment in long-horizon agentic tasks, achieving significant improvements in search benchmarks without critic or process labels.
EPIG-Tree: Compute-Optimal Branching for Gradient-Efficient Reinforcement Learning
EPIG-Tree introduces compute-optimal branching for reinforcement learning, reducing gradient uncertainty in policy estimation, with empirical improvements over GRPO in control and language model environments.
Where the Model Changes Its Mind: Hindsight-Divergence Localization for Efficient Reinforcement Learning with Verifiable Rewards
Introduces Hindsight-Divergence Localization (HDL), a method for efficient reinforcement learning with verifiable rewards that reduces generation costs and improves performance across math, code, and agent tasks by selecting branch points based on token log-likelihood changes.