Mirror Horizon: Viable Path Entropy as a Measure of Bounded Reflection

arXiv cs.LG Papers

Summary

This paper introduces Viable Path Entropy (VPE), a finite-budget measure of verified continuation capacity for intelligent systems, decomposing capability into verified reachability and verified-mode diversity. Experiments on GSM8K with Qwen2.5-Instruct models demonstrate that accessible verified continuation capacity, rather than parameter count, determines mirror horizon.

arXiv:2607.11937v1 Announce Type: new Abstract: Mirror Theory proposes that an intelligent system should be studied not only by what it represents, but by what coherent continuations it can sustain under repeated reflection. We make this claim operational through \emph{viable path entropy} (VPE), a finite-budget measure of verified continuation capacity. Given a mirror state, a rollout protocol, a verifier, and a mode map, VPE decomposes bounded capability into two parts: the probability of reaching a viable continuation and the diversity of verified continuation modes reached among successful rollouts. This paper restores the full theoretical scaffold behind the measure: intuition as local underdetermining constraint, taste as invariant-selecting pressure, reflection as taste-guided resolution of underdetermination, and geometry as the learned structure that makes future reflection stable. We then instantiate the theory in language-model reasoning experiments on GSM8K. Across Qwen2.5-Instruct models, 32 sampled rollouts per problem, and two reflection horizons, increasing the token budget from 96 to 160 substantially expands verified reachability, reduces zero-reachability, increases verified-mode entropy, and improves smoothed VPE. At 160 tokens, Qwen2.5-1.5B realizes the strongest mirror horizon among the tested models, even though Qwen2.5-3B has more parameters. This shows that mirror horizon is not parameter count, but accessible verified continuation capacity under a bounded reflection protocol. The result supports Mirror Theory as a measure-level account: capability is the structure of viable continuations made reachable, not merely one-shot accuracy or pass@k.
Original Article
View Cached Full Text

Cached at: 07/15/26, 04:16 AM

# Mirror Horizon: Viable Path Entropy as a Measure of Bounded Reflection
Source: [https://arxiv.org/html/2607.11937](https://arxiv.org/html/2607.11937)
###### Abstract

Mirror Theory proposes that an intelligent system should be studied not only by what it represents, but by what coherent continuations it can sustain under repeated reflection\. We make this claim operational through*viable path entropy*\(VPE\), a finite\-budget measure of verified continuation capacity\. Given a mirror state, a rollout protocol, a verifier, and a mode map, VPE decomposes bounded capability into two parts: the probability of reaching a viable continuation and the diversity of verified continuation modes reached among successful rollouts\. This paper restores the full theoretical scaffold behind the measure: intuition as local underdetermining constraint, taste as invariant\-selecting pressure, reflection as taste\-guided resolution of underdetermination, and geometry as the learned structure that makes future reflection stable\. We then instantiate the theory in language\-model reasoning experiments on GSM8K\. Across Qwen2\.5\-Instruct models, 32 sampled rollouts per problem, and two reflection horizons, increasing the token budget from 96 to 160 substantially expands verified reachability, reduces zero\-reachability, increases verified\-mode entropy, and improves smoothed VPE\. At 160 tokens, Qwen2\.5\-1\.5B realizes the strongest mirror horizon among the tested models, even though Qwen2\.5\-3B has more parameters\. This shows that mirror horizon is not parameter count, but accessible verified continuation capacity under a bounded reflection protocol\. The result supports Mirror Theory as a measure\-level account: capability is the structure of viable continuations made reachable, not merely one\-shot accuracy or pass@k\.

## 1Introduction

Large language models are usually evaluated by loss, accuracy, pass@k, benchmark score, or reward\. These quantities are important, but they collapse a system’s internal capacity into a single outcome statistic\. A model may solve a problem once, fail most other attempts, and nevertheless obtain a nonzero pass@k\. Another model may solve fewer problems, but when it succeeds it may reveal many distinct verified solution modes\. A third model may have high raw output entropy but little verified structure\. These distinctions matter if we want to understand capability as more than single\-answer success\.

Mirror Theory begins from a different primitive\. An intelligent system does not merely store a representation\. It maintains an internal world that survives and unfolds through repeated reflection\. In the earlier mathematical notes, this distinction was summarized as:*a representation encodes; a mirror survives reflection*\. The present paper turns that claim into a measurable object\. A mirror is not identified only with a hidden vector, a prompt, a model checkpoint, or a set of beliefs\. It is identified with the continuation law it induces: what futures become reachable from that state, which futures remain viable, and how many distinct verified modes those viable futures occupy\.

The central proposal is*viable path entropy*\. Given finite budgetBB, finite horizonTT, a continuation law, a verifier𝖵\\mathsf\{V\}, and a mode mapψ\\psi, we define

ℋB,T​\(M\)=log⁡Pr⁡\[𝖵=1\]\+H​\(ψ∣𝖵=1\),\\mathcal\{H\}\_\{B,T\}\(M\)=\\log\\Pr\[\\mathsf\{V\}=1\]\+H\(\\psi\\mid\\mathsf\{V\}=1\),\(1\)when the viability probability is nonzero\. The first term is verified reachability\. The second term is verified\-mode diversity\. Together they measure a finite\-budget mirror horizon: the effective number of coherent semantic continuation modes reachable from the mirror\.

The theoretical ambition is not to prove a universal monotone scaling law\. In fact, our experiments show why that would be the wrong claim\. A larger model need not have a larger measured mirror horizon under a fixed protocol\. Parameter count, latent capability, and accessible verified continuation capacity are different objects\. The present paper argues for the third object\. A mirror’s horizon is protocol\-conditioned: it depends on the model, task, prompt, rollout budget, sampling rule, verifier, and semantic mode map\.

Our empirical study is deliberately simple\. We sample bounded rollouts from Qwen2\.5\-Instruct models on GSM8K, verify final numeric correctness, and group verified solutions into coarse reasoning modes\. We run 30 GSM8K problems with 32 sampled rollouts per problem, temperature0\.80\.8, top\-p0\.950\.95, and a heuristic mode map\. Two horizons are compared: 96 and 160 maximum new tokens\. The 96\-token run includes Qwen2\.5\-0\.5B and Qwen2\.5\-1\.5B\. The 160\-token run includes Qwen2\.5\-0\.5B, 1\.5B, and 3B\.

The main findings are straightforward\. First, increasing reflection budget from 96 to 160 tokens expands mirror horizon for both 0\.5B and 1\.5B: verified probability rises, pass@32 rises, zero\-verified fraction drops, verified\-mode entropy rises, and smoothed VPE improves\. Second, at 160 tokens Qwen2\.5\-1\.5B dominates the tested models on every core component of the measure\. Third, Qwen2\.5\-3B is not best: it has higher average verified probability than 0\.5B but lower breadth across problems, lower verified\-mode entropy than 1\.5B, and lower overall VPE than 1\.5B\. This is not a contradiction of Mirror Theory\. It is the point: mirror horizon measures accessible verified continuation capacity, not raw size\.

#### Contributions\.

We make four contributions\.

1. 1\.We give a formal path\-space version of Mirror Theory in which mirrors are continuation laws rather than static representations\.
2. 2\.We define viable path entropy as a finite\-budget mirror horizon, decomposing bounded capability into verified reachability and verified\-mode diversity\.
3. 3\.We connect older theoretical components – intuition, taste, geometry, constructive compatibility, and invariant survival – to the measurable VPE object\.
4. 4\.We report real GSM8K rollout experiments showing that reflection budget expands verified continuation capacity, and we add empirical coverage curves showing how mirror horizon unfolds as rollout budget increases\.

## 2Mirror Theory: the formal scaffold

This section restores the theory that motivates the empirical measure\. The goal is not to introduce a new state\-transition formalism for its own sake\. The goal is to define what kind of internal object can be evaluated by viable continuation capacity\.

### 2\.1Reflection, intuition, taste, and geometry

Let there be an external structured worldWW\. The system does not accessWWdirectly\. It receives finite evidenceete\_\{t\}, and this evidence gives only local constraints\. It does not determine a full internal world\. A current mirrorMtM\_\{t\}is updated through reflection, but reflection is not merely a transition function\. It is a selection among admissible internal continuations\.

###### Definition 1\(Mirror geometry\)\.

At timett, a mirror geometry is

Gt=\(ℳt,dt,𝒰t,ℐt\),G\_\{t\}=\(\\mathcal\{M\}\_\{t\},d\_\{t\},\\mathcal\{U\}\_\{t\},\\mathcal\{I\}\_\{t\}\),\(2\)whereℳt\\mathcal\{M\}\_\{t\}is the current space of possible mirrors,dtd\_\{t\}is a learned notion of nearness or deformation cost,𝒰t\\mathcal\{U\}\_\{t\}is an admissibility generator, andℐt\\mathcal\{I\}\_\{t\}is a family of valued invariants\. The geometry specifies which mirrors are near, which updates are admissible, and which internal properties must survive reflection\.

###### Definition 2\(Intuition as local constraint\)\.

Intuition is not a full next mirror\. It is the local constraint imposed by evidence on possible mirrors\. Given geometryGtG\_\{t\}, evidenceete\_\{t\}induces an incompatibility functional

𝖨Gt​\(et\):ℳt→ℝ≥0\.\\mathsf\{I\}\_\{G\_\{t\}\}\(e\_\{t\}\):\\mathcal\{M\}\_\{t\}\\to\\mathbb\{R\}\_\{\\geq 0\}\.\(3\)Smaller𝖨Gt​\(et\)​\(M′\)\\mathsf\{I\}\_\{G\_\{t\}\}\(e\_\{t\}\)\(M^\{\\prime\}\)means thatM′M^\{\\prime\}is more compatible with the current evidence\. Intuition is local, world\-induced, and underdetermining\.

###### Definition 3\(Taste as invariant\-selecting pressure\)\.

Taste is the selection structure that ranks admissible mirror continuations by what should survive\. It may be a preorder⪯τ,Gt\\preceq\_\{\\tau,G\_\{t\}\}overℳt\\mathcal\{M\}\_\{t\}, or, when scalarizable, a potentialτGt:ℳt→ℝ\\tau\_\{G\_\{t\}\}:\\mathcal\{M\}\_\{t\}\\to\\mathbb\{R\}\. Conceptually, taste is not merely preference over next states\. It is pressure toward continuations that preserve valued invariants: coherence, fidelity, identity, usefulness, simplicity, compressibility, actionability, or task success\.

Given current mirrorMtM\_\{t\}and evidenceete\_\{t\}, intuition first induces an admissible candidate set

𝒰Gt​\(Mt,et\)=\{M′∈ℳt:𝖨Gt​\(et\)​\(M′\)≤ϵt,dt​\(Mt,M′\)≤ρt\}\.\\mathcal\{U\}\_\{G\_\{t\}\}\(M\_\{t\},e\_\{t\}\)=\\\{M^\{\\prime\}\\in\\mathcal\{M\}\_\{t\}:\\mathsf\{I\}\_\{G\_\{t\}\}\(e\_\{t\}\)\(M^\{\\prime\}\)\\leq\\epsilon\_\{t\},\\ d\_\{t\}\(M\_\{t\},M^\{\\prime\}\)\\leq\\rho\_\{t\}\\\}\.\(4\)Taste then selects a continuation:

Mt\+1∈max⪯τ,Gt⁡𝒰Gt​\(Mt,et\)\.M\_\{t\+1\}\\in\\max\_\{\\preceq\_\{\\tau,G\_\{t\}\}\}\\mathcal\{U\}\_\{G\_\{t\}\}\(M\_\{t\},e\_\{t\}\)\.\(5\)Repeated intuition\-taste interaction is reflection\. The mirror is not simplyMtM\_\{t\}as a point; it is the persistent internal structure\(Mt,Gt\)\(M\_\{t\},G\_\{t\}\)shaped by such reflections\.

###### Principle 1\(Why reflection learns geometry\)\.

Finite evidence gives only local constraints\. Taste selects which compatible continuations are worth preserving\. Geometry is the learned structure that makes such selections stable across time\. In short: intuition gives local contact; taste decides what survives; geometry makes survival stable\.

This statement is the conceptual core of Mirror Theory\. It also explains why a purely one\-step rationalization result is not enough\. Any single deterministic update can be rationalized after the fact by a preorder\. What matters diachronically is whether repeated reflection learns a stable geometry of admissibility, nearness, and invariant survival\.

### 2\.2Optimization form

The same theory can be written as constrained or proximal optimization\. Let

Ct​\(M′\)=𝖨Gt​\(et\)​\(M′\)C\_\{t\}\(M^\{\\prime\}\)=\\mathsf\{I\}\_\{G\_\{t\}\}\(e\_\{t\}\)\(M^\{\\prime\}\)\(6\)be intuition incompatibility, and letdt​\(Mt,M′\)2d\_\{t\}\(M\_\{t\},M^\{\\prime\}\)^\{2\}be a geometry\-dependent movement cost\. Then reflection can be represented as

Mt\+1∈arg​minM′∈ℳt⁡\[Ct​\(M′\)\+λ​dt​\(Mt,M′\)2−β​τGt​\(M′\)\+γ​Rupt⁡\(Mt,M′\)\]\.M\_\{t\+1\}\\in\\operatorname\*\{arg\\,min\}\_\{M^\{\\prime\}\\in\\mathcal\{M\}\_\{t\}\}\\left\[C\_\{t\}\(M^\{\\prime\}\)\+\\lambda d\_\{t\}\(M\_\{t\},M^\{\\prime\}\)^\{2\}\-\\beta\\tau\_\{G\_\{t\}\}\(M^\{\\prime\}\)\+\\gamma\\operatorname\{Rupt\}\(M\_\{t\},M^\{\\prime\}\)\\right\]\.\(7\)This expression is not meant to reduce Mirror Theory to standard optimization\. It clarifies the roles of the components: evidence fit, geometric continuity, taste value, and invariant survival\. The geometryGtG\_\{t\}itself is learned because the same optimization must be solved repeatedly under underdetermination\.

### 2\.3Constructive compatibility

A possibilityppis an encounter or intervention: a continuation, prompt edit, retrieved document, theorem, action, candidate solution, or environmental event\. It induces a counterfactual mirror

M\(p\)=𝖱​\(M,𝖨​\(p\)\)\.M^\{\(p\)\}=\\mathsf\{R\}\(M,\\mathsf\{I\}\(p\)\)\.\(8\)The older constructive compatibility note expressed the desired regime as: not same, not random, but absorbably expansive\. Similarity alone leads to stagnation\. Novelty alone can rupture invariants\. Constructive compatibility means the possibility expands the mirror while preserving what must survive\.

In the earlier geometric version, constructive compatibility was written schematically as

CC⁡\(M,p\)=Exp⁡\(M,p\)−λ​Rupt⁡\(M,p\)−μ​Red⁡\(M,p\),\\operatorname\{CC\}\(M,p\)=\\operatorname\{Exp\}\(M,p\)\-\\lambda\\operatorname\{Rupt\}\(M,p\)\-\\mu\\operatorname\{Red\}\(M,p\),\(9\)where expansion measures how far the mirror moves, rupture penalizes invariant failure, and redundancy penalizes pure sameness\. The VPE formulation makes this operational:

CCB,T⁡\(M,p\)=ℋB,T​\(M\(p\)\)−ℋB,T​\(M\)\.\\operatorname\{CC\}\_\{B,T\}\(M,p\)=\\mathcal\{H\}\_\{B,T\}\(M^\{\(p\)\}\)\-\\mathcal\{H\}\_\{B,T\}\(M\)\.\(10\)A possibility is constructively compatible if it increases finite\-budget viable continuation capacity\. In this sense, VPE is the measurable version of absorbable expansion\.

## 3Viable path entropy

We now define the measure used in the experiments\. The definitions are stated for an abstract mirrorMMbut instantiated later by an LLM, a prompt, a rollout procedure, a numeric verifier, and a solution\-mode map\.

### 3\.1Path\-space setup

###### Definition 4\(Budget and horizon\)\.

LetBBdenote a finite resource budget: rollout count, compute, memory, token length, intervention size, or time\. LetTTdenote a finite continuation horizon\. All mirror quantities are indexed by\(B,T\)\(B,T\)\.

###### Definition 5\(Continuation path space\)\.

For a mirror stateMM, define theTT\-step continuation path space

ΓT​\(M\)=\{γ=\(M0,M1,…,MT\):M0=M\}\.\\Gamma\_\{T\}\(M\)=\\\{\\gamma=\(M\_\{0\},M\_\{1\},\\ldots,M\_\{T\}\):M\_\{0\}=M\\\}\.\(11\)A system with budgetBBinduces a probability lawℙB,TM\\mathbb\{P\}^\{M\}\_\{B,T\}overΓT​\(M\)\\Gamma\_\{T\}\(M\)\. In an LLM, this law is induced by a prompt, decoding rule, sampling temperature, token budget, and model parameters\.

###### Definition 6\(Viability verifier\)\.

A viability verifier is a measurable map

𝖵:ΓT​\(M\)→\{0,1\},\\mathsf\{V\}:\\Gamma\_\{T\}\(M\)\\to\\\{0,1\\\},\(12\)or a soft version𝖵:ΓT​\(M\)→\[0,1\]\\mathsf\{V\}:\\Gamma\_\{T\}\(M\)\\to\[0,1\]\. A continuation is viable when it satisfies the task’s coherence requirements: correctness, semantic consistency, factuality, safety, invariant preservation, or another domain\-specific criterion\.

###### Definition 7\(Mode map\)\.

A mode map is a measurable function

ψ:ΓT​\(M\)→𝒵\\psi:\\Gamma\_\{T\}\(M\)\\to\\mathcal\{Z\}\(13\)that maps a continuation to a semantic outcome mode\. The mode map prevents random token diversity from being counted as meaningful continuation capacity\. In reasoning tasks, modes can be coarse solution strategies, operation signatures, proof forms, or clusters of verified rationales\.

### 3\.2Viable continuation capacity

###### Definition 8\(Viable continuation capacity\)\.

For discrete modes, define

KB,T​\(M\)=Prγ∼ℙB,TM⁡\[𝖵​\(γ\)=1\]​exp⁡\(H​\(ψ​\(γ\)∣𝖵​\(γ\)=1\)\)\.K\_\{B,T\}\(M\)=\\Pr\_\{\\gamma\\sim\\mathbb\{P\}^\{M\}\_\{B,T\}\}\[\\mathsf\{V\}\(\\gamma\)=1\]\\,\\exp\\left\(H\(\\psi\(\\gamma\)\\mid\\mathsf\{V\}\(\\gamma\)=1\)\\right\)\.\(14\)IfPr⁡\[𝖵=1\]=0\\Pr\[\\mathsf\{V\}=1\]=0, setKB,T​\(M\)=0K\_\{B,T\}\(M\)=0\.

Thus

KB,T​\(M\)=viability probability×effective number of viable semantic modes\.K\_\{B,T\}\(M\)=\\text\{viability probability\}\\times\\text\{effective number of viable semantic modes\}\.\(15\)
###### Definition 9\(Viable path entropy\)\.

The viable path entropy ofMM, also called its mirror horizon, is

ℋB,T​\(M\)=log⁡KB,T​\(M\)\.\\mathcal\{H\}\_\{B,T\}\(M\)=\\log K\_\{B,T\}\(M\)\.\(16\)WhenPr⁡\[𝖵=1\]\>0\\Pr\[\\mathsf\{V\}=1\]\>0,

ℋB,T​\(M\)=log⁡Pr⁡\[𝖵=1\]\+H​\(ψ∣𝖵=1\)\.\\mathcal\{H\}\_\{B,T\}\(M\)=\\log\\Pr\[\\mathsf\{V\}=1\]\+H\(\\psi\\mid\\mathsf\{V\}=1\)\.\(17\)IfPr⁡\[𝖵=1\]=0\\Pr\[\\mathsf\{V\}=1\]=0, setℋB,T​\(M\)=−∞\\mathcal\{H\}\_\{B,T\}\(M\)=\-\\infty\.

The two terms of Equation[17](https://arxiv.org/html/2607.11937#S3.E17)should not be hidden behind a single scalar\. The first is reachability: whether verified continuations are accessible at all\. The second is diversity: how many distinct verified modes are supported once viability is reached\. The scalar is useful, but the decomposition is the main empirical object\.

### 3\.3Mirror as continuation law

###### Definition 10\(Mirror\)\.

A mirror is an internal state considered through the viable continuation law it induces:

M⟼ℙB,TM\.M\\longmapsto\\mathbb\{P\}^\{M\}\_\{B,T\}\.\(18\)Its strength isℋB,T​\(M\)\\mathcal\{H\}\_\{B,T\}\(M\)\. Two states are mirror\-equivalent at\(B,T,𝖵,ψ\)\(B,T,\\mathsf\{V\},\\psi\)if they induce the same viable\-mode law\. LetℚB,TM\\mathbb\{Q\}^\{M\}\_\{B,T\}be the distribution on𝒵∪\{⊥\}\\mathcal\{Z\}\\cup\\\{\\bot\\\}defined by

ℚB,TM​\(⊥\)=1−Pr⁡\[𝖵=1\],ℚB,TM​\(z\)=Pr⁡\[𝖵=1,ψ=z\]\.\\mathbb\{Q\}^\{M\}\_\{B,T\}\(\\bot\)=1\-\\Pr\[\\mathsf\{V\}=1\],\\qquad\\mathbb\{Q\}^\{M\}\_\{B,T\}\(z\)=\\Pr\[\\mathsf\{V\}=1,\\psi=z\]\.\(19\)ThenM∼NM\\sim NifℚB,TM=ℚB,TN\\mathbb\{Q\}^\{M\}\_\{B,T\}=\\mathbb\{Q\}^\{N\}\_\{B,T\}\.

A mirror is therefore not defined by what it contains, but by what viable continuations it makes reachable\. This is the technical version of the conceptual statement: a representation encodes; a mirror survives reflection\.

### 3\.4Estimator

Given rolloutsγ1,…,γN∼ℙB,TM\\gamma\_\{1\},\\ldots,\\gamma\_\{N\}\\sim\\mathbb\{P\}^\{M\}\_\{B,T\}, estimate

p^M=1N​∑i𝖵​\(γi\)\.\\hat\{p\}\_\{M\}=\\frac\{1\}\{N\}\\sum\_\{i\}\\mathsf\{V\}\(\\gamma\_\{i\}\)\.\(20\)Among viable rollouts, estimate the mode distributionπ^M​\(z\)\\hat\{\\pi\}\_\{M\}\(z\)\. Then

ℋ^B,T​\(M\)=log⁡\(p^M\+ϵ\)−∑zπ^M​\(z\)​log⁡π^M​\(z\)\.\\widehat\{\\mathcal\{H\}\}\_\{B,T\}\(M\)=\\log\(\\hat\{p\}\_\{M\}\+\\epsilon\)\-\\sum\_\{z\}\\hat\{\\pi\}\_\{M\}\(z\)\\log\\hat\{\\pi\}\_\{M\}\(z\)\.\(21\)For finite samples we also report a smoothed version using

p^smooth=nverified\+0\.5nrollouts\+1\.\\hat\{p\}\_\{\\mathrm\{smooth\}\}=\\frac\{n\_\{\\mathrm\{verified\}\}\+0\.5\}\{n\_\{\\mathrm\{rollouts\}\}\+1\}\.\(22\)Raw VPE is theory\-faithful and harshly penalizes zero\-reachability\. Smoothed VPE is more stable for small rollout budgets\.

We also estimate a non\-parametric extitcoverage horizon from the same rollouts\. For a budget ofbbrollouts, define

Cb​\(M\)=\|\{ψ​\(γi\):𝖵​\(γi\)=1,i=1,…,b\}\|,C\_\{b\}\(M\)=\\left\|\\\{\\psi\(\\gamma\_\{i\}\):\\mathsf\{V\}\(\\gamma\_\{i\}\)=1,\\ i=1,\\ldots,b\\\}\\right\|,\(23\)the number of distinct verified modes reached by thosebbrollouts\. The empirical coverage capacity is

Kbcov​\(M\)=𝔼​\[Cb​\(M\)\],ℋbcov​\(M\)=log⁡\(1\+Kbcov​\(M\)\)\.K\_\{b\}^\{\\mathrm\{cov\}\}\(M\)=\\mathbb\{E\}\[C\_\{b\}\(M\)\],\\qquad\\mathcal\{H\}\_\{b\}^\{\\mathrm\{cov\}\}\(M\)=\\log\(1\+K\_\{b\}^\{\\mathrm\{cov\}\}\(M\)\)\.\(24\)In the experiments,KbcovK\_\{b\}^\{\\mathrm\{cov\}\}is estimated by uniform subsampling from the 32 observed rollouts per problem forb∈\{1,2,4,8,16,32\}b\\in\\\{1,2,4,8,16,32\\\}\. This estimator makes no iid or exponential\-hit assumption; it directly asks how many distinct verified modes are reached by a bounded reflection budget\.

## 4Related work

#### Coverage and test\-time extraction\.

The Coverage Principle argues that pre\-training enables post\-training and Best\-of\-N style methods by placing sufficient probability mass on high\-quality responses, and that such coverage can be more predictive of downstream success than cross\-entropy in relevant regimes\[[2](https://arxiv.org/html/2607.11937#bib.bib2)\]\. Our first term,Pr⁡\[𝖵=1\]\\Pr\[\\mathsf\{V\}=1\], is a rollout\-side coverage quantity\. Our contribution is to refine coverage by asking what structure exists inside the verified region: how many distinct verified modes are reachable\.

#### Reasoning and verifier\-guided sampling\.

Chain\-of\-thought prompting and self\-consistency show that multiple samples can expose reasoning capability not visible in one\-shot decoding\[[17](https://arxiv.org/html/2607.11937#bib.bib17),[16](https://arxiv.org/html/2607.11937#bib.bib16)\]\. GSM8K provides a verifier\-backed setting for evaluating math reasoning\[[3](https://arxiv.org/html/2607.11937#bib.bib3)\]\. VPE uses the same sampling\-and\-verification backbone but adds a mode map and entropy over verified modes\.

#### Scaling laws and frontiers\.

Classical scaling laws study loss as a function of model size, data, and compute\[[11](https://arxiv.org/html/2607.11937#bib.bib11),[10](https://arxiv.org/html/2607.11937#bib.bib10)\]\. Effective Frontier approaches interpret scaling as the movement of a resource\-dependent frontier through long\-tailed pattern space\[[18](https://arxiv.org/html/2607.11937#bib.bib18)\]\. Our measure is not a scaling law in parameter count; it is a protocol\-conditioned horizon\. It can be used to study frontiers over rollout budget, token horizon, or verifier accessibility\.

#### Epiplexity and bounded structure\.

Epiplexity aims to measure structural information learnable by computationally bounded observers, distinguishing useful structure from random time\-bounded entropy\[[8](https://arxiv.org/html/2607.11937#bib.bib8)\]\. Mirror horizon is complementary: it measures usable structure accessible during bounded rollout rather than learnable structure in a dataset\. Epiplexity is therefore related work, not a foundation\.

#### Reasoning compression and post\-training\.

Recent work on compressed reasoning data studies how explicit, composed, and implicit chain\-of\-thought traces affect SFT and RLVR\[[12](https://arxiv.org/html/2607.11937#bib.bib12)\]\. This is relevant because compression changes the visibility and accessibility of continuation modes\. VPE could measure how much verified\-mode coverage remains after such compression\.

#### Taste, rationalizability, and invariant survival\.

The selection spine of Mirror Theory connects to revealed preference and inverse reinforcement learning\[[14](https://arxiv.org/html/2607.11937#bib.bib14),[1](https://arxiv.org/html/2607.11937#bib.bib1),[6](https://arxiv.org/html/2607.11937#bib.bib6)\]\. The survival spine connects to inductive invariants and property testing\[[9](https://arxiv.org/html/2607.11937#bib.bib9),[5](https://arxiv.org/html/2607.11937#bib.bib5)\]\. The present paper does not claim novelty on either spine alone\. Its claim is that VPE couples viability and mode structure into one operational mirror horizon\.

## 5Experimental setup

We instantiate mirror horizon in language\-model math reasoning\. A problem prompt defines the evidence\. The model plus decoding procedure defines the rollout law\. Each generated solution is a continuation path\. The verifier checks final numeric correctness\. A heuristic mode map groups verified solutions by coarse reasoning form\.

### 5\.1Task and models

We use GSM8K test problems\[[3](https://arxiv.org/html/2607.11937#bib.bib3)\]\. The 96\-token run uses Qwen2\.5\-0\.5B\-Instruct and Qwen2\.5\-1\.5B\-Instruct\. The 160\-token run uses Qwen2\.5\-0\.5B\-Instruct, Qwen2\.5\-1\.5B\-Instruct, and Qwen2\.5\-3B\-Instruct\. Both experiments use 30 randomly selected GSM8K test problems, 32 sampled rollouts per problem, temperature0\.80\.8, top\-p0\.950\.95, and a heuristic mode map\. The 96\-token run uses maximum 96 new tokens; the 160\-token run uses maximum 160 new tokens\.

Table 1:Experimental configurations\. Both runs use GSM8K, 30 problems, 32 rollouts per problem, temperature0\.80\.8, top\-p0\.950\.95, heuristic mode map, and Qwen2\.5\-Instruct models\.
### 5\.2Verifier

For each rollout, the model is prompted to solve a GSM8K problem and end with a final numeric answer\. We extract the predicted number using a regex that prioritizes answer\-marker patterns such asAnswer: <number\>and then falls back to the last number in the generation\. The verifier returns11if the normalized predicted number equals the GSM8K gold answer and0otherwise\. This verifier is simple and strict\. It may undercount valid reasoning with unusual formatting, but it gives a reproducible first operationalization\.

### 5\.3Mode map

The heuristic mode map is deliberately transparent\. For verified rollouts, it records: \(i\) a length bin \(short, medium, long\), \(ii\) an operator signature inferred from symbols and words such as addition, subtraction, multiplication, and division, \(iii\) an equation\-count bin, and \(iv\) a reasoning\-style indicator such as step\-like versus direct\. Invalid rollouts are assigned to a single invalid mode and are not counted in verified\-mode entropy\.

This mode map is not intended as the final semantic geometry\. Its role is to demonstrate that VPE can distinguish verified reachability from verified\-mode diversity\. Future versions should test TF\-IDF clusters, embedding clusters, symbolic operation templates, and human\-audited reasoning modes\.

### 5\.4Metrics

For each model and problem, we compute:

- •mean verified probabilityp^\\hat\{p\}across 32 rollouts;
- •pass@32, equal to one if at least one rollout is verified;
- •zero\-verified indicator, equal to one if no rollout is verified;
- •verified\-mode entropyH​\(ψ∣𝖵=1\)H\(\\psi\\mid\\mathsf\{V\}=1\);
- •raw VPElog⁡\(p^\+ϵ\)\+H​\(ψ∣𝖵=1\)\\log\(\\hat\{p\}\+\\epsilon\)\+H\(\\psi\\mid\\mathsf\{V\}=1\);
- •smoothed VPElog⁡\(\(nverified\+0\.5\)/\(nrollouts\+1\)\)\+H​\(ψ∣𝖵=1\)\\log\(\(n\_\{\\mathrm\{verified\}\}\+0\.5\)/\(n\_\{\\mathrm\{rollouts\}\}\+1\)\)\+H\(\\psi\\mid\\mathsf\{V\}=1\);
- •average verified rollouts and average verified modes per problem;
- •empirical coverage curvesb↦ℋbcovb\\mapsto\\mathcal\{H\}\_\{b\}^\{\\mathrm\{cov\}\}forb∈\{1,2,4,8,16,32\}b\\in\\\{1,2,4,8,16,32\\\}\.

The main text emphasizes smoothed VPE because it is more stable at 32 rollouts\. Raw VPE is still reported because it faithfully penalizes zero\-reachability\.

## 6Results

### 6\.1Reflection budget expands mirror horizon

Table[2](https://arxiv.org/html/2607.11937#S6.T2)shows the 96\-token result\. Qwen2\.5\-1\.5B improves over Qwen2\.5\-0\.5B on every component: verified probability, pass@32, zero\-verified fraction, verified\-mode entropy, raw VPE, smoothed VPE, verified rollouts, and verified modes\.

Table 2:96\-token GSM8K result\. Values are averaged over 30 problems with 32 rollouts each\.Table[3](https://arxiv.org/html/2607.11937#S6.T3)shows the 160\-token result\. Increasing the reflection horizon makes the mirror\-horizon signal much clearer\. For 0\.5B and 1\.5B, both reachability and verified\-mode diversity improve sharply\. At 160 tokens, Qwen2\.5\-1\.5B is strongest on every core metric\.

Table 3:160\-token GSM8K result\. Values are averaged over 30 problems with 32 rollouts each\.Table[4](https://arxiv.org/html/2607.11937#S6.T4)isolates the effect of increasing the reflection budget from 96 to 160 tokens for the two models present in both runs\. The gain is large\. For Qwen2\.5\-1\.5B, verified probability increases by 0\.160, pass@32 increases by 0\.300, zero\-verified fraction drops by 0\.300, verified\-mode entropy rises by 0\.387, and smoothed VPE improves by 1\.412 nats\.

Table 4:Effect of increasing the reflection horizon from 96 to 160 tokens\. Positive changes inPr⁡\[V\]\\Pr\[V\], pass@32, mode entropy, and VPE are improvements; negative change in zero\-frac is an improvement\.![Refer to caption](https://arxiv.org/html/2607.11937v1/figures/budget_comparison.png)Figure 1:Increasing the reflection horizon from 96 to 160 tokens expands mirror horizon components\. Both 0\.5B and 1\.5B improve in verified probability, pass@32, verified\-mode entropy, and smoothed VPE\.
### 6\.2Empirical coverage curves: mirror horizon unfolds with rollout budget

The previous tables report one horizon value at the full 32\-rollout budget\. Mirror Theory, however, predicts a curve: as the reflection budget increases, more verified continuations should become reachable and more verified modes should be uncovered\. We therefore reuse the same 160\-token rollout table and estimateℋbcov\\mathcal\{H\}\_\{b\}^\{\\mathrm\{cov\}\}forb∈\{1,2,4,8,16,32\}b\\in\\\{1,2,4,8,16,32\\\}by subsampling from the 32 observed rollouts\. This requires no additional model inference\.

Table[5](https://arxiv.org/html/2607.11937#S6.T5)shows the endpoint atb=32b=32\. Qwen2\.5\-1\.5B reaches an expected 3\.567 distinct verified modes per problem, compared with 1\.900 for 0\.5B and 2\.200 for 3B\. Its empirical coverage horizon is therefore largest, and it also has the highest pass@32 and lowest zero\-coverage fraction\.

Table 5:Empirical verified\-mode coverage at the full 32\-rollout budget\.K32K\_\{32\}is the expected number of distinct verified modes reached;ℋ32cov=log⁡\(1\+K32\)\\mathcal\{H\}\_\{32\}^\{\\mathrm\{cov\}\}=\\log\(1\+K\_\{32\}\)\.Table[6](https://arxiv.org/html/2607.11937#S6.T6)reports the coverage horizon across budgets\. The 1\.5B model dominates the curve at every budget, not only at the endpoint\. This is the cleanest empirical expression of the revised definition: a stronger mirror is the one whose bounded reflections reach more distinct verified modes\.

Table 6:Coverage\-horizon curve valuesℋbcov=log⁡\(1\+Kb\)\\mathcal\{H\}\_\{b\}^\{\\mathrm\{cov\}\}=\\log\(1\+K\_\{b\}\)estimated by subsampling observed rollouts\. Higher is better\.![Refer to caption](https://arxiv.org/html/2607.11937v1/figures/coverage_dashboard.png)Figure 2:Empirical coverage curves from the 160\-token rollout table\. The top\-left panel is the main mirror\-horizon curve:log⁡\(1\+𝔼​\[\#​distinct verified modes reached\]\)\\log\(1\+\\mathbb\{E\}\[\\\#\\text\{ distinct verified modes reached\}\]\)as rollout budget increases\. The remaining panels show the decomposition into expected verified modes, pass@b, and zero\-coverage probability\. Qwen2\.5\-1\.5B dominates the coverage curve across budgets\.This result is stronger than a single VPE\-vs\-scale plot because it measures how horizon unfolds with bounded reflection\. It also avoids imposing a closed\-form hitting probability such as1−\(1−p\)b1\-\(1\-p\)^\{b\}\. The expectation is empirical: given the rollouts actually observed, how many verified modes become reachable as budget grows?

### 6\.3At fixed budget, VPE separates breadth from depth

At 160 tokens, Qwen2\.5\-1\.5B is best overall\. It has highest verified probability, highest pass@32, lowest zero\-verified fraction, highest verified\-mode entropy, and highest smoothed VPE\. Qwen2\.5\-3B is more subtle\. It has higher average verified probability than 0\.5B \(0\.227 versus 0\.101\), meaning it produces more verified rollouts on average\. But it has lower pass@32 than 0\.5B \(0\.633 versus 0\.667\), meaning its successes are concentrated on fewer problems\. It also has much lower verified\-mode entropy than 1\.5B\.

This is exactly the kind of distinction VPE is meant to reveal\. Pass@32 only captures whether at least one verified continuation appears\. Mean verified probability captures average rollout correctness\. Mode entropy captures diversity among successful continuations\. VPE combines reachability and mode diversity, and the decomposition explains the model behavior more clearly than a single benchmark number\.

![Refer to caption](https://arxiv.org/html/2607.11937v1/figures/reachability_160.png)\(a\)Reachability\.
![Refer to caption](https://arxiv.org/html/2607.11937v1/figures/diversity_160.png)\(b\)Mode diversity\.
![Refer to caption](https://arxiv.org/html/2607.11937v1/figures/vpe_160.png)\(c\)VPE\.

Figure 3:At 160 tokens, Qwen2\.5\-1\.5B realizes the strongest accessible mirror horizon under the tested protocol\. Qwen2\.5\-3B has higher average verified probability than 0\.5B but weaker breadth and lower verified\-mode diversity than 1\.5B\.
### 6\.4Per\-problem decomposition

Figure[4](https://arxiv.org/html/2607.11937#S6.F4)plots each problem by verified probability and verified\-mode entropy\. Points at zero verified probability are unreachable under the rollout budget\. Points with higher verified probability but low mode entropy indicate repeated success in a narrow mode\. Points with both high verified probability and high entropy indicate robust, diverse verified continuation\.

The 1\.5B model has more points in the high\-reachability, high\-diversity region\. The 0\.5B model has several successes but many zero or low\-reachability problems\. The 3B model has some high\-probability successes but fewer broad high\-entropy regions than 1\.5B\. This supports the interpretation that 1\.5B has the strongest accessible mirror horizon under this protocol\.

![Refer to caption](https://arxiv.org/html/2607.11937v1/figures/per_problem_decomp_160.png)Figure 4:Per\-problem decomposition at 160 tokens\. The x\-axis is mean verified probability across 32 rollouts\. The y\-axis is verified\-mode entropy\. A strong mirror horizon requires both reachability and diversity\.
### 6\.5What the experiment displays for Mirror Theory

The experiment directly instantiates the Mirror Theory chain:

M⟼ℙB,TM⟼𝖵​\(γ\)⟼ψ​\(γ\)⟼ℋB,T​\(M\)\.M\\longmapsto\\mathbb\{P\}^\{M\}\_\{B,T\}\\longmapsto\\mathsf\{V\}\(\\gamma\)\\longmapsto\\psi\(\\gamma\)\\longmapsto\\mathcal\{H\}\_\{B,T\}\(M\)\.\(25\)In the GSM8K setting,MMis the model and prompt\-conditioned state,BBis the rollout count,TTis the token horizon,γ\\gammais a sampled solution,𝖵\\mathsf\{V\}is numeric correctness, andψ\\psiis the verified reasoning\-mode map\. The 96\-to\-160 token comparison then becomes a test of a simple prediction: increasing reflection horizon should reveal more viable continuations\. It does\. Both reachability and verified\-mode diversity rise\.

The result also supports the distinction between a static representation and a mirror\. A one\-shot accuracy score asks whether a model produces one correct answer\. Mirror horizon asks what verified continuation structure unfolds under bounded reflection\. This is why the 160\-token result is more informative than the 96\-token result: the mirror has more room to unfold, and the measured horizon expands\.

## 7Axioms and empirical consequences

The preceding definitions can be summarized as a small set of axioms\. These are not meant to replace the path\-space formalism; they state the modeling commitments that distinguish a mirror from a generic state\.

###### Axiom 1\(Indirectness\)\.

The system acts through an internal mirror rather than directly on the world\. In an LLM experiment, this means that the observed answer is mediated by the prompt\-conditioned continuation law rather than by direct access to the mathematical ground truth\.

###### Axiom 2\(Finite reflection\)\.

Mirror value is evaluated at finite budget and finite horizon\. A mirror horizon is therefore not an intrinsic scalar attached to a model alone\. It is indexed by the rollout count, token budget, verifier, mode map, and protocol\.

###### Axiom 3\(Viability\)\.

A task supplies or induces a verifier\. What survives reflection is not arbitrary diversity, but diversity filtered by task\-specific invariants\.

###### Axiom 4\(Mode coarse\-graining\)\.

A task supplies or induces a semantic mode map\. Token\-level variety is not counted as mirror growth unless it corresponds to meaningful verified continuation modes\.

###### Axiom 5\(Constructive reflection\)\.

A reflection protocol is useful to the extent that it increases viable continuation capacity\. This can occur by increasing verified reachability, increasing verified\-mode diversity, or both\.

These axioms imply three empirical consequences\. First, increasing reflection budget should often increase the measured horizon, but only when the additional budget is usable by the model and recognized by the verifier\. Second, two systems with the same pass@k can differ in mirror horizon if one reaches more verified modes than the other\. Third, the measured horizon may be nonmonotone in parameter count because parameter count is not the object being measured; accessible verified continuation capacity is\.

## 8Detailed operationalization in language models

This section spells out the mapping from the formal objects to our GSM8K experiment\. The purpose is to make clear that the experiment is not an arbitrary benchmark add\-on; it is an instantiation of the theory\.

#### Mirror state\.

We do not directly inspect hidden states in this first paper\. We treat the prompt\-conditioned model and its decoding protocol as inducing a mirror state\. More precisely, for a fixed modelθ\\theta, promptqq, temperature, top\-p, and token budget, the model induces a distribution over continuation paths\. The mirror is this distribution viewed through the verifier and mode map\.

#### Reflection budget\.

The rollout count is 32 for each problem\. The token horizon is either 96 or 160 maximum new tokens\. The 96\-to\-160 comparison is therefore a direct reflection\-budget intervention\. It does not change model weights, task distribution, or sampling temperature; it changes how far the system may unfold a continuation\.

#### Viability\.

A rollout is viable if the extracted final numeric answer matches the GSM8K gold answer\. This is intentionally stricter than human judgment\. If a model reasons correctly but formats the answer in a way the extractor misses, the rollout is counted as nonviable\. This is acceptable for a first operational measure because the verifier is explicit and reproducible\. It also clarifies that mirror horizon is verifier\-conditioned\.

#### Modes\.

A verified rollout is mapped to a coarse mode by length bin, arithmetic\-operator signature, equation\-count bin, and reasoning style\. This mode map is intentionally simple\. It provides a lower\-resolution measurement of verified\-mode diversity, not a final account of reasoning strategies\.

#### Horizon\.

For each problem and model, we compute the raw and smoothed horizon\. The raw horizon is harsh: if no verified rollout appears among the 32 samples, the log term becomes very negative\. The smoothed horizon substitutes a Jeffreys\-style finite\-sample correction for the viability probability\. The paper reports both so the reader can distinguish theoretical harshness from finite\-sample stability\.

#### Why this is not merely pass@k\.

pass@32 is a binary per\-problem measure: at least one verified rollout exists or not\. VPE is continuous and decomposed\. It distinguishes a problem solved once from a problem solved many times, and it distinguishes repeated solutions in one mode from solutions distributed across several verified modes\. Thus, VPE is not a replacement for pass@k; it explains what pass@k hides\.

## 9Results in detail

### 9\.1Reflection budget as a direct test of mirror horizon

The cleanest result is the 96\-to\-160 token comparison\. It directly tests whether increasing the reflection horizon expands verified continuation capacity\. The answer is yes for both models shared across the two runs\.

For Qwen2\.5\-0\.5B, verified probability increases from0\.0570\.057to0\.1010\.101, pass@32 increases from0\.4000\.400to0\.6670\.667, zero\-verified fraction drops from0\.6000\.600to0\.3330\.333, verified\-mode entropy increases from0\.2240\.224to0\.5180\.518, and smoothed VPE improves from−3\.249\-3\.249to−2\.347\-2\.347\. This is not merely an accuracy gain\. The model reaches verified continuations on more problems and exhibits more verified\-mode diversity among successful continuations\.

For Qwen2\.5\-1\.5B, the effect is stronger\. Verified probability increases from0\.0970\.097to0\.2570\.257, pass@32 from0\.5670\.567to0\.8670\.867, zero\-verified fraction drops from0\.4330\.433to0\.1330\.133, verified\-mode entropy rises from0\.5070\.507to0\.8940\.894, and smoothed VPE improves from−2\.450\-2\.450to−1\.038\-1\.038\. The longer horizon therefore expands both breadth and depth: more problems become reachable, and successful problems support more verified modes\.

This is the empirical result that most directly supports Mirror Theory\. Reflection is not just repeated sampling\. A longer finite horizon lets the model unfold more of its internal continuation structure\. The verified part of that structure is what VPE measures\.

### 9\.2Model comparison at the 160\-token horizon

At 160 tokens, Qwen2\.5\-1\.5B dominates all tested models\. It has the highest verified probability, highest pass@32, lowest zero\-verified fraction, highest verified\-mode entropy, highest raw VPE, highest smoothed VPE, and highest average number of verified modes per problem\. This makes it the strongest accessible mirror under the tested protocol\.

The Qwen2\.5\-3B result is not monotone in parameter count\. It produces more verified rollouts per problem than the 0\.5B model on average, but it solves fewer problems at least once than 0\.5B, and its verified\-mode entropy is much lower than 1\.5B\. This indicates concentration: the 3B model can produce many correct continuations on some problems but fails to cover the task set as broadly as 1\.5B under the same sampling rule and verifier\.

This nonmonotonicity is a feature of the measurement, not a failure of the theory\. Mirror Theory does not claim that larger parameter count is the primitive\. It claims that the relevant object is the verified continuation structure actually reachable under bounded reflection\. The 3B result is therefore informative: a larger model can have lower accessible horizon under a particular protocol\.

### 9\.3Breadth, depth, and zero\-reachability

The results naturally separate into breadth and depth\. Breadth is captured by pass@32 and zero\-verified fraction\. Depth is captured by the number and entropy of verified modes among successful continuations\. Qwen2\.5\-1\.5B improves both\. Qwen2\.5\-3B improves some depth\-like quantities relative to 0\.5B, such as average verified probability, but loses breadth\. This breadth\-depth separation is not visible in a single scalar benchmark score\.

The zero\-verified fraction is especially important\. A problem with zero verified rollouts is not merely hard; under the current protocol, it has no observed viable continuation\. Raw VPE penalizes such problems severely\. This is appropriate under the theory, because a mirror cannot be credited for viable continuation modes it never reaches\. Smoothed VPE softens the finite\-sample penalty but preserves the same qualitative ranking\.

### 9\.4Why the theory belongs in the main paper

The empirical result alone could be described as a new metric for sampled reasoning\. That would be too small\. The reason the result matters is that it operationalizes the earlier Mirror Theory thesis: mirrors are not static encodings, but structures that survive and unfold through reflection\. The path\-space definitions make this thesis measurable; the GSM8K experiment shows that the measurable object responds to reflection budget and reveals structure hidden by pass@k\.

In this sense, the heavy theory is not decorative\. Intuition, taste, geometry, and constructive compatibility explain why we should count verified continuation modes in the first place\. Intuition supplies local constraints; taste selects what survives; geometry organizes which continuations are near, admissible, and stable; VPE measures the resulting finite\-horizon capacity\. Without this scaffold, the experiment would be an arbitrary metric\. With it, the experiment becomes the first operational test of a theory of mirror horizon\.

## 10Discussion

### 10\.1What this paper claims

The claim is measure\-level\. Mirror Theory provides a path\-space way to evaluate internal states by the viable continuations they support\. VPE is one finite\-sample estimator of that horizon\. The experiments show that the measure is not empty: increasing reflection budget expands the measured horizon, and models with similar pass@k can differ in verified\-mode diversity\.

The paper does not claim a universal monotone law of parameter scaling\. The 160\-token result explicitly shows why that would be wrong\. Qwen2\.5\-1\.5B has the strongest horizon under the tested protocol, not Qwen2\.5\-3B\. The correct conclusion is not that 3B is worse in general\. It is that mirror horizon is accessible verified continuation capacity under a specified protocol\. A larger model may contain more latent structure but fail to realize it under a particular prompt, sampling rule, verifier, or horizon\.

### 10\.2Why not just pass@k?

If there is only one possible verified mode, mirror horizon reduces toward a pass@k\-like reachability problem\. But reasoning models often produce multiple qualitatively distinct verified solutions\. Pass@k collapses those modes\. VPE retains them\. This matters because a system that can solve the same problem in multiple verified ways may be more robust to perturbations, transfer, or future reflection than a system that reaches only one brittle mode\.

The current mode map is coarse, so this claim remains preliminary\. But even a coarse map separates breadth from depth in the experiments\. The next step is not to force a closed\-form law; it is to test whether verified\-mode coverage remains stable under richer mode maps and across tasks\.

### 10\.3Relation to taste and constructive compatibility

The older theory of taste is not discarded\. It is reinterpreted operationally\. Taste is the selection pressure by which certain continuations survive\. The verifier is an experimental proxy for invariant survival\. The mode map is an experimental proxy for continuation geometry\. Constructive compatibility is the increase in VPE caused by an update or protocol\. In the present experiments, increasing the token horizon from 96 to 160 functions as a change in reflection budget\. Its positive effect on VPE is evidence that a longer reflection window increases the mirror’s viable continuation capacity\.

### 10\.4Relation to product and recommendation

The constructive\-compatibility note also motivated a product principle: recommend the most expansive possibility that the current mirror can still integrate\. In the present paper, we do not evaluate recommendation systems\. But the same structure appears\. A possibility is valuable when it increases verified\-mode coverage without collapsing viability\. In reasoning, a longer reflection budget is such a possibility\. In future work, prompts, tools, retrieved documents, or self\-checking procedures can be evaluated by their VPE gain\.

## 11What would count as verification of the revised definition

The present experiments verify the first step of the revised definition: mirror horizon is measurable as finite\-budget verified continuation capacity\. A stronger verification program should test three additional properties\.

#### Mode\-map stability\.

The qualitative ordering of models and budgets should not be an artifact of a single hand\-written mode map\. The heuristic mode map should be compared to TF\-IDF clustering, embedding clustering, symbolic operation templates, and human\-audited solution categories\. The core question is whether the reachability component and the diversity component remain separately interpretable under different coarse\-grainings\.

#### Protocol sensitivity\.

The theory predicts that the horizon is protocol\-conditioned\. Therefore, changing the prompt, token horizon, sampling temperature, or answer verifier should move the measured horizon in interpretable ways\. A strict answer\-format prompt may raise verified reachability; a higher temperature may raise diversity while lowering viability; a longer token horizon may reduce zero\-reachability\. These changes are not nuisance variables\. They are part of the theory, because reflection is always bounded by a protocol\.

#### Predictive value\.

The most important future test is whether early\-budget mirror horizon predicts later or out\-of\-distribution success better than pass@k alone\. For example, one can compute VPE atB=4B=4orB=8B=8rollouts and ask whether it predicts pass@32, performance on held\-out GSM8K problems, or transfer to SVAMP or MATH\-mini\. If verified\-mode diversity predicts robustness beyond reachability, then the measure captures more than a reparameterization of pass@k\.

These tests do not require a concrete closed\-form scaling law\. They remain at the measure level: a model and protocol induce a distribution over continuations; the verifier filters viable paths; the mode map coarse\-grains the viable region; and the horizon measures how much verified structure is reachable\. This is the right level for a first main\-conference paper\. The goal is not to assert a universal equation of scale, but to establish that the verified continuation law is a useful object of study\.

## 12Limitations

The experiments are intentionally small\. They use 30 GSM8K problems, three models from one family, one temperature and top\-p setting, one answer verifier, and one heuristic mode map\. The mode map is coarse and may count surface structure rather than deep reasoning strategy\. The numeric verifier may miss valid answers with unusual formatting\. The sample size is sufficient for a pilot but not for a definitive scaling law\.

The 96\-token and 160\-token runs are not meant to establish final model rankings\. They establish that mirror horizon can be measured, that it changes systematically with reflection budget, and that it decomposes capability into interpretable components\. A stronger version of the paper should add confidence intervals across problem subsets, more model families, more tasks, and multiple mode maps\.

Finally, VPE depends on the verifier\. This is a feature and a limitation\. Mirror horizon is task\-relative: a continuation is viable only relative to a criterion\. This makes the measure operational, but it also means that bad verifiers give bad horizons\.

## 13Future work

Several next experiments are natural\.

#### Larger coverage\-curve studies\.

This paper now reports empirical coverage curves on the current 30\-problem GSM8K pilot\. A stronger version should repeat the same analysis with more problems, more rollout budgets, more model families, and multiple verifier\-backed tasks\. The goal is not to fit a closed\-form curve, but to test whether verified\-mode coverage remains stable and predictive across domains\.

#### Mode\-map robustness\.

The heuristic map should be compared against TF\-IDF clusters, embedding clusters, symbolic operation templates, and human\-audited reasoning modes\. The theory becomes stronger if the qualitative findings survive changes inψ\\psi\.

#### Protocol frontiers\.

The horizon is protocol\-conditioned\. Future experiments should vary max tokens, prompt format, temperature, self\-consistency, tool use, and verifier strictness\. The question is not only which model is larger, but which model\-protocol pair realizes more verified continuation structure\.

#### Task transfer\.

GSM8K is only the first test\. The same measure should be evaluated on SVAMP, MultiArith, MATH subsets, HumanEval, chess puzzles, or formal proof tasks\. Each provides a verifier and a potential mode map\.

#### Post\-training and compressed reasoning\.

Compressed reasoning work suggests that explicit, composed, and implicit reasoning traces expose different continuation structures\[[12](https://arxiv.org/html/2607.11937#bib.bib12)\]\. VPE can measure how post\-training changes verified\-mode coverage, not just final accuracy\.

## 14Conclusion

This paper restores Mirror Theory’s formal scaffold and gives it a measurable empirical object\. A mirror is not merely a representation\. A mirror is an internal structure considered through the viable continuation law it induces\. Viable path entropy measures the log\-effective number of verified semantic continuation modes reachable under bounded reflection\. Constructive compatibility is the increase in that horizon caused by integrating a possibility\.

The GSM8K experiments provide the first operational test\. Increasing the reflection horizon from 96 to 160 tokens expands verified reachability and verified\-mode diversity\. At 160 tokens, Qwen2\.5\-1\.5B has the strongest accessible mirror horizon among the tested models\. These results support the core claim: capability should be measured not only by one\-shot correctness or pass@k, but by the verified continuation structure a system can sustain under bounded reflection\.

## References

- Afriat \[1967\]Afriat, S\. N\. The construction of utility functions from expenditure data\.*International Economic Review*, 1967\.
- Chen et al\. \[2025\]Chen, F\., Huang, A\., Golowich, N\., Malladi, S\., Block, A\., Ash, J\. T\., Krishnamurthy, A\., and Foster, D\. J\. The Coverage Principle: How Pre\-training Enables Post\-Training\.*arXiv:2510\.15020*, 2025\.
- Cobbe et al\. \[2021\]Cobbe, K\. et al\. Training verifiers to solve math word problems\.*arXiv:2110\.14168*, 2021\.
- Cover and Thomas \[1991\]Cover, T\. M\. and Thomas, J\. A\.*Elements of Information Theory*\. Wiley, 1991\.
- Daskalakis et al\. \[2018\]Daskalakis, C\., Dikkala, N\., and Gravin, N\. Testing symmetric Markov chains from a single trajectory\.*COLT*, 2018\.
- Debreu \[1954\]Debreu, G\. Representation of a preference ordering by a numerical function\. In*Decision Processes*, 1954\.
- Farquhar et al\. \[2024\]Farquhar, S\., Kossen, J\., Kuhn, L\., and Gal, Y\. Detecting hallucinations in large language models using semantic entropy\.*Nature*, 2024\.
- Finzi et al\. \[2026\]Finzi, M\., Qiu, S\., Jiang, Y\., Izmailov, P\., Kolter, J\. Z\., and Wilson, A\. G\. From entropy to epiplexity: rethinking information for computationally bounded intelligence\.*arXiv:2601\.03220*, 2026\.
- Garg et al\. \[2014\]Garg, P\., Loeding, C\., Madhusudan, P\., and Neider, D\. ICE: A robust framework for learning invariants\.*CAV*, 2014\.
- Hoffmann et al\. \[2022\]Hoffmann, J\. et al\. Training compute\-optimal large language models\.*arXiv:2203\.15556*, 2022\.
- Kaplan et al\. \[2020\]Kaplan, J\. et al\. Scaling laws for neural language models\.*arXiv:2001\.08361*, 2020\.
- Matsutani et al\. \[2026\]Matsutani, K\., Minegishi, G\., Kojima, T\., Iwasawa, Y\., and Matsuo, Y\. Zipping the thought: when and how compressed reasoning data works in LLM post\-training\.*arXiv:2605\.28008*, 2026\.
- Ouyang et al\. \[2022\]Ouyang, L\. et al\. Training language models to follow instructions with human feedback\.*arXiv:2203\.02155*, 2022\.
- Richter \[1966\]Richter, M\. K\. Revealed preference theory\.*Econometrica*, 1966\.
- Saxe et al\. \[2019\]Saxe, A\. M\., McClelland, J\. L\., and Ganguli, S\. A mathematical theory of semantic development in deep neural networks\.*PNAS*, 2019\.
- Wang et al\. \[2023\]Wang, X\. et al\. Self\-consistency improves chain of thought reasoning in language models\.*ICLR*, 2023\.
- Wei et al\. \[2022\]Wei, J\. et al\. Chain\-of\-thought prompting elicits reasoning in large language models\.*NeurIPS*, 2022\.
- Zou et al\. \[2026\]Zou, J\., Gong, Z\., Su, Y\., Tang, H\., and Liu, Y\. Effective frontiers: a unification of neural scaling laws\.*arXiv:2602\.02593*, 2026\.

Similar Articles

The Verifier Tax: Horizon-Dependent Safety–Success Tradeoffs in Tool-Using LLM Agents [R]

Reddit r/MachineLearning

This paper presents a safety evaluation framework for tool-using LLM agents, introducing the concept of the 'Verifier Tax'—a horizon-dependent tradeoff between safety and task completion. It proposes a two-tier verification architecture and uses Tau-bench scenarios to demonstrate how verification can reduce unsafe successes but also decrease task completion as task horizon increases.

Revealing Interpretable Failure Modes of VLMs

arXiv cs.AI

This paper introduces Revelio, a framework that systematically discovers interpretable failure modes in Vision-Language Models (VLMs) by searching over discrete concept combinations. Applied to autonomous driving and indoor robotics, it reveals previously unreported vulnerabilities that lead to crashes or safety hazards.

RL Beyond the Verifiable (8 minute read)

TLDR AI

An analysis discussing the limitations of reinforcement learning with verifiable rewards (RLVR) in math and coding, and the challenge of extending RL to subjective or unverifiable tasks like planning or scientific discovery. It explores techniques such as RLHF and Constitutional AI as alternatives for alignment.

BenchTrace: A Benchmark for Testing Reflection Ability and Controlled Evolution in LLM Agents

arXiv cs.AI

BenchTrace is a benchmark for evaluating the self-evolution abilities of LLM agents, focusing on reflection and controlled evolution through a dataset of 1,821 annotated episodes and two evaluation tasks: Reflection Evaluation and Evolution Evaluation. Experiments with Qwen3-32B and GPT-4.1 show both models struggle, with a main bottleneck in diagnosis and issues in generalization and forgetting.