NeuralCert: certified computational discovery of extremal mathematical constructions

arXiv cs.LG Papers

Summary

NeuralCert presents a discovery-to-certification framework using neural networks to find and verify extremal mathematical constructions, enabling exact and independently verifiable proofs.

arXiv:2609.30296v1 Announce Type: new Abstract: Neural networks are becoming popular in solving mathematical problems, but stochastic models do not provide mathematical exactness by themselves. This study introduces a discovery-to-certification framework in which high-dimensional variational trial functions are learned in a compact separable representation, spectrally diagnosed and pruned, and then certified exactly through multimodular evaluation. Exact certification makes the numerical proofs fully explicit and independently verifiable. This framework can be run on a standard personal computer. Across three extremal problems, we show that neural optimization can contribute to rigorous mathematics in three distinct ways: by discovering improved constructions, by exposing empirical invariants that lead to proofs, and by revealing optimization barriers whose geometry motivates new analytic or numerical representations. More broadly, these results suggest a path toward AI-assisted mathematics in which flexible computational discovery and exact certification become complementary components of a single rigorous workflow.
Original Article
View Cached Full Text

Cached at: 09/29/26, 09:34 AM

# NeuralCert: certified computational discovery of extremal mathematical constructions
Source: [https://arxiv.org/html/2609.30296](https://arxiv.org/html/2609.30296)
Mark Patrick Roeling††thanks:Corresponding author:[mp\.roeling@mindef\.nl](mailto:[email protected])\. Code and certified data:[https://github\.com/mproeling/neuralcert](https://github.com/mproeling/neuralcert)

###### Abstract

Neural networks are becoming popular in solving mathematical problems, but stochastic models do not provide mathematical exactness by themselves\. This study introduces a discovery\-to\-certification framework in which high\-dimensional variational trial functions are learned in a compact separable representation, spectrally diagnosed and pruned, and then certified exactly through multimodular evaluation\. Exact certification makes the numerical proofs fully explicit and independently verifiable\. This framework can be run on a standard personal computer\.

Across three extremal problems, we show that neural optimization can contribute to rigorous mathematics in three distinct ways: by discovering improved constructions, by exposing empirical invariants that lead to proofs, and by revealing optimization barriers whose geometry motivates new analytic or numerical representations\.

More broadly, these results suggest a path toward AI\-assisted mathematics in which flexible computational discovery and exact certification become complementary components of a single rigorous workflow\.

Machine learning can contribute to mathematical discovery by identifying latent relations between mathematical objects, guiding search over vast combinatorial or algebraic spaces, and generating candidate constructions or algorithms\. Davies*et al\.*used learned predictors and attribution methods to expose latent mathematical structure and guide conjecture formation\[[6](https://arxiv.org/html/2609.30296#bib.bib29)\]\. AlphaTensor showed that reinforcement learning can search large spaces of tensor decompositions and recover exact matrix\-multiplication algorithms, including previously unknown constructions\[[8](https://arxiv.org/html/2609.30296#bib.bib30)\]\. FunSearch coupled neural generation to an executable evaluator, demonstrating that iterative proposal, evaluation and selection can produce new mathematical constructions\[[26](https://arxiv.org/html/2609.30296#bib.bib31)\]\. Also, neural function approximators can transform a mixed discrete–continuous geometric problem into a differentiable optimization problem\[[20](https://arxiv.org/html/2609.30296#bib.bib32)\]\.

Despite these advances, a methodological gap persists for problems in which the unknown object is a function or extremizer in a large or infinite\-dimensional space\. Such problems are commonly reduced computationally by choosing a finite ansatz,

f⁡\(x\)=∑j=1maj​ϕj​\(x\),f\(x\)=\\sum\_\{j=1\}^\{m\}a\_\{j\}\\phi\_\{j\}\(x\),\(1\)and optimizing over the coefficientsaja\_\{j\}\. While this makes the problem tractable, it also imposes an a priori structural restriction,

f⋆≈span⁡\{ϕ1,…,ϕm\},f^\{\\star\}\\approx\\operatorname\{span\}\\\{\\phi\_\{1\},\\ldots,\\phi\_\{m\}\\\},\(2\)so that improvements in optimization accuracy cannot remove approximation error induced by an inadequate representation\. However, neural parameterizations offer a complementary strategy: instead of fixing the shape of a candidate solution in advance, they provide flexible differentiable families in which large\-scale functional search can be conducted\. For example, in geometric colouring a mixed discrete–continuous search space was replaced by a probabilistic differentiable representation amenable to gradient optimization\[[20](https://arxiv.org/html/2609.30296#bib.bib32)\]\.

Here we introduce NeuralCert, a discovery\-to\-certification framework that deliberately separates the representation used for computational search from the representation used for mathematical proof\. Flexible neural parameterizations are used to learn one\-dimensional factors within a separable trial family, without fixing the polynomial dictionary subsequently used for exact certification\. Discovered candidates are subsequently compressed into explicit mathematical representations from which rigorous certificates can be constructed and independently verified\. Here, certification means recording a claimed bound together with the exact data needed to recompute it independently, so that the inequality can be confirmed without trusting, or re\-running, any part of the discovery pipeline\. Across three extremal problems, we show that this separation enables neural optimization to contribute to rigorous mathematics through three distinct modes: discovering improved constructions, exposing structure that leads to analytic results, and identifying optimization barriers that motivate alternative representations\.

Since a floating\-point neural candidate does not constitute a proof, the discovery candidate is reduced to an independently checkable object\[[8](https://arxiv.org/html/2609.30296#bib.bib30),[26](https://arxiv.org/html/2609.30296#bib.bib31)\]such that verification establishes the mathematical claim\. This idea builds on previous work where neural search produced tensor decompositions which could be checked algebraically\[[8](https://arxiv.org/html/2609.30296#bib.bib30)\], and candidate programs judged by an explicit executable evaluator\[[26](https://arxiv.org/html/2609.30296#bib.bib31)\]\. These examples suggest that discovery and verification impose fundamentally different computational requirements\. Discovery benefits from \(over\)parameterization, approximation and exploration, whereas verification benefits from low\-dimensional structure, exactness and controlled numerical error\.

Our central methodological hypothesis is not that neural optimization should replace mathematical proof, but that the representation best suited for discovery need not be the representation best suited for proof\. A flexible computational model can search a substantially larger functional space than is convenient for exact symbolic treatment, after which the useful structure can be transferred into an explicit mathematical object\. Conversely, repeated failure or instability of such a search can reveal invariants or unsuitable coordinates and thereby guide analytic reformulation\[[6](https://arxiv.org/html/2609.30296#bib.bib29)\]\. NeuralCert is designed around this asymmetry: discovery is deliberately flexible and expendable, whereas every reported mathematical claim must be supported by a representation that can be reconstructed and verified independently\.

## NeuralCert overview

NeuralCert is a Python framework for neural discovery of candidate functions followed by independent certification and verification\. The architecture is deliberately separated into three components; discovery, certification and verification, which do not share numerical evaluation routines\. Discovery uses flexible neural parameterizations to explore problem\-specific function spaces and exports candidate functions and metadata, typically as NumPy\.npzfiles\. Certification is independent of the neural model: no network parameters, optimizer state or training data are required\. Instead, the candidate is reconstructed in a problem\-specific mathematical representation and converted into an explicit certificate\. Verification is implemented as a standalone minimal code path that ingests only this certificate, typically as\.json, and independently reconstructs the inequalities underlying the claimed result\. Thus, neural optimization serves only as a proposal mechanism; final mathematical claims depend exclusively on independently reconstructed and verifiable proof objects\.

## Discovery and certification of sieve constants improve prime gap bounds

The Maynard–Tao sieve reduces bounded gaps between primes to lower bounds on a variational constantMkM\_\{k\}, which we approached through two independent computational routes\. The first uses neural search over separable functional families, followed by reconstruction in a polynomial representation and rigorous Rayleigh–Ritz certification\. At smallkk, this approach recovers known extremal behaviour \(M5≥2\.0071443298M\_\{5\}\\geq 2\.0071443298andM20≥3\.12755795M\_\{20\}\\geq 3\.12755795\); atk=25k=25it improves the Polymath8b reference \(certified3\.32215113\.3221511versus3\.32214263\.3221426\)\. Thus, high\-quality extremizers can be discovered without prescribing the polynomial ansatz subsequently used for certification \(Supplementary Tables[7](https://arxiv.org/html/2609.30296#S11.T7)and[8](https://arxiv.org/html/2609.30296#S11.T8)\)\. Neural discovery further benefits from overparameterization: a redundant channel representation supplies extra variational degrees of freedom during optimization, after which unnecessary channels are removed by backward elimination with re\-optimization at each step, subject to a cumulative relative Rayleigh\-quotient loss of at mostτprune=10−7\\tau\_\{\\mathrm\{prune\}\}=10^\{\-7\}\.

At largekk, an early neural run neark=3600k=3600produced a numerically stable Rayleigh value above88that remained below the Cauchy–Schwarz ceiling and was stable under grid refinement\. Exact reconstruction nevertheless failed: the optimizer had collapsed onto a sharply localized single channel whose discovery score was inflated by an inconsistent quadrature frame\. The discovered profile approaches1/\(log⁡k⋅t\)1/\(\\log k\\cdot t\), a form the polynomial channels cannot represent across the increasingly flat Rayleigh landscape \(Supplementary Methods[5\.18](https://arxiv.org/html/2609.30296#S5.SS18)\)\. The artifact survived the available heuristic gates because its value did not cross the upper bound; only independent reconstruction of the candidate for certification exposed the inconsistency, which motivated NeuralCert’s scalar cross\-checks and support\-aware thresholds\.

To access much largerkk, we developed a second computational route using the rational profileg⁡\(t\)=1/\(c\+\(k−1\)​t\)g\(t\)=1/\(c\+\(k\-1\)t\), already employed in Polymath8b \(Theorem 6\.7\)\. Our contribution is a direct, rigorously certified Fourier\-domain evaluation of the associated product trial function on the simplex\. This makes it possible to optimize and certify the construction at a computational cost nearly independent ofkk, obtain improved explicit prime\-gap bounds, and analyse its asymptotic deficit through a stable\-law limit\. This yields

M3655≥8\.000064869,M208910≥12\.00000199,M11655069≥16\.0000003924,M644589002≥20\.0000009143,M\_\{3655\}\\geq 8\.000064869,\\quad M\_\{208910\}\\geq 12\.00000199,\\quad M\_\{11655069\}\\geq 16\.0000003924,\\quad M\_\{644589002\}\\geq 20\.0000009143,for33,44,55and66primes in a tuple, together with the uniform boundMk≥log⁡k−0\.307M\_\{k\}\\geq\\log k\-0\.307for100≤k≤1\.88×109100\\leq k\\leq 1\.88\\times 10^\{9\}\. These results require only the Bombieri–Vinogradov theorem, assuming neither the Elliott–Halberstam conjecture nor Deligne\-type distributional input\.

Combining these thresholds with explicit admissible tuples yields

H2≤33 118,H3≤2 718 108,H4≤214 099 720,H5≤14 541 349 288,H\_\{2\}\\leq 33\\,118,\\quad H\_\{3\}\\leq 2\\,718\\,108,\\quad H\_\{4\}\\leq 214\\,099\\,720,\\quad H\_\{5\}\\leq 14\\,541\\,349\\,288,where the tuple fork=3655k=3655is taken from the Engelsma–Sutherland database\[[7](https://arxiv.org/html/2609.30296#bib.bib6)\], the tuple fork=208910k=208910was constructed by a hybrid shifted\-Schinzel/greedy sieve, and the two largest use the firstkkprimes exceedingkk; all four are checked by a standalone verifier that shares no code with the construction \(Supplementary Methods\)\. Measured against the corresponding records established without the Elliott–Halberstam conjecture and without Deligne’s theorems\[polymath\_wiki,[23](https://arxiv.org/html/2609.30296#bib.bib27)\], these improve the published bounds by factors of14\.314\.3,11\.911\.9,9\.59\.5and8\.68\.6respectively\. They also improve on the stronger records that do invoke Deligne\-type input\[[28](https://arxiv.org/html/2609.30296#bib.bib28)\], by factors of12\.012\.0,9\.09\.0,6\.56\.5and5\.35\.3\(Figure[1](https://arxiv.org/html/2609.30296#Sx2.F1)and Supplementary Data\)\. This distinction matters because, as noted in\[[29](https://arxiv.org/html/2609.30296#bib.bib7)\], the previously best\-known upper boundsHmH\_\{m\}form≥2m\\geq 2relies on equidistribution estimates of Zhang type; the bounds above are obtained without them\.

An asymptotic analysis of the single\-channel rational family identifies its limiting convolution profile with a spectrally positive11\-stable law and provesMk≥log⁡k−C∗−o⁡\(1\)M\_\{k\}\\geq\\log k\-C\_\{\\ast\}\-o\(1\), whereC∗=infa∈ℝC⁡\(a\)C\_\{\\ast\}=\\inf\_\{a\\in\\mathbb\{R\}\}C\(a\); numerical evaluation suggestsC∗≈0\.3343C\_\{\\ast\}\\approx 0\.3343\(Supplementary Methods\)\. This explains why the rank\-one structure remains asymptotically competitive, attaining the optimallog⁡k\\log kgrowth rate with only a constant\-order deficit, without implying global optimality of rank one\. Consequently

0≤lim infk→∞\(log⁡k−Mk\)≤C∗,0\\leq\\liminf\_\{k\\to\\infty\}\(\\log k\-M\_\{k\}\)\\leq C\_\{\\ast\},where the lower endpoint follows from the classical boundMk≤kk−1​log⁡kM\_\{k\}\\leq\\frac\{k\}\{k\-1\}\\log k\. The finite\-range constant0\.3070\.307is smaller thanC∗C\_\{\\ast\}because, within this rational family, the finite\-range deficit is numerically consistent withC∗−1\.047/log⁡kC\_\{\\ast\}\-1\.047/\\log k; the directly evaluated deficit remains below0\.290\.29throughout the certified range\.

Despite sharing no numerical evaluation, the two computational pipelines agree throughout their overlapping range\. While the finite\-kkresults remain below Bogaert’s independently computed \(Krylov subspace\) values fork≠25k\\neq 25,k≤30k\\leq 30, indicating that the Krylov method is powerful for modestkk, NeuralCert improves31≤k≤10031\\leq k\\leq 100\. Every reported finite\-kkinequality is accompanied by a machine\-checkable certificate that can be independently verified\.

![Refer to caption](https://arxiv.org/html/2609.30296v1/Section1/images/Rplot-Asymptote-allmethods.jpeg)Figure 1:Certified lower bounds for the Maynard variational constant across computational regimes\. Certified values obtained with Bogaert’s Krylov method are compared with NeuralCert polynomial constructions for the vanilla problem \(ε=0\\varepsilon=0\) and enlarged support \(ε=1/25\\varepsilon=1/25\), together with the certified rank\-one rational family\. The polynomial construction first exceeds the Krylov bounds atk=25k=25, but loses efficiency askkincreases, whereas the rank\-one rational representation becomes competitive and provides the continuation to largerkk\. The Cauchy–Schwarz bound is shown as a reference\. Lines connecting discrete certified values are guides to the eye and do not represent interpolated certificates\. Beyondk=500k=500, the rank\-one rational family is the only representation that remained reliably optimizable and certifiable in the regime studied here\.We also implemented theε\\varepsilon\-enlarged variant proposed in Polymath8b\. Atk=49k=49, optimization within our certified trial family saturates at3\.988673\.98867\. The certified value is a lower bound, so this does not establishM49≤4M\_\{49\}\\leq 4; it shows that enlarging the sieve weight alone, at Bombieri–Vinogradov level, did not reach the threshold in our search\. Concurrent work is consistent with this reading:H1≤240H\_\{1\}\\leq 240has since been established by combining the Bombieri–Vinogradov support with equidistribution estimates for smooth moduli\[[29](https://arxiv.org/html/2609.30296#bib.bib7)\], and subsequent refinements of that construction giveH1≤212H\_\{1\}\\leq 212\[[2](https://arxiv.org/html/2609.30296#bib.bib9)\]andH1≤186H\_\{1\}\\leq 186\[[21](https://arxiv.org/html/2609.30296#bib.bib10)\]; all three use distributional input beyond Bombieri–Vinogradov\. The additional ingredient is therefore on the arithmetic side rather than the variational one, which is where our saturation result locates it\. We report the first certifiedε\\varepsilon\-values atk=49,52k=49,52and abovek=54k=54\.

## Higher\-order Delsarte discovery exposes a tensorization obstruction

Higher\-order Delsarte hierarchies strengthen the classical LP for linear codes and show substantial finite\-length gains, but asymptotic progress requires an explicit dual family rather than additional numerical optimization\[[18](https://arxiv.org/html/2609.30296#bib.bib14),[5](https://arxiv.org/html/2609.30296#bib.bib13)\]\. We therefore treated the higher\-order dual as a discovery problem, initially asking whether a low\-complexity correction to the lifted level\-one solution could improve the MRRW bound\. Extending the symmetry\-reduced hierarchy tor=3r=3confirmed that higher levels contain genuine information, including the level\-3 boundV3​\(12,5\)≤16V\_\{3\}\(12,5\)\\leq 16, against the level\-2 valueV2​\(12,5\)=24\.260255​…V\_\{2\}\(12,5\)=24\.260255\\ldotslocated numerically\. However, comparison with the explicit lift showed that the source of the improvement changes with hierarchy level: atr=2r=2it can reside entirely in the partial\-Fourier constraints, whereas atr=3r=3the full\-transform rows already improve on the lift\. A single learned correction was therefore not a stable asymptotic target\.

We redirected discovery to the spectral construction of\[[5](https://arxiv.org/html/2609.30296#bib.bib13)\]where the blocklength dependence can be removed analytically\. For a normalized configurationGGon𝔽2ℓ\\mathbb\{F\}\_\{2\}^\{\\ell\}, the relevant objective isJℓ​\(G\)=𝖧⁡\(G\)/ℓJ\_\{\\ell\}\(G\)=\\mathsf\{H\}\(G\)/\\ell\. Although we initially intended to parameterizeGGneurally, symmetry and analytic reduction collapsed the problem faster than a generic network became useful: forℓ≤4\\ell\\leq 4the largest search space has only1515free probability coordinates and no remaining dependence onnn\. We therefore optimized the reduced configuration directly\. Across2828instances and13,60013\{,\}600randomized local starts, no configuration improved MRRW; instead, the recovered optima repeatedly matched the quasirandom tensor\-product family to numerical precision\.

This repeated null result exposed a structural obstruction\. Defining the translation affinity

Φ⁡\(G,v\)=∑uG⁡\(u\)​G​\(u\+v\),\\Phi\(G,v\)=\\sum\_\{u\}\\sqrt\{G\(u\)G\(u\+v\)\},the finite spectral walk condition impliesΦ⁡\(G,v\)≥ε\\Phi\(G,v\)\\geq\\varepsilonon a spanning set of shifts \(Methods, hypotheses \(A1\)–\(A2\)\)\. For

h⁡\(ε\)=𝖧⁡\(1−1−ε22\),h\(\\varepsilon\)=\\mathsf\{H\}\\\!\\left\(\\frac\{1\-\\sqrt\{1\-\\varepsilon^\{2\}\}\}\{2\}\\right\),convexity ofhh, conditioning and Jensen’s inequality give

𝖧⁡\(G\)≥∑i=1ℓh⁡\(Φ⁡\(G,ei\)\)≥ℓ​h​\(ε\)\.\\mathsf\{H\}\(G\)\\geq\\sum\_\{i=1\}^\{\\ell\}h\\\!\\left\(\\Phi\(G,e\_\{i\}\)\\right\)\\geq\\ell h\(\\varepsilon\)\.Hence, withδ=\(1−ε\)/2\\delta=\(1\-\\varepsilon\)/2,

Jℓ​\(G\)≥h⁡\(ε\)=𝖧⁡\(12−δ⁡\(1−δ\)\),J\_\{\\ell\}\(G\)\\geq h\(\\varepsilon\)=\\mathsf\{H\}\\\!\\left\(\\frac\{1\}\{2\}\-\\sqrt\{\\delta\(1\-\\delta\)\}\\right\),exactly the first MRRW expression\. The tensor\-product configurationBernoulli​\(pε\)⊗ℓ\\mathrm\{Bernoulli\}\(p\_\{\\varepsilon\}\)^\{\\otimes\\ell\}saturates the entropy–affinity inequality, identifying the object repeatedly recovered computationally\. Hence, within the fixed single\-orbit spectral ansatz, optimizing the configurationGGcannot improve MRRW Theorem[18](https://arxiv.org/html/2609.30296#Thmtheorem18)and Corollary[19](https://arxiv.org/html/2609.30296#Thmtheorem19)\)\.

The null result is informative because it is highly structured rather than merely negative\. Across 13\.600 local optimizations, the recovered configurations repeatedly collapse onto the same tensor\-product family to numerical precision \(Supplementary Methods[7\.5](https://arxiv.org/html/2609.30296#S7.SS5)\)\. This suggested that the observed optimum reflected an invariant of the reduced problem instead of optimization failure\. The entropy–affinity argument supports this interpretation: the tensor\-product configuration saturates the inequality and the numerically recovered object is an extremizer of the constrained configuration problem\.

Furthermore, within the single\-orbit spectral construction, additional optimization ofGGcannot improve the first MRRW expression; any improvement must instead exploit degrees of freedom not controlled by the obstruction, such as alternative sign functions, support on multiple configuration orbits, the partial\-Fourier dual families or more general dual feasible solutions \(Supplementary Methods[7\.6](https://arxiv.org/html/2609.30296#S7.SS6)and[7\.7](https://arxiv.org/html/2609.30296#S7.SS7)\)\. Thus, the computation eliminates an apparently natural search direction while identifying the components in which non\-trivial asymptotic freedom remains\.

## Convex discovery improves sign\-uncertainty bounds

For the\+1\+1Bourgain–Clozel–Kahane sign\-uncertainty problem, radial polynomial–Gaussian Fourier eigenfunctions are the canonical ansatz\. Cohn, Dong and Gonçalves\[[3](https://arxiv.org/html/2609.30296#bib.bib12)\]showed that sublinear\-degree members of this family encounter an asymptotic ceiling, motivating our initial search for more flexible Fourier\-eigenfunction structures\. We therefore treated the choice of envelope, scales and contact locations as a discovery problem, with learned proposals followed by independent high\-precision evaluation\.

The flexible search improved rapidly but stalled in dimensiond=1d=1\. Gaussian\-mixture constructions reachedρ=0\.578532\\rho=0\.578532, above the publishedA\+​\(1\)≤0\.572990A\_\{\+\}\(1\)\\leq 0\.572990bound\. Neural optimization did not remove the plateau: apparently improving trajectories systematically entered coefficient systems with condition numbers101710^\{17\}–101810^\{18\}, where no reliable digits could be guaranteed from the floating\-point solve\. Multiprecision reevaluation rejected these candidates\. Multistart runs also converged to distinct plateaus rather than a common value, consistent with changes in root multiplicity and contact structure in the last\-sign\-change functional\. We therefore treated the failure as a diagnostic of the search representation rather than evidence that the function class itself had been exhausted\.

The decisive simplification was to return to the Fourier\-eigenfunction constraint\. Withu=π​\|x\|2u=\\pi\|x\|^\{2\}, every radial\+1\+1polynomial–Gaussian eigenfunction in the truncated Laguerre space has the form

f⁡\(x\)=P⁡\(u\)​e−u,P∈span⁡\{L0\(d/2−1\)​\(2​u\),L2\(d/2−1\)​\(2​u\),…,L2​N\(d/2−1\)​\(2​u\)\}\.f\(x\)=P\(u\)e^\{\-u\},\\qquad P\\in\\operatorname\{span\}\\left\\\{L\_\{0\}^\{\(d/2\-1\)\}\(2u\),L\_\{2\}^\{\(d/2\-1\)\}\(2u\),\\ldots,L\_\{2N\}^\{\(d/2\-1\)\}\(2u\)\\right\\\}\.The Fourier constraint is linear in the Laguerre coefficients ande−u\>0e^\{\-u\}\>0\. For fixedu0u\_\{0\}, admissibility therefore reduces to

P\(0\)=0,P\(u\)≥0\(u≥u0\)\.P\(0\)=0,\\qquad P\(u\)\\geq 0\\quad\(u\\geq u\_\{0\}\)\.To exclude the zero polynomial and fix the otherwise arbitrary scale, we impose the linear normalization

∫u0∞P⁡\(u\)​e−u​𝑑u=1\.\\int\_\{u\_\{0\}\}^\{\\infty\}P\(u\)e^\{\-u\}\\,du=1\.Thus the search can be reformulated as a convex semi\-infinite feasibility problem in the coefficients, with bisection inu0u\_\{0\}, rather than a non\-convex optimization over forced contacts\. An adaptive exchange LP locates the active tangencies; these are then rationalized, used to reconstruct an exact Laguerre polynomial overℚ\\mathbb\{Q\}, and verified on the complete ray by exact factorization and a Sturm chain overℤ\\mathbb\{Z\}\. The numerical search determines the quality of the construction found, whereas admissibility of each quoted bound follows independently from the rational certificate\. No global optimality claim is made\.

Ind=1d=1the degree ladder first reproduces the published configuration and then improves it, with approximate last\-sign\-change radii

ρ22≈0\.572989678,ρ30≈0\.572706699,ρ38≈0\.572588699\.\\rho\_\{22\}\\approx 0\.572989678,\\qquad\\rho\_\{30\}\\approx 0\.572706699,\\qquad\\rho\_\{38\}\\approx 0\.572588699\.The degree\-2222search numerically resolves the same five\-contact configuration as the published construction to its stated precision; the new improvement occurs at higher degree\. The degree\-3838candidate is certified by an explicit rational polynomial, givingA\+​\(1\)≤0\.572588700\\boxed\{A\_\{\+\}\(1\)\\leq 0\.572588700\}\. The same pipeline givesA\+​\(2\)≤0\.756206237A\_\{\+\}\(2\)\\leq 0\.756206237, which is approximately7\.6×10−77\.6\\times 10^\{\-7\}below the published0\.7562070\.756207value and is therefore treated as a certified reproduction with a marginal gain rather than a substantive improvement \(see Figure[2](https://arxiv.org/html/2609.30296#Sx4.F2)\)\.

![Refer to caption](https://arxiv.org/html/2609.30296v1/Section3/images/figure_landscape3d.png)Figure 2:Objective landscape and trajectory search\.The last\-sign\-change functionalρ⁡\(τ1,τ2\)\\rho\(\\tau\_\{1\},\\tau\_\{2\}\)on a slice through the certified record configuration \(d=1d=1, degree 38,m=9m=9contacts,ρ=0\.5725887\\rho=0\.5725887\), obtained by varying the two innermost contacts with the remaining seven held atτ∗\\tau^\{\\ast\}\. The surface is drawn from the raw grid without smoothing or interpolation, and is evaluated in the damped Laguerre basisψn​\(u\)=Ln\(α\)​\(2​u\)​e−u\\psi\_\{n\}\(u\)=L^\{\(\\alpha\)\}\_\{n\}\(2u\)e^\{\-u\}: at this degree the underlying polynomial reaches∼1069\{\\sim\}10^\{69\}on the grid, so a monomial evaluation loses every digit at the outer tangencies and splits them into spurious sign changes\. The landscape has two sheets, a narrow valley \(colour\) and an upper sheet atρ≈3\\rho\\approx 3\(translucent\), with no intermediate values — crossing the dashed curve, a different root ofPPbecomes the last sign change andρ\\rholeaps\. Black markers show a simplex/Adam trajectory, its polyline broken wherever the objective jumps\. The observed jumps illustrate the sensitivity of direct optimization to changes in the last\-sign\-change root\. The convex reformulation avoids differentiating this discontinuous objective by using bisection inu0=π​ρ2u\_\{0\}=\\pi\\rho^\{2\}and solving a coefficient\-feasibility problem at each step\. The dashed curve is drawn as a high\-gradient contour and indicates whereρ\\rhojumps rather than resolving the discontinuity set exactly\.Higher\-dimensional runs approach, rather than cross, the known polynomial–Gaussian asymptotic ceiling; their interpretation is limited by the numerical degree ceiling of the floating\-point discovery stage \(Methods\)\. Thus, the new small\-dimensional bounds do not evade the Cohn–Dong–Gonçalves obstruction\. Instead, they show that the finite polynomial–Gaussian family had not been numerically exhausted\.

## Discussion

Can modern neural optimization rediscover and improve extremal constructions in hard analytic and combinatorial mathematics? Across the three problems studied here, the answer takes several distinct forms\. In the Maynard–Tao sieve, neural optimization identifies stronger extremizers, but at largekkthe neural representation ultimately gives way to the rational familyg⁡\(t\)=\(c\+\(k−1\)​t\)−1g\(t\)=\(c\+\(k\-1\)t\)^\{\-1\}\. The large\-kkresults illustrate how computational progress can arise from a new evaluation and certification strategy for an established mathematical family\. In the higher\-order Delsarte hierarchy, computational failure is itself informative: the search collapses onto a tensor\-product family, and the resulting invariant yields an obstruction to improving the first MRRW expression through the entropy\-based rate estimate under the spanning\-affinity and finite\-walk hypotheses considered here\. The result leaves open constructions that bypass these hypotheses or use alternative spectral supports or dual feasible solutions\. In sign uncertainty, neural optimization encounters a plateau; independent verification reveals severe ill\-conditioning, and the failure exposes a convex reformulation that produces new certified bounds\. In each case, the discovery representation is provisional: its purpose is to expose mathematical structure rather than to become part of the final proof\[[6](https://arxiv.org/html/2609.30296#bib.bib29),[25](https://arxiv.org/html/2609.30296#bib.bib11),[32](https://arxiv.org/html/2609.30296#bib.bib5)\]\. Computation therefore contributes not only by finding better objects, but also by identifying obstructions and more appropriate representations\. The numerical degree ceiling encountered in the sign uncertainty problem is a case in point: asymptotic theory indicates that the strongest bounds lie in the high\-degree regime, yet that regime is precisely where the discovery representation loses numerical resolution, so progress required changing the representation rather than pushing the degree further\.

Separating heuristic proposal generation from mathematically controlled inference also underlies neuro\-symbolic systems such as AlphaGeometry, in which learned construction proposals are coupled to a symbolic deduction engine\[[31](https://arxiv.org/html/2609.30296#bib.bib4)\]\. NeuralCert applies a related separation not to theorem search, but to continuous variational discovery: the learned component proposes mathematical objects, whereas the final inequality depends only on the reconstructed certificate\. The computational model need not itself be interpretable, exact or even retained; it can use floating\-point arithmetic, stochastic optimization, approximate quadrature and overparameterized models\. Discovery output is instead regarded as an exploratory and untrusted proposal that must cross into a mathematically explicit representation before it can support a claim\. Certification and verification then operate on that representation independently of the optimizer\. This design follows certificate\-based principles in verified computation\[[30](https://arxiv.org/html/2609.30296#bib.bib3)\]\. The failed large\-kkMaynard candidate demonstrated the value of this design\. A candidate neark=3600k=3600remained stable under grid refinement and below the ceiling, yet independent reconstruction revealed a scale inconsistency with inflation of the objective\. The optimizer had exploited the numerical error, generating a plausible extremizer that failed certification\. Conversely, the later rational rank\-one construction was accepted because its Rayleigh quotient can be independently reconstructed and certified\. Hence, structural similarity discovered numerically becomes mathematically relevant only when it survives the certification boundary, in line with the certificate discipline of computer\-assisted proof\[[13](https://arxiv.org/html/2609.30296#bib.bib2)\]\. The methodological contribution is therefore the transition between search representations and independently verifiable mathematical objects, with neural optimization serving as one component of that process\.

The same separation has appeared concurrently on this very problem\. Recently, the boundH1≤246H\_\{1\}\\leq 246was improved to240240\[[29](https://arxiv.org/html/2609.30296#bib.bib7)\], then to212212\[[2](https://arxiv.org/html/2609.30296#bib.bib9)\]and186186\[[21](https://arxiv.org/html/2609.30296#bib.bib10)\], the last two accompanied by Lean certificates and, in one case, a proof attributed to an automated system\. The present work automates different objects\. Instead of a proof generated through a proof assistant at a single small dimension \(k=49k=49,4545and4040respectively\), we discovered a trial function numerically and then removed from the argument entirely: what is certified is the Rayleigh quotient of one fixed, exactly specified function, at dimensions up tok≈6×108k\\approx 6\\times 10^\{8\}and uniformly tok=1\.88×109k=1\.88\\times 10^\{9\}\. Formal verification is the stronger guarantee, and extending certificates into that setting is a potent direction\. Conversely, the variational constant is shared infrastructure: earlier work requiresk≈3\.5×104k\\approx 3\.5\\times 10^\{4\}to reachm=2m=2because the available lower bound onMkM\_\{k\}falls well short oflog⁡k\\log k, and combining the sharper boundMk≥log⁡k−0\.307M\_\{k\}\\geq\\log k\-0\.307with the improved equidistribution estimates used in\[[29](https://arxiv.org/html/2609.30296#bib.bib7),[2](https://arxiv.org/html/2609.30296#bib.bib9),[21](https://arxiv.org/html/2609.30296#bib.bib10)\]would be expected to improveHmH\_\{m\}form≥2m\\geq 2beyond either ingredient alone\.

The framework also has clear limitations\. NeuralCert is not a general mechanism for converting numerical optimization into proof\. Certification requires that a discovery can be transferred to an explicit mathematical representation, and this transfer may itself become a bottleneck\. A failed reconstruction is not hard evidence that the discovered function is mathematically uninteresting; it only establishes that the proposed claim has not crossed the certification boundary\. Likewise, certificates for variational constructions establish the performance of explicit admissible trial functions, not their global optimality\. A further limitation concerns the standard of verification itself\. The certificates here are rigorous numerical objects; interval enclosures and exact modular arithmetic, not machine\-checked proofs, and a Lean or Coq development of the reduction and error bounds would be a strictly stronger guarantee\. That the same problem has recently attracted both approaches suggests the combination is within reach\. More generally, flexible neural \(over\)parameterizations are particularly useful when their additional representational freedom assists discovery\. Once lower\-dimensional structure, convexity or an explicit analytic family has been identified, retaining a neural representation offers limited advantage; the mathematically simpler representation should replace it\.

## Methods

## 1Discovery and certification of sieve constants improve prime gap bounds

### 1\.1Maynard–Tao variational formulation

The Maynard–Tao sieve reduces bounded gaps between primes to the construction of admissible trial functions with large variational quotient\[[19](https://arxiv.org/html/2609.30296#bib.bib1),[23](https://arxiv.org/html/2609.30296#bib.bib27)\]\. For the separable constructions considered here, the relevant denominator and numerator functionals can be expressed through pairwise one\-dimensional convolution integrals\. We write the trial function as

Fθ,c​\(t1,…,tk\)=∑j=1mcj​∏i=1kgθ,j​\(ti\),F\_\{\\theta,c\}\(t\_\{1\},\\ldots,t\_\{k\}\)=\\sum\_\{j=1\}^\{m\}c\_\{j\}\\prod\_\{i=1\}^\{k\}g\_\{\\theta,j\}\(t\_\{i\}\),\(3\)wheregθ,jg\_\{\\theta,j\}are one\-dimensional channels andc∈ℝmc\\in\\mathbb\{R\}^\{m\}is a linear mixing vector\. For each pair,

hj​ℓ​\(x\)=gθ,j​\(x\)​gθ,ℓ​\(x\),h\_\{j\\ell\}\(x\)=g\_\{\\theta,j\}\(x\)g\_\{\\theta,\\ell\}\(x\),\(4\)and thekk\-dimensional functionals reduce to quadratic forms

Ik,ε​\(Fθ,c\)=c⊤​A​\(θ\)​c,Jk,ε\(1\)​\(Fθ,c\)=c⊤​B​\(θ\)​c,I\_\{k,\\varepsilon\}\(F\_\{\\theta,c\}\)=c^\{\\top\}A\(\\theta\)c,\\qquad J^\{\(1\)\}\_\{k,\\varepsilon\}\(F\_\{\\theta,c\}\)=c^\{\\top\}B\(\\theta\)c,\(5\)whose entries depend on one\-dimensional integrals involvinghj​ℓ∗\(k−1\)h\_\{j\\ell\}^\{\*\(k\-1\)\}\. At fixed channels the objective is therefore the generalized Rayleigh quotient

R⁡\(θ\)=k​λmax​\(B⁡\(θ\),A⁡\(θ\)\)=k​maxc≠0​c⊤​B​\(θ\)​cc⊤​A​\(θ\)​c\.R\(\\theta\)=k\\,\\lambda\_\{\\max\}\\\!\\left\(B\(\\theta\),A\(\\theta\)\\right\)=k\\max\_\{c\\neq 0\}\\frac\{c^\{\\top\}B\(\\theta\)c\}\{c^\{\\top\}A\(\\theta\)c\}\.\(6\)The nonlinear optimizer consequently searches only over the channel manifold; the optimal linear combination is re\-solved as a generalized Ritz vector at each objective evaluation\[[22](https://arxiv.org/html/2609.30296#bib.bib21),[11](https://arxiv.org/html/2609.30296#bib.bib19)\]\. The precise normalization ofMkM\_\{k\}, itsε\\varepsilon\-enlarged analogue, and the conversion of thresholds inMkM\_\{k\}to bounds onHmH\_\{m\}are given in Supplementary Methods, Section[5\.1](https://arxiv.org/html/2609.30296#S5.SS1)\.

### 1\.2Neural discovery of separable trial functions

The discovery stage is used only to identify strong explicit trial functions and does not establish a rigorous inequality\. To expose the increasingly sharp boundary structure at largekk, each channel is parameterized as

gθ,j​\(x\)=rθ,j​\(x\)​e−ρj​x,ρj\>0,g\_\{\\theta,j\}\(x\)=r\_\{\\theta,j\}\(x\)e^\{\-\\rho\_\{j\}x\},\\qquad\\rho\_\{j\}\>0,\(7\)where the residual functions are produced by a shared multilayer perceptron with channel\-specific outputs\. The network receives\(x,e−k​x\)\(x,e^\{\-kx\}\)as input, explicitly exposing the naturalO⁡\(k−1\)O\(k^\{\-1\}\)boundary\-layer scale\. In the positive\-channel solver,rθ,j​\(x\)=exp⁡\(fθ,j​\(x,e−k​x\)\)r\_\{\\theta,j\}\(x\)=\\exp\(f\_\{\\theta,j\}\(x,e^\{\-kx\}\)\), allowing the convolution recursion to be performed without cancellation in logarithmic coordinates\.

The leading generalized Ritz vector is profiled out rather than learned jointly with the channel parameters\. If

B⁡\(θ\)​vθ=λθ​A​\(θ\)​vθ,vθ⊤​A​\(θ\)​vθ=1,B\(\\theta\)v\_\{\\theta\}=\\lambda\_\{\\theta\}A\(\\theta\)v\_\{\\theta\},\\qquad v\_\{\\theta\}^\{\\top\}A\(\\theta\)v\_\{\\theta\}=1,\(8\)we freezevθv\_\{\\theta\}and differentiate

qvθ​\(θ\)=k​vθ⊤​B​\(θ\)​vθvθ⊤​A​\(θ\)​vθ\.q\_\{v\_\{\\theta\}\}\(\\theta\)=k\\frac\{v\_\{\\theta\}^\{\\top\}B\(\\theta\)v\_\{\\theta\}\}\{v\_\{\\theta\}^\{\\top\}A\(\\theta\)v\_\{\\theta\}\}\.\(9\)Under the usual nondegeneracy condition, the Hellmann–Feynman/envelope identity\[[9](https://arxiv.org/html/2609.30296#bib.bib15),[27](https://arxiv.org/html/2609.30296#bib.bib24)\]gives the derivative of the profiled leading eigenvalue without differentiating through the eigendecomposition\. Adam\[[16](https://arxiv.org/html/2609.30296#bib.bib23)\]is used for exploration and a block L\-BFGS minorize–maximize procedure for polishing\[[17](https://arxiv.org/html/2609.30296#bib.bib16)\]\. Channel\-rank and grid\-resolution continuation are used during search\. Network architecture, optimization schedules, finite\-difference gradient controls and continuation rules are specified in Supplementary Methods, Sections[5\.4](https://arxiv.org/html/2609.30296#S5.SS4)–[5\.10](https://arxiv.org/html/2609.30296#S5.SS10)\.

### 1\.3Stable evaluation of high\-order convolution powers

Direct recursion onh∗ph^\{\*p\}becomes unstable asppgrows because algebraic concentration, exponential decay and global amplitude must otherwise be represented simultaneously\. For each channel pair we write

hj​ℓ​\(x\)=rj​ℓ​\(x\)​e−aj​ℓ​x,aj​ℓ=ρj\+ρℓ,h\_\{j\\ell\}\(x\)=r\_\{j\\ell\}\(x\)e^\{\-a\_\{j\\ell\}x\},\\qquad a\_\{j\\ell\}=\\rho\_\{j\}\+\\rho\_\{\\ell\},\(10\)and factor its convolution powers as

νp\(j​ℓ\)​\(s\)=sp−1​e−aj​ℓ​s​eσp\(j​ℓ\)​ψp\(j​ℓ\)​\(s\)\.\\nu^\{\(j\\ell\)\}\_\{p\}\(s\)=s^\{p\-1\}e^\{\-a\_\{j\\ell\}s\}e^\{\\sigma^\{\(j\\ell\)\}\_\{p\}\}\\psi^\{\(j\\ell\)\}\_\{p\}\(s\)\.\(11\)The algebraic and exponential terms are handled analytically, the scalarσp\\sigma\_\{p\}carries the global amplitude, and only the residual shapeψp\\psi\_\{p\}is represented numerically\. Substitution into the convolution identity and the change of variablesx=s​ux=suyield a beta\-weighted recursion,

ψp\+q​\(s\)=B⁡\(p,q\)​∫01ψp​\(s​u\)​ψq​\(s⁡\(1−u\)\)​up−1​\(1−u\)q−1B⁡\(p,q\)​𝑑u,\\psi\_\{p\+q\}\(s\)=B\(p,q\)\\int\_\{0\}^\{1\}\\psi\_\{p\}\(su\)\\psi\_\{q\}\(s\(1\-u\)\)\\frac\{u^\{p\-1\}\(1\-u\)^\{q\-1\}\}\{B\(p,q\)\}\\,du,\(12\)which is evaluated with Gauss–Jacobi quadrature\[[12](https://arxiv.org/html/2609.30296#bib.bib22)\]\. The concentration associated with increasingppandqqis thereby moved into a known quadrature weight rather than left as unresolved structure in the represented residual\.

For positive channels we storeξp=log⁡ψp\\xi\_\{p\}=\\log\\psi\_\{p\}and evaluate the positive quadrature sum by log\-sum\-exp, preventing underflow across repeated convolutions\. The required powerh∗\(k−1\)h^\{\*\(k\-1\)\}is assembled through a binary addition chain, reducing convolution depth fromO⁡\(k\)O\(k\)toO⁡\(log⁡k\)O\(\\log k\)\. Representation, convolution, primitive\-integration and outer\-integration grids are kept distinct\. Full recursion formulas, grid definitions and quadrature construction are given in Supplementary Methods, Sections[5\.5](https://arxiv.org/html/2609.30296#S5.SS5)–[5\.7](https://arxiv.org/html/2609.30296#S5.SS7)\.

### 1\.4Structure\-preserving Rayleigh–Ritz optimization

The denominator matrixAAis mathematically a Gram matrix and must satisfyA⪰0A\\succeq 0\. Entrywise accurate discretization does not automatically preserve this property\. We therefore use positive two\-point interpolation in the convolution recursion and a shared positive\-weight outer integration mesh across all channel pairs\. Whenever the inner quadrature nodes are common to all pairs, the discrete denominator has the form

Aj​ℓ\(N\)=∑αwα​Φα​j​Φα​ℓ,wα\>0,A^\{\(N\)\}\_\{j\\ell\}=\\sum\_\{\\alpha\}w\_\{\\alpha\}\\Phi\_\{\\alpha j\}\\Phi\_\{\\alpha\\ell\},\\qquad w\_\{\\alpha\}\>0,\(13\)and therefore retains the positive\-semidefinite Gram structure up to floating\-point round\-off\. In the saturated regime, where the inner rule for the channel primitives becomes pair\-dependent, this common representation is preserved only up to inner\-quadrature error; numerical adequacy is then assessed by the positive\-semidefiniteness, two\-grid and refinement diagnostics of Section[1\.5](https://arxiv.org/html/2609.30296#S1.SS5), which are safeguards rather than rigorous bounds on that error\. No bound reported here depends on this discrete structure: the certified values are produced by the independent exact certifier\.

Because channel norms may span hundreds of orders of magnitude, the Gram pencil is diagonally equilibrated by the exact congruence transformation

A↦D−1/2AD−1/2,B↦D−1/2BD−1/2,D=diag\(A11,…,Am​m\)\.A\\mapsto D^\{\-1/2\}AD^\{\-1/2\},\\qquad B\\mapsto D^\{\-1/2\}BD^\{\-1/2\},\\qquad D=\\operatorname\{diag\}\(A\_\{11\},\\ldots,A\_\{mm\}\)\.\(14\)Numerically unresolved directions are removed by rank\-revealing spectral whitening; no ridge term is added to the variational denominator\. The interpolation, shared\-mesh construction, equilibration and rank\-revealing Ritz solve are detailed in Supplementary Methods, Sections[5\.8](https://arxiv.org/html/2609.30296#S5.SS8)–[5\.9](https://arxiv.org/html/2609.30296#S5.SS9)\.

### 1\.5Validation of neural candidates

Validation is embedded into model selection because a flexible optimizer can exploit structured discretization error\. Candidate states are therefore required to reproduce the independently recomputed Rayleigh quotient, preserve numerical positive semidefiniteness of the denominator, retain sufficient effective rank, remain stable under a more conservative rank threshold, respect applicable rigorous upper bounds for both individual channels and the full quotient \(rejecting numerical states that spuriously exceed the Cauchy–Schwarz ceiling\[[23](https://arxiv.org/html/2609.30296#bib.bib27)\]\), and reproduce on a finer representation grid\. Closed\-form controls based ong⁡\(x\)=1g\(x\)=1andg⁡\(x\)=xg\(x\)=xexercise the convolution and integration paths before optimization\. Final candidates are re\-evaluated on multiple grids; extrapolation is used only diagnostically and never as a certified lower bound\. Exact validation criteria, controls and refinement procedures are provided in Supplementary Methods, Sections[5\.11](https://arxiv.org/html/2609.30296#S5.SS11)–[5\.12](https://arxiv.org/html/2609.30296#S5.SS12)\.

### 1\.6Exact certification of neural candidates

A discovered candidate is projected onto rational polynomial channels,

F⁡\(t1,…,tk\)=∑j=1mcj​∏i=1kqj​\(ti1\+ε\),F\(t\_\{1\},\\ldots,t\_\{k\}\)=\\sum\_\{j=1\}^\{m\}c\_\{j\}\\prod\_\{i=1\}^\{k\}q\_\{j\}\\\!\\left\(\\frac\{t\_\{i\}\}\{1\+\\varepsilon\}\\right\),\(15\)and the mixing vector is transferred from the preconditioned discovery frame to the corresponding exact channel frame\. For fixed rational channels and rational coefficients the associated Maynard quotient is rational; certification therefore reduces to exact evaluation of the two quadratic forms\.

Rather than constructing and reconstructing every exact Gram\-matrix entry, the certifier accumulates the pair contributions toc⊤​A​cc^\{\\top\}Acandc⊤​B​cc^\{\\top\}Bcmodulo machine\-word primes\. For each primeppit computes only two aggregate residues,

SA​\(p\)=∑j≤ℓκj​ℓ​cj​cℓ​Aj​ℓint\(modp\),SB​\(p\)=∑j≤ℓκj​ℓ​cj​cℓ​Bj​ℓint\(modp\)\.S\_\{A\}\(p\)=\\sum\_\{j\\leq\\ell\}\\kappa\_\{j\\ell\}c\_\{j\}c\_\{\\ell\}A^\{\\mathrm\{int\}\}\_\{j\\ell\}\\pmod\{p\},\\qquad S\_\{B\}\(p\)=\\sum\_\{j\\leq\\ell\}\\kappa\_\{j\\ell\}c\_\{j\}c\_\{\\ell\}B^\{\\mathrm\{int\}\}\_\{j\\ell\}\\pmod\{p\}\.\(16\)Only the aggregate integersSAS\_\{A\}andSBS\_\{B\}are reconstructed, with analytically known positive denominatorsDA,DBD\_\{A\},D\_\{B\}:

c⊤​A​c=SADA,c⊤​B​c=SBDB\.c^\{\\top\}Ac=\\frac\{S\_\{A\}\}\{D\_\{A\}\},\\qquad c^\{\\top\}Bc=\\frac\{S\_\{B\}\}\{D\_\{B\}\}\.\(17\)Rigorous a priori integer bounds\|SA\|≤ℬA\\lvert S\_\{A\}\\rvert\\leq\\mathcal\{B\}\_\{A\}and\|SB\|≤ℬB\\lvert S\_\{B\}\\rvert\\leq\\mathcal\{B\}\_\{B\}determine the reconstruction budget\. IfMP=∏p∈PpM\_\{P\}=\\prod\_\{p\\in P\}pand

MP\>2​max⁡\(ℬA,ℬB\),M\_\{P\}\>2\\max\(\\mathcal\{B\}\_\{A\},\\mathcal\{B\}\_\{B\}\),\(18\)the centered Chinese\-remainder representative is unique\. The final quotient is then formed exactly from the reconstructed integers and known denominators\. Thus the neural optimizer, numerical quadrature and grid extrapolation no longer enter the proof once the rational trial function has been fixed\.

Before exact evaluation, the candidate is compressed by loss\-controlled backward elimination: channels are removed only while the re\-optimized Rayleigh quotient remains within a prescribed relative loss budget\. The exact backend certifies the explicit reduced function, so pruning affects certificate strength and computational cost but not validity\. Detailed denominator formulas, coefficient\-extraction identities, rational representation transfer, pruning rules, integer bounds and deterministic CRT reconstruction are given in Supplementary Methods, Sections[5\.13](https://arxiv.org/html/2609.30296#S5.SS13)–[5\.17](https://arxiv.org/html/2609.30296#S5.SS17)\.

### 1\.7Large\-kkrational construction

At largekk, neural optimization is used as a structural discovery instrument rather than as the source of a certified numerical value\. The neural candidate is compressed into clustered rational families

grat​\(t\)=∑a=1r∑j=1pmaxua,j​\(ca\+n​t\)−j,n=k−1,g\_\{\\mathrm\{rat\}\}\(t\)=\\sum\_\{a=1\}^\{r\}\\sum\_\{j=1\}^\{p\_\{\\max\}\}u\_\{a,j\}\(c\_\{a\}\+nt\)^\{\-j\},\\qquad n=k\-1,\(19\)whereca\>0c\_\{a\}\>0are pole locations and the higher powers represent confluent directions\. For fixedcac\_\{a\}, the coefficientsua,ju\_\{a,j\}are eliminated by variable projection, so only the nonlinear pole locations are optimized\. The confluent basis has a direct geometric interpretation because

∂q∂cq​1c\+n​t=\(−1\)q​q\!​\(c\+n​t\)−\(q\+1\)\.\\frac\{\\partial^\{q\}\}\{\\partial c^\{q\}\}\\frac\{1\}\{c\+nt\}=\(\-1\)^\{q\}q\!\(c\+nt\)^\{\-\(q\+1\)\}\.\(20\)
The quality of a distilled representation is not determined by pointwise fitting error alone\. Each candidate is rebuilt in an independent deterministic evaluator, the corresponding Gram pencil is recomputed, and the generalized Rayleigh problem is solved afresh\. Candidates are rejected if they fail probability\-mass conservation, analytic moment checks, Fourier\-inversion stability, Gram\-rank diagnostics, cross\-cluster conditioning, or the rigorous ceiling

Rk≤kk−1​log⁡k\.R\_\{k\}\\leq\\frac\{k\}\{k\-1\}\\log k\.\(21\)Surviving pole locations are then locally refined under the deterministic Rayleigh objective\.

In the large\-kkregime, additional rational clusters can reduce the pointwise approximation error while failing these distributional checks\. The stable deterministic representation collapses to the first\-order family

g⁡\(t\)=1c\+n​t,n=k−1,g\(t\)=\\frac\{1\}\{c\+nt\},\\qquad n=k\-1,\(22\)which is subsequently optimized independently of the neural model\. This distillation step therefore uses neural search to identify structure, while the numerical value submitted for certification is obtained from a separate rational representation\. The clustered fit, deterministic Fourier evaluator, failure diagnostics and local refinement are detailed in Supplementary Methods, Section[5\.18](https://arxiv.org/html/2609.30296#S5.SS18)\.

### 1\.8Certified large\-kkevaluation in ball arithmetic

For the first\-order family in Eq\. \([22](https://arxiv.org/html/2609.30296#S1.E22)\), certification can be reduced to a one\-dimensional probabilistic Fourier calculation\. Letw=g2w=g^\{2\},m0=∫01w⁡\(t\)​𝑑t=\[c⁡\(c\+n\)\]−1m\_\{0\}=\\int\_\{0\}^\{1\}w\(t\)\\,dt=\[c\(c\+n\)\]^\{\-1\}, and letSSbe the sum ofn=k−1n=k\-1independent draws from the probability densityw/m0w/m\_\{0\}on\[0,1\]\[0,1\]\. With

G⁡\(ρ\)=∫0ρg⁡\(t\)​𝑑t,H⁡\(ρ\)=∫0ρw⁡\(t\)​𝑑t,G\(\\rho\)=\\int\_\{0\}^\{\\rho\}g\(t\)\\,dt,\\qquad H\(\\rho\)=\\int\_\{0\}^\{\\rho\}w\(t\)\\,dt,\(23\)the separable Maynard quotient reduces exactly to

R=k​ND,N=𝔼⁡\[G​\(1−S\)2​𝟏S<1\],D=𝔼⁡\[H⁡\(1−S\)​𝟏S<1\]\.R=\\frac\{kN\}\{D\},\\qquad N=\\mathbb\{E\}\\\!\\left\[G\(1\-S\)^\{2\}\\mathbf\{1\}\_\{S<1\}\\right\],\\qquad D=\\mathbb\{E\}\\\!\\left\[H\(1\-S\)\\mathbf\{1\}\_\{S<1\}\\right\]\.\(24\)Writing

φ⁡\(θ\)=\(w^​\(θ\)m0\)n,ΨX​\(θ\)=∫01X⁡\(ρ\)​ei​θ​ρ​𝑑ρ,\\varphi\(\\theta\)=\\left\(\\frac\{\\widehat\{w\}\(\\theta\)\}\{m\_\{0\}\}\\right\)^\{n\},\\qquad\\Psi\_\{X\}\(\\theta\)=\\int\_\{0\}^\{1\}X\(\\rho\)e^\{i\\theta\\rho\}\\,d\\rho,\(25\)forX∈\{G2,H\}X\\in\\\{G^\{2\},H\\\}, Fourier inversion gives

IX=12​π​∫ℝφ⁡\(θ\)​e−i​θ​ΨX​\(θ\)​𝑑θ\.I\_\{X\}=\\frac\{1\}\{2\\pi\}\\int\_\{\\mathbb\{R\}\}\\varphi\(\\theta\)e^\{\-i\\theta\}\\Psi\_\{X\}\(\\theta\)\\,d\\theta\.\(26\)
The certifier evaluates a truncated trapezoidal form of Eq\. \([26](https://arxiv.org/html/2609.30296#S1.E26)\) using Arb ball arithmetic\[[15](https://arxiv.org/html/2609.30296#bib.bib8)\]\. The difference from the exact integral is bounded by three separately controlled contributions: periodic aliasing, omitted Fourier mass, and rigorous numerical enclosure of the transforms and auxiliary integrals\. Aliasing is one\-sided because the corresponding convolution density is nonnegative and compactly supported; its residual mass is bounded by certified moment inequalities\. The omitted Fourier range is divided into a finite band controlled by Arb wide\-ball suprema and an analytic tail obtained from integration\-by\-parts decay\.

All error budgets remain in ball arithmetic until the final quotient is assembled\. IfNballN\_\{\\rm ball\}andDballD\_\{\\rm ball\}are the certified truncated\-sum enclosures andEN,EDE\_\{N\},E\_\{D\}the rigorous aliasing and truncation budgets, the exported lower bound is

Rlow=lb⁡\(k​lb⁡\(Nball−EN\)ub⁡\(Dball\+ED\)\),R\_\{\\rm low\}=\\operatorname\{lb\}\\\!\\left\(\\frac\{k\\,\\operatorname\{lb\}\(N\_\{\\rm ball\}\-E\_\{N\}\)\}\{\\operatorname\{ub\}\(D\_\{\\rm ball\}\+E\_\{D\}\)\}\\right\),\(27\)with all conversions explicitly rounded outward\. The stored claim is rounded downward, so changes in valid Arb subdivision or enclosure radii cannot strengthen the reported inequality\. The complete reduction, Fourier lemmas, moment bounds, tail estimates, special\-function evaluation and directed rounding procedure are given in Supplementary Methods, Sections[5\.20](https://arxiv.org/html/2609.30296#S5.SS20)–[5\.23](https://arxiv.org/html/2609.30296#S5.SS23)\.

### 1\.9Finite\-range certification and stable\-law asymptotics

The first\-order rational family also permits a certified bound over a continuous range of integer dimensions\. LetRk\(1\)R\_\{k\}^\{\(1\)\}denote the Rayleigh quotient ofgc​\(t\)=\(c\+\(k−1\)​t\)−1g\_\{c\}\(t\)=\(c\+\(k\-1\)t\)^\{\-1\}after selection ofcc\. To avoid optimizingccindependently at every integerkk, we use the standardized simplex headroom

η⁡\(c,k\)=1−𝔼⁡\[S\]sd⁡\(S\)\\eta\(c,k\)=\\frac\{1\-\\mathbb\{E\}\[S\]\}\{\\operatorname\{sd\}\(S\)\}\(28\)as a scale coordinate\. Its moments are available in closed form, and the large\-kkoptimizers empirically concentrate near a fixed value ofη\\eta\. A log\-uniform ladder beginning atk=100k=100was therefore generated by solvingη⁡\(c,k\)=−0\.4746\\eta\(c,k\)=\-0\.4746and certifying the resulting fixed rank\-one trial at each rung\. For consecutive certified valueski<ki\+1k\_\{i\}<k\_\{i\+1\}, monotonicity ofMkM\_\{k\}gives, for every integerk∈\[ki,ki\+1\]k\\in\[k\_\{i\},k\_\{i\+1\}\],

Mk≥Rcert\(1\)​\(ki\)≥log⁡k−\[log⁡ki\+1−Rcert\(1\)​\(ki\)\]\.M\_\{k\}\\geq R\_\{\\rm cert\}^\{\(1\)\}\(k\_\{i\}\)\\geq\\log k\-\\left\[\\log k\_\{i\+1\}\-R\_\{\\rm cert\}^\{\(1\)\}\(k\_\{i\}\)\\right\]\.\(29\)The maximum interval deficit over the complete ladder is0\.3067549360\.306754936, yielding the uniform certified statement

Mk≥log⁡k−0\.307,100≤k≤1\.88×109\.M\_\{k\}\\geq\\log k\-0\.307,\\qquad 100\\leq k\\leq 1\.88\\times 10^\{9\}\.\(30\)The ladder construction, scale surrogate and interval\-by\-interval calculation are given in Supplementary Methods, Sections[5\.24](https://arxiv.org/html/2609.30296#S5.SS24)–[5\.25](https://arxiv.org/html/2609.30296#S5.SS25)\.

The asymptotic behavior of the same rational family can be analysed independently of the numerical discovery pipeline\. Putn=k−1n=k\-1and choose the fixed\-shift scaling

ck−1=log⁡n\+a,a∈ℝ\.c\_\{k\}^\{\-1\}=\\log n\+a,\\qquad a\\in\\mathbb\{R\}\.\(31\)After rescaling the squared rational channel, the associated triangular array has Lévy density converging tox−2​d​xx^\{\-2\}\\,dx\. The centered boundary variable converges in distribution to

Xa=\(a\+1\)−Λ,X\_\{a\}=\(a\+1\)\-\\Lambda,\(32\)whereΛ\\Lambdais the spectrally positive11\-stable random variable with characteristic function

𝔼ei​s​Λ=exp\{∫0∞\(ei​s​x−1−isx𝟏\{x≤1\}\)d​xx2\}\.\\mathbb\{E\}e^\{is\\Lambda\}=\\exp\\left\\\{\\int\_\{0\}^\{\\infty\}\\left\(e^\{isx\}\-1\-isx\\mathbf\{1\}\_\{\\\{x\\leq 1\\\}\}\\right\)\\frac\{dx\}\{x^\{2\}\}\\right\\\}\.\(33\)Uniform density and logarithmic\-integrability estimates permit passage from weak convergence to the logarithmic observables in the Rayleigh quotient\. The resulting asymptotic defect is

R⁡\(Fk,ck\)=log⁡k−𝒞⁡\(a\)\+o⁡\(1\),𝒞⁡\(a\)=a−2​𝔼​\[log⁡Xa∣Xa\>0\]\.R\(F\_\{k,c\_\{k\}\}\)=\\log k\-\\mathcal\{C\}\(a\)\+o\(1\),\\qquad\\mathcal\{C\}\(a\)=a\-2\\,\\mathbb\{E\}\[\\log X\_\{a\}\\mid X\_\{a\}\>0\]\.\(34\)Thus, with

𝒞∗=infa∈ℝ𝒞⁡\(a\),\\mathcal\{C\}\_\{\*\}=\\inf\_\{a\\in\\mathbb\{R\}\}\\mathcal\{C\}\(a\),\(35\)the Maynard constant satisfies

Mk≥log⁡k−𝒞∗−o⁡\(1\)\.M\_\{k\}\\geq\\log k\-\\mathcal\{C\}\_\{\*\}\-o\(1\)\.\(36\)Numerical evaluation of the limiting expression suggestsC∗≈0\.3343C^\{\\ast\}\\approx 0\.3343\. This is a lower\-bound construction obtained from fixed\-shift scalings and is not asserted here to be the asymptotic optimum over all possible sequencesckc\_\{k\}\. The stable\-limit proof, logarithmic uniform\-integrability argument and numerical evaluation of𝒞∗\\mathcal\{C\}\_\{\*\}are detailed in Supplementary Methods, Section[5\.26](https://arxiv.org/html/2609.30296#S5.SS26)\.

### 1\.10ε\\varepsilon\-enlarged sieve calculations

We additionally evaluate the enlarged\-domain variant of the Maynard variational problem\. For each computed pair\(k,ε\)\(k,\\varepsilon\), the same discovery\-to\-certification separation is retained: a candidate is optimized numerically, transferred to the explicit representation used by theε\\varepsilon\-certifier, and the resulting lower bound is recomputed independently\. To compare the enlarged and vanilla searches we define

Δcert​\(k,ε\)=Lk,ε−Lk,0,\\Delta\_\{\\rm cert\}\(k,\\varepsilon\)=L\_\{k,\\varepsilon\}\-L\_\{k,0\},\(37\)whereLk,εL\_\{k,\\varepsilon\}andLk,0L\_\{k,0\}are the certified lower bounds obtained by the two pipelines\. Across the reliable sweep,Δcert\\Delta\_\{\\rm cert\}decreases withkkand changes sign for the tested values ofε\\varepsilon\. Linear interpolation of the reliable sign\-bracketing pairs gives the crossover estimates reported in the Results \(Supplementary Table[9](https://arxiv.org/html/2609.30296#S11.T9)\)\.

Because Eq\. \([37](https://arxiv.org/html/2609.30296#S1.E37)\) compares two lower bounds rather than the unknown exact variational constants, a negative value does not proveMk,ε<MkM\_\{k,\\varepsilon\}<M\_\{k\}\. It establishes only that the certified vanilla construction found by our pipeline outperforms the certified enlarged\-domain construction at thatkk\. We therefore treat the finite crossover as an empirical property of the certified construction ladder and state its extension to the exact variational constants as a conjecture\. In particular, the observed crossover locations motivate

ε​k∗​\(ε\)⟶∞\(ε→0\),\\varepsilon k^\{\*\}\(\\varepsilon\)\\longrightarrow\\infty\\qquad\(\\varepsilon\\to 0\),\(38\)but the available data do not determine a unique finer scaling law\. The sign brackets, excluded optimization failures, interpolation procedure and scope of the conjecture are given in Supplementary Methods, Section[5\.27](https://arxiv.org/html/2609.30296#S5.SS27)\.

### 1\.11Independent certificate verification

Large\-kkcertificates are self\-contained records containing the exact rational trial parameter, Fourier discretization parameters, error summaries and the claimed lower bound\. The standalone verifier does not trust stored intermediate numerical bounds\. It reconstructs the analytic constants, regenerates the far\-field partition, recomputes wide\-ball suprema, certified moments, Fourier nodes, trapezoidal sums, aliasing and truncation errors, and assembles a newRlowR\_\{\\rm low\}in ball arithmetic\. The claim is accepted only when this independently recomputed lower bound supports the stored inequality\.

The certificate proves the Rayleigh quotient of one fixed admissible trial function; it does not establish optimality of its parameter, re\-prove the external Maynard–Tao sieve theorem, or certify the diameter of the admissiblekk\-tuple used to convert a certifiedMkM\_\{k\}value to a prime\-gap bound\. Monte–Carlo and small\-kkcomparisons are falsification tests rather than proof components\. Certificate schema and trust boundary are specified in Supplementary Methods, Section[5\.31](https://arxiv.org/html/2609.30296#S5.SS31)\.

## 2Higher\-order Delsarte discovery and certification

##### Symmetry\-reduced hierarchy\.

For hierarchy levelrr, blocklengthnnand minimum distanceddwe construct the symmetry\-reduced higher\-order Delsarte LP over type vectors

Ir,n=\{α∈ℕ2r:∑u∈𝔽2rαu=n\},I\_\{r,n\}=\\Bigl\\\{\\alpha\\in\\mathbb\{N\}^\{2^\{r\}\}:\\textstyle\\sum\_\{u\\in\\mathbb\{F\}\_\{2\}^\{r\}\}\\alpha\_\{u\}=n\\Bigr\\\},where a type records the multiplicity of each columnu∈𝔽2ru\\in\\mathbb\{F\}\_\{2\}^\{r\}in anr×nr\\times nbinary matrix\. For linear codes the program is invariant underGL⁡\(r,2\)\\mathrm\{GL\}\(r,2\), so variables are indexed by group orbits\. Full Fourier constraints are represented by multivariate Krawtchouk coefficientsKα​\(β\)K\_\{\\alpha\}\(\\beta\), and the Loyfer–Linial strengthening adds the corresponding partial transforms\[[18](https://arxiv.org/html/2609.30296#bib.bib14)\]\. We implemented these transforms recursively and extended the symmetry reduction tor=3r=3, where sorting\-based orbit representatives are no longer valid; construction, orbit enumeration and validation anchors are given in Supplementary Methods[7\.1](https://arxiv.org/html/2609.30296#S7.SS1)\.

##### Exact dual certification\.

Floating\-point optimization is used only to locate competitive dual solutions\. Candidates are reconstructed at high precision, rounded to dyadic rationals and checked against every dual inequality in exact arithmetic\. Residual rounding violations are repaired using

∑β∈Ir,n\(nβ\)Kα\(β\)=2r​n𝟏\{α=nε0\},\\sum\_\{\\beta\\in I\_\{r,n\}\}\\binom\{n\}\{\\beta\}K\_\{\\alpha\}\(\\beta\)=2^\{rn\}\\mathbf\{1\}\_\{\\\{\\alpha=n\\varepsilon\_\{0\}\\\}\},which supplies an exactly feasible repair direction with a rational cost\. Finite\-nnbounds therefore rest on exact dual feasibility and weak duality alone, not on the numerical LP tolerance, and are upper bounds on the level\-rroptimum \(Supplementary Methods[7\.2](https://arxiv.org/html/2609.30296#S7.SS2)\)\.

### Reduction to an asymptotic configuration search

We next considered the spectral dual construction of Coregliano et al\., in which the complementary spectral function is supported on a single configuration orbit\[[5](https://arxiv.org/html/2609.30296#bib.bib13)\]\. A normalized configuration is a probability distributionGGon𝔽2ℓ\\mathbb\{F\}\_\{2\}^\{\\ell\}\. Because a level\-ℓ\\elldual bounds\|C\|ℓ\|C\|^\{\\ell\}, the asymptotically relevant rate objective is

Jℓ​\(G\)=𝖧⁡\(G\)ℓ,J\_\{\\ell\}\(G\)=\\frac\{\\mathsf\{H\}\(G\)\}\{\\ell\},with𝖧⁡\(⋅\)\\mathsf\{H\}\(\\cdot\)the Shannon entropy in bits; this normalization, and two discrepancies in the cited preprint that it exposes, are recorded in Supplementary Methods[7\.4](https://arxiv.org/html/2609.30296#S7.SS4)\. For a non\-zero shiftvvwe use the Bhattacharyya translation affinity

Φ⁡\(G,v\)=∑uG⁡\(u\)​G​\(u\+v\)\.\\Phi\(G,v\)=\\sum\_\{u\}\\sqrt\{G\(u\)G\(u\+v\)\}\.WritingG=p2G=p^\{2\}withp≥0p\\geq 0and‖p‖2=1\\\|p\\\|\_\{2\}=1givesΦ⁡\(G,v\)=p𝖳​Sv​p\\Phi\(G,v\)=p^\{\\mathsf\{T\}\}S\_\{v\}p, whereSvS\_\{v\}is the permutation induced by translation byvv, so discovery reduces to entropy minimization on the positive unit sphere under quadratic affinity constraints\. The analytic and symmetry reductions leave2ℓ−12^\{\\ell\}\-1free probability coordinates, at most1515forℓ≤4\\ell\\leq 4, with no remaining dependence onnn\. We had intended a neural parameterization ofGG, but at this point a generic network was unnecessary, so we searched the reduced space directly by multistart sequential quadratic programming:2828instance solves overℓ≤4\\ell\\leq 4andε∈\{0\.30,0\.20,0\.10,0\.05\}\\varepsilon\\in\\\{0\.30,0\.20,0\.10,0\.05\\\},13,60013\{,\}600local optimizations in total\. No run produced a configuration below the first MRRW benchmark, and the recovered optima matched the quasirandom tensor\-product family to within10−1110^\{\-11\}coordinatewise \(Supplementary Methods[7\.5](https://arxiv.org/html/2609.30296#S7.SS5)\)\.

### The tensorization obstruction

The repeated null result is explained by an inequality that holds for every admissible configuration\. Letε∈\(0,1\)\\varepsilon\\in\(0,1\), letVVbe the set of shifts at which the spectral walk condition of\[[5](https://arxiv.org/html/2609.30296#bib.bib13)\]is verified, and let

h⁡\(ε\)=H2​\(12​\(1−1−ε2\)\)=H2​\(12−δ⁡\(1−δ\)\),δ=1−ε2,h\(\\varepsilon\)=H\_\{2\}\\\!\\left\(\\tfrac\{1\}\{2\}\\bigl\(1\-\\sqrt\{1\-\\varepsilon^\{2\}\}\\bigr\)\\right\)=H\_\{2\}\\\!\\left\(\\tfrac\{1\}\{2\}\-\\sqrt\{\\delta\(1\-\\delta\)\}\\right\),\\qquad\\delta=\\tfrac\{1\-\\varepsilon\}\{2\},whereH2H\_\{2\}is the binary entropy function; the second form is the first MRRW expression\. Two properties of the construction enter: the admissible shifts span𝔽2ℓ\\mathbb\{F\}\_\{2\}^\{\\ell\}, and at each such shift a finite even spectral\-walk ordermmsatisfies a lower bound of the formWm​\(G,v\)≥c​εmW\_\{m\}\(G,v\)\\geq c\\,\\varepsilon^\{m\}withc≥1c\\geq 1, whereWmW\_\{m\}is the paired walk weight appearing in the configuration asymptotics of\[[5](https://arxiv.org/html/2609.30296#bib.bib13)\]\. ComparingWmW\_\{m\}termwise with the multinomial expansion ofΦ​\(G,v\)m\\Phi\(G,v\)^\{m\}givesWm​\(G,v\)≤Φ​\(G,v\)mW\_\{m\}\(G,v\)\\leq\\Phi\(G,v\)^\{m\}, so the walk condition forces the affinity floorΦ⁡\(G,v\)≥ε\\Phi\(G,v\)\\geq\\varepsilonat a fixedmm, with no exchange of then→∞n\\to\\inftyandm→∞m\\to\\inftylimits\. Conditioning each coordinate on the others turns that floor into an entropy floor: the conditional entropy ofXiX\_\{i\}givenX−iX\_\{\-i\}equals𝔼⁡\[h⁡\(2​pY​\(1−pY\)\)\]\\mathbb\{E\}\[h\(2\\sqrt\{p\_\{Y\}\(1\-p\_\{Y\}\)\}\)\]whileΦ⁡\(G,ei\)\\Phi\(G,e\_\{i\}\)equals𝔼⁡\[2​pY​\(1−pY\)\]\\mathbb\{E\}\[2\\sqrt\{p\_\{Y\}\(1\-p\_\{Y\}\)\}\]for the same conditional marginals, so convexity ofhhand Jensen’s inequality give𝖧⁡\(Xi∣X−i\)≥h⁡\(Φ⁡\(G,ei\)\)\\mathsf\{H\}\(X\_\{i\}\\mid X\_\{\-i\}\)\\geq h\(\\Phi\(G,e\_\{i\}\)\)\. The entropy chain rule then yields

𝖧⁡\(G\)≥∑i=1ℓh⁡\(Φ⁡\(G,ei\)\)≥ℓ​h​\(ε\),henceJℓ​\(G\)≥h⁡\(ε\)\.\\mathsf\{H\}\(G\)\\;\\geq\\;\\sum\_\{i=1\}^\{\\ell\}h\\bigl\(\\Phi\(G,e\_\{i\}\)\\bigr\)\\;\\geq\\;\\ell\\,h\(\\varepsilon\),\\qquad\\text\{hence\}\\qquad J\_\{\\ell\}\(G\)\\;\\geq\\;h\(\\varepsilon\)\.The tensor\-product configurationGε=Bernoulli⁡\(pε\)⊗ℓG\_\{\\varepsilon\}=\\operatorname\{Bernoulli\}\(p\_\{\\varepsilon\}\)^\{\\otimes\\ell\}attains equality in the relaxed affinity\-constrained problem\. Under the stronger finite\-walk hypotheses \(A1\)–\(A2\), however, the entropy\-based rate estimate is strictly larger thanh⁡\(ε\)h\(\\varepsilon\)\. Thus the computationally recovered tensor\-product structure is explained by the tensorization inequality within the analyzed relaxation\. This does not establish feasibility ofGεG\_\{\\varepsilon\}for the original finite\-walk construction, nor does it rule out improvements through configurations or dual constructions that lie outside these hypotheses\. Full hypotheses, proofs and scope are given in Supplementary Methods[7\.6](https://arxiv.org/html/2609.30296#S7.SS6)\(Theorem[18](https://arxiv.org/html/2609.30296#Thmtheorem18)and Corollary[19](https://arxiv.org/html/2609.30296#Thmtheorem19)\); the algebraic reductions underlying the theorem were verified against independent brute\-force implementations \(Supplementary Methods[7\.7](https://arxiv.org/html/2609.30296#S7.SS7)\)\.

The obstruction constrains only the configuration\. It does not constrain alternative sign functionsϕ\\phi, spectral functions supported on several configuration orbits, other uses of the partial\-Fourier dual families, or general dual feasible solutions, and the higher\-order hierarchy itself remains complete\. The computational search therefore localized the missing asymptotic freedom in those components rather than inGG\.

## 3Sign\-uncertainty discovery and certification

### Fourier\-eigenfunction formulation

We use

f^​\(ξ\)=∫ℝdf⁡\(x\)​e−2​π​i​⟨x,ξ⟩​𝑑x\\widehat\{f\}\(\\xi\)=\\int\_\{\\mathbb\{R\}^\{d\}\}f\(x\)e^\{\-2\\pi i\\langle x,\\xi\\rangle\}\\,dxand consider the\+1\+1sign\-uncertainty constantA\+​\(d\)A\_\{\+\}\(d\)over non\-zero even integrable functions satisfyingf^=f\\widehat\{f\}=f,f⁡\(0\)=0f\(0\)=0, andf⁡\(x\)≥0f\(x\)\\geq 0for all sufficiently large\|x\|\|x\|\[[1](https://arxiv.org/html/2609.30296#bib.bib26),[4](https://arxiv.org/html/2609.30296#bib.bib25)\]\. To construct upper bounds we restrict the search to finite\-dimensional radial Laguerre polynomial–Gaussian eigenspaces\. With

u=π​\|x\|2,α=d2−1,u=\\pi\|x\|^\{2\},\\qquad\\alpha=\\frac\{d\}\{2\}\-1,the functions

ψn​\(u\)=Ln\(α\)​\(2​u\)​e−u\\psi\_\{n\}\(u\)=L\_\{n\}^\{\(\\alpha\)\}\(2u\)e^\{\-u\}satisfyψn^=\(−1\)n​ψn\\widehat\{\\psi\_\{n\}\}=\(\-1\)^\{n\}\\psi\_\{n\}\. Hence the truncated\+1\+1eigenspace is

VN=span⁡\{ψ0,ψ2,…,ψ2​N\}\.V\_\{N\}=\\operatorname\{span\}\\\{\\psi\_\{0\},\\psi\_\{2\},\\ldots,\\psi\_\{2N\}\\\}\.Writingf=P⁡\(u\)​e−uf=P\(u\)e^\{\-u\}converts the Fourier constraint into the linear requirement thatPPlie in the corresponding even Laguerre span\. For integerddthese generalized Laguerre polynomials have rational coefficients, which enables exact reconstruction\.

### From neural proposals to a convex search

The initial discovery stage searched a broader Fourier\-eigenfunction family with trainable Gaussian scales and contact locations\. Local and neural optimization rapidly improved the objective but became unreliable when the coefficient systems approached condition numbers of101710^\{17\}–101810^\{18\}\. Independent multiprecision reevaluation showed that apparently improving trajectories were not reproducible\. We therefore treated learned widths and contacts only as proposals and used this failure to reconsider the search coordinates rather than to infer that the function class had been exhausted\.

For the pure Laguerre family the rescaled function is simply the polynomialP⁡\(u\)=eu​f​\(u\)P\(u\)=e^\{u\}f\(u\)\. Sincee−u\>0e^\{\-u\}\>0, fixing a candidate last\-sign pointu0u\_\{0\}reduces admissibility to

P\(0\)=0,P\(u\)≥0\(u≥u0\)\.P\(0\)=0,\\qquad P\(u\)\\geq 0\\quad\(u\\geq u\_\{0\}\)\.The feasible set in the Laguerre coefficients is convex and monotone inu0u\_\{0\}\. Thus, the non\-convex optimization over roots and contact locations can be replaced by a one\-dimensional bisection inu0u\_\{0\}with a convex semi\-infinite feasibility problem at each step\.

### Adaptive semi\-infinite feasibility

At each bisection point we solve a margin\-maximizing LP on an adaptive grid\. A positive linear normalization,

∫u0∞P⁡\(u\)​e−u​𝑑u=1,\\int\_\{u\_\{0\}\}^\{\\infty\}P\(u\)e^\{\-u\}\\,du=1,removes the conic scale, and the leading Laguerre coefficient is constrained to have the sign required for positivity at infinity\. The grid is only a relaxation\. After each LP solve, roots ofP′P^\{\\prime\}are located at high precision; any missed negative minimum is inserted as a cutting plane and the LP is resolved\. Near the fixed\-degree feasibility boundary, the active constraints identify candidate double contacts\.

The exchange loop is used to locate the boundary and its contact structure, not to certify the quoted upper bound\. The underlying space is not a Haar system, so uniqueness of the active contact set is not assumed\. Instead, the convex semi\-infinite formulation supplies a numerical bracket for sharpness within the chosen finite\-dimensional space, while exact admissibility is established independently downstream\. Failure of the numerical oracle can therefore reduce sharpness but cannot make an invalid certificate valid\.

### Exact rational reconstruction and verification

Letτ1<⋯<τm\\tau\_\{1\}<\\cdots<\\tau\_\{m\}denote the active tangencies returned by the convex search, rounded to nearby rational values\. We reconstructPPexactly in the Laguerre basis by imposing

P\(0\)=0,P\(τi\)=P′\(τi\)=0\(1≤i≤m\),P\(0\)=0,\\qquad P\(\\tau\_\{i\}\)=P^\{\\prime\}\(\\tau\_\{i\}\)=0\\quad\(1\\leq i\\leq m\),together with a fixed positive top coefficient\. In the square certificates,nb=2​m\+2n\_\{\\mathrm\{b\}\}=2m\+2Laguerre coordinates are retained\. Solving overℚ\\mathbb\{Q\}gives

P⁡\(u\)=∏i=1m\(u−τi\)2​R​\(u\),R∈ℚ⁡\[u\],P\(u\)=\\prod\_\{i=1\}^\{m\}\(u\-\\tau\_\{i\}\)^\{2\}R\(u\),\\qquad R\\in\\mathbb\{Q\}\[u\],with every polynomial division checked to have zero remainder\.

Positivity on the complete ray is then decided independently of the LP grid\. We verify that the leading coefficient ofRRis positive, thatR⁡\(u¯\)\>0R\(\\bar\{u\}\)\>0, and that a Sturm chain has the same sign\-variation count atu¯\\bar\{u\}and at\+∞\+\\infty\. HenceRRhas no zero on\[u¯,∞\)\[\\bar\{u\},\\infty\)and therefore

P⁡\(u\)≥0\(u≥u¯\)\.P\(u\)\\geq 0\\qquad\(u\\geq\\bar\{u\}\)\.Sincef^=f\\widehat\{f\}=fandf⁡\(0\)=0f\(0\)=0hold exactly by construction,f=P⁡\(u\)​e−uf=P\(u\)e^\{\-u\}is admissible and

A\+​\(d\)≤u¯/π\.A\_\{\+\}\(d\)\\leq\\sqrt\{\\bar\{u\}/\\pi\}\.
For thed=1d=1degree\-3838certificate,m=9m=9and

u¯=514 997 856 7615⋅1011,\\bar\{u\}=\\frac\{514\\,997\\,856\\,761\}\{5\\cdot 10^\{11\}\},which yields

A\+​\(1\)≤0\.572588699​…\.A\_\{\+\}\(1\)\\leq 0\.572588699\\ldots\.Thed=2d=2degree\-2222certificate is constructed analogously and givesA\+​\(2\)≤0\.756206236​…A\_\{\+\}\(2\)\\leq 0\.756206236\\ldots\. The final certificate contains only rational Laguerre data and integer/rational Sturm arithmetic; floating\-point computation is confined to discovery of a useful contact structure\.

### Scope

The finite\-dimensional convex reformulation does not contradict the Cohn–Dong–Gonçalves asymptotic ceiling for sublinear\-degree polynomial–Gaussian families\[[3](https://arxiv.org/html/2609.30296#bib.bib12)\]\. It shows instead that, at small dimension, the classical family had not been numerically exhausted by the previous root\-parameterized searches\. Detailed tripwires, discovery experiments, convergence diagnostics, contact degeneracies and higher\-dimensional numerical limits are given in the Supplementary Methods\.

## 4Use of large language models

A large language model \(Claude, Anthropic; Opus 4\.5, 4\.8 and Fable5\) was used throughout the exploratory development of the computational pipeline as an interactive programming and research\-assistance tool\. Its contributions were concentrated in software engineering \(making the scripts into a package\) and numerical debugging: proposing and revising implementation strategies, generating and refactoring code, and diagnosing numerical failure modes in the certification pipeline, including floating\-point underflow \(e\.g\. float\-64 versus float\-32 in GPU representations\) and dynamic\-range collapse, basis\-conditioning and rank\-instability effects, quadrature overflow, and \(formulation\) inconsistencies between the discovery and certification stages \(see Supplementary Tables[2](https://arxiv.org/html/2609.30296#S5.T2)and[3](https://arxiv.org/html/2609.30296#S5.T3)\)\. The model also served as a discussant for \(formal\) mathematical framing\.

The mathematical content of this work is the author’s\. The variational formulation, the reduction underlying the pipeline, the certification scheme, and all theorem statements were conceived, derived, and verified by the author; the model’s role with respect to the mathematics was limited to exposition, sanity\-checking, and the surfacing of relevant literature, all of which the author confirmed\. Where the model proposed an algorithmic optimisation, the proposal was accepted only after the author verified it\. For reference; PyTorch and Flint are not installed on Anthropic servers so no computations were done by a LLM\.

All model outputs were critically evaluated by the author, who bears sole responsibility for the correctness of every claim\. All reported numerical bounds are the output of deterministic software\. Crucially, the language model is not invoked by the final certification pipeline: the validity of certified bound rests entirely on the exact\-arithmetic certificate and its independent verification, which can be checked without reference to any part of the development process\. A structured contribution\-provenance record, and the complete certification and verification software are provided in the Supplementary Methods\.

## References

- \[1\]J\. Bourgain, L\. Clozel, and J\. Kahane\(2010\)Principe d’heisenberg et fonctions positives\.InAnnales de l’institut Fourier,Vol\.60,pp\. 1215–1232\.Cited by:[§3](https://arxiv.org/html/2609.30296#S3.SSx1.p1.2),[Rigorous tripwires](https://arxiv.org/html/2609.30296#Sx7.SSx1.p1.2)\.
- \[2\]F\. Charton, L\. Hong, K\. Lau, K\. Ono, G\. Remy, H\. C\. Siu, A\. A\. Swaminathan, J\. Thorner, and Y\. Xie\(2026\)A new bound for small gaps between primes\.Note:Preliminary draft, Axiom MathDated 3 September 2026\. Lean formalization at[https://github\.com/AxiomMath/PrimeGapsLib](https://github.com/AxiomMath/PrimeGapsLib)\. Accessed 14 September 2026Cited by:[Discovery and certification of sieve constants improve prime gap bounds](https://arxiv.org/html/2609.30296#Sx2.p7.1),[Discussion](https://arxiv.org/html/2609.30296#Sx5.p3.1)\.
- \[3\]H\. Cohn, D\. Dong, and F\. Gonçalves\(2022\)Sign uncertainty principles and low\-degree polynomials\.arXiv preprint arXiv:2210\.01684\.Cited by:[§3](https://arxiv.org/html/2609.30296#S3.SSx5.p1.1),[Convex discovery improves sign\-uncertainty bounds](https://arxiv.org/html/2609.30296#Sx4.p1.1),[Rigorous tripwires](https://arxiv.org/html/2609.30296#Sx7.SSx1.p1.2)\.
- \[4\]H\. Cohn and F\. Gonçalves\(2019\)An optimal uncertainty principle in twelve dimensions via modular forms: h\. cohn, f\. gonçalves\.Inventiones mathematicae217\(3\),pp\. 799–831\.Cited by:[§3](https://arxiv.org/html/2609.30296#S3.SSx1.p1.2)\.
- \[5\]L\. N\. Coregliano, F\. G\. Jeronimo, C\. Jones, N\. Linial, and E\. Loyfer\(2025\)Higher\-order delsarte dual lps: lifting, constructions and completeness\.arXiv preprint arXiv:2501\.04854\.Cited by:[§2](https://arxiv.org/html/2609.30296#S2.SSx1.p1.1),[§2](https://arxiv.org/html/2609.30296#S2.SSx2.p1.1),[§2](https://arxiv.org/html/2609.30296#S2.SSx2.p1.2),[§7\.3](https://arxiv.org/html/2609.30296#S7.SS3.p1.1),[§7\.6\.1](https://arxiv.org/html/2609.30296#S7.SS6.SSS1.p2.2),[§7\.6\.1](https://arxiv.org/html/2609.30296#S7.SS6.SSS1.p5.1),[Higher\-order Delsarte discovery exposes a tensorization obstruction](https://arxiv.org/html/2609.30296#Sx3.p1.1),[Higher\-order Delsarte discovery exposes a tensorization obstruction](https://arxiv.org/html/2609.30296#Sx3.p2.1),[Remark 21](https://arxiv.org/html/2609.30296#Thmtheorem21.p1.1),[Remark 24](https://arxiv.org/html/2609.30296#Thmtheorem24),[Remark 24](https://arxiv.org/html/2609.30296#Thmtheorem24.p1.1)\.
- \[6\]A\. Davies, P\. Veličković, L\. Buesing, S\. Blackwell, D\. Zheng, N\. Tomašev, R\. Tanburn, P\. Battaglia, C\. Blundell, A\. Juhász, M\. Lackenby, and G\. Williamson\(2021\)Advancing mathematics by guiding human intuition with AI\.Nature600,pp\. 70–74\.External Links:[Document](https://dx.doi.org/10.1038/s41586-021-04086-x)Cited by:[Discussion](https://arxiv.org/html/2609.30296#Sx5.p1.1),[NeuralCert: certified computational discovery of extremal mathematical constructions](https://arxiv.org/html/2609.30296#p1.1),[NeuralCert: certified computational discovery of extremal mathematical constructions](https://arxiv.org/html/2609.30296#p5.1)\.
- \[7\]T\. J\. Engelsma and A\. V\. Sutherland\(2013\)Narrow admissible tuples\.Note:[https://math\.mit\.edu/~primegaps/](https://math.mit.edu/~primegaps/)Entry fork=3655k=3655submitted by A\. V\. Sutherland, 27 June 2013\.Cited by:[Discovery and certification of sieve constants improve prime gap bounds](https://arxiv.org/html/2609.30296#Sx2.p4.2)\.
- \[8\]A\. Fawzi, M\. Balog, A\. Huang, T\. Hubert, B\. Romera\-Paredes, M\. Barekatain, A\. Novikov, F\. J\. R\. Ruiz, J\. Schrittwieser, G\. Swirszcz, D\. Silver, D\. Hassabis, and P\. Kohli\(2022\)Discovering faster matrix multiplication algorithms with reinforcement learning\.Nature610,pp\. 47–53\.External Links:[Document](https://dx.doi.org/10.1038/s41586-022-05172-4)Cited by:[NeuralCert: certified computational discovery of extremal mathematical constructions](https://arxiv.org/html/2609.30296#p1.1),[NeuralCert: certified computational discovery of extremal mathematical constructions](https://arxiv.org/html/2609.30296#p4.1)\.
- \[9\]R\. P\. Feynman\(1939\)Forces in molecules\.Physical Review56,pp\. 340–343\.External Links:[Document](https://dx.doi.org/10.1103/PhysRev.56.340)Cited by:[§1\.2](https://arxiv.org/html/2609.30296#S1.SS2.p2.3)\.
- \[10\]M\. Ghadimi\(2025\)Heuristic bounded prime gaps via a chaotic multidimensional sieve and random matrix theory\.Note:arXiv:2507\.17986v1External Links:2507\.17986Cited by:[§5\.2](https://arxiv.org/html/2609.30296#S5.SS2.p5.1)\.
- \[11\]G\. H\. Golub and C\. F\. Van Loan\(2013\)Matrix computations\.4 edition,Johns Hopkins University Press\.Cited by:[§1\.1](https://arxiv.org/html/2609.30296#S1.SS1.p1.5)\.
- \[12\]G\. H\. Golub and J\. H\. Welsch\(1969\)Calculation of gauss quadrature rules\.Mathematics of Computation23\(106\),pp\. 221–230\.External Links:[Document](https://dx.doi.org/10.1090/S0025-5718-69-99647-1)Cited by:[§1\.3](https://arxiv.org/html/2609.30296#S1.SS3.p1.4)\.
- \[13\]T\. Hales, M\. Adams, G\. Bauer, T\. D\. Dang, J\. Harrison, L\. T\. Hoang, C\. Kaliszyk, V\. Magron, S\. McLaughlin, T\. T\. Nguyen,et al\.\(2017\)A formal proof of the kepler conjecture\.InForum of mathematics, Pi,Vol\.5,pp\. e2\.Cited by:[Discussion](https://arxiv.org/html/2609.30296#Sx5.p2.1)\.
- \[14\]W\. Hart, F\. Johansson, and S\. Pancratz\(2012\)FLINT: fast library for number theory\.ACM Communications in Computer Algebra46\(3/4\),pp\. 88–93\.Cited by:[§5\.13](https://arxiv.org/html/2609.30296#S5.SS13.p2.1)\.
- \[15\]F\. Johansson\(2017\)Arb: efficient arbitrary\-precision midpoint\-radius interval arithmetic\.IEEE Transactions on Computers66\(8\),pp\. 1281–1292\.Cited by:[§1\.8](https://arxiv.org/html/2609.30296#S1.SS8.p2.1)\.
- \[16\]D\. P\. Kingma and J\. Ba\(2015\)Adam: a method for stochastic optimization\.InInternational Conference on Learning Representations,Cited by:[§1\.2](https://arxiv.org/html/2609.30296#S1.SS2.p2.3)\.
- \[17\]D\. C\. Liu and J\. Nocedal\(1989\)On the limited memory bfgs method for large scale optimization\.Mathematical Programming45,pp\. 503–528\.External Links:[Document](https://dx.doi.org/10.1007/BF01589116)Cited by:[§1\.2](https://arxiv.org/html/2609.30296#S1.SS2.p2.3)\.
- \[18\]E\. Loyfer and N\. Linial\(2023\)New lp\-based upper bounds in the rate\-vs\.\-distance problem for binary linear codes\.IEEE Transactions on Information Theory69\(5\),pp\. 2886–2899\.Cited by:[§2](https://arxiv.org/html/2609.30296#S2.SS0.SSS0.Px1.p1.2),[Higher\-order Delsarte discovery exposes a tensorization obstruction](https://arxiv.org/html/2609.30296#Sx3.p1.1)\.
- \[19\]J\. Maynard\(2015\)Small gaps between primes\.Annals of mathematics,pp\. 383–413\.Cited by:[§1\.1](https://arxiv.org/html/2609.30296#S1.SS1.p1.1),[§5\.2](https://arxiv.org/html/2609.30296#S5.SS2.p1.1)\.
- \[20\]K\. Mundinger, M\. Zimmer, A\. Kiem, C\. Spiegel, and S\. Pokutta\(2025\)Neural discovery in mathematics: do machines dream of colored planes?\.InProceedings of the 42nd International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.267\.Cited by:[NeuralCert: certified computational discovery of extremal mathematical constructions](https://arxiv.org/html/2609.30296#p1.1),[NeuralCert: certified computational discovery of extremal mathematical constructions](https://arxiv.org/html/2609.30296#p2.3)\.
- \[21\]OpenAI\(2026\)Improved short gaps between primes\.Note:[https://cdn\.openai\.com/pdf/51126fac\-1b68\-4128\-9666\-c908bcc16033/short\_gaps\.pdf](https://cdn.openai.com/pdf/51126fac-1b68-4128-9666-c908bcc16033/short_gaps.pdf)Dated 30 August 2026\. Proof attributed to GPT\-6 Astra; Lean 4 formalization and numerical certificate at[https://github\.com/openai/PrimeGaps186](https://github.com/openai/PrimeGaps186)\. Accessed 14 September 2026Cited by:[Discovery and certification of sieve constants improve prime gap bounds](https://arxiv.org/html/2609.30296#Sx2.p7.1),[Discussion](https://arxiv.org/html/2609.30296#Sx5.p3.1)\.
- \[22\]B\. N\. Parlett\(1998\)The symmetric eigenvalue problem\.SIAM\.Cited by:[§1\.1](https://arxiv.org/html/2609.30296#S1.SS1.p1.5)\.
- \[23\]D\. H\. J\. Polymath\(2014\)Variants of the selberg sieve, and bounded intervals containing many primes\.External Links:1407\.4897,[Link](https://arxiv.org/abs/1407.4897)Cited by:[§1\.1](https://arxiv.org/html/2609.30296#S1.SS1.p1.1),[§1\.5](https://arxiv.org/html/2609.30296#S1.SS5.p1.1),[§5\.1](https://arxiv.org/html/2609.30296#S5.SS1.p1.6),[Discovery and certification of sieve constants improve prime gap bounds](https://arxiv.org/html/2609.30296#Sx2.p4.2)\.
- \[24\]D\. K\. Probst and V\. S\. Alagar\(1979\)A family of algorithms for powering sparse polynomials\.SIAM Journal on Computing8\(4\),pp\. 626–644\.External Links:[Document](https://dx.doi.org/10.1137/0208050)Cited by:[§5\.13](https://arxiv.org/html/2609.30296#S5.SS13.p6.1)\.
- \[25\]G\. Raayoni, S\. Gottlieb, Y\. Manor, G\. Pisha, Y\. Harris, U\. Mendlovic, D\. Haviv, Y\. Hadad, and I\. Kaminer\(2021\)Generating conjectures on fundamental constants with the ramanujan machine\.Nature590\(7844\),pp\. 67–73\.Cited by:[Discussion](https://arxiv.org/html/2609.30296#Sx5.p1.1)\.
- \[26\]B\. Romera\-Paredes, M\. Barekatain, A\. Novikov, M\. Balog, M\. P\. Kumar, E\. Dupont, F\. J\. R\. Ruiz, J\. S\. Ellenberg, P\. Wang, O\. Fawzi, P\. Kohli, and A\. Fawzi\(2024\)Mathematical discoveries from program search with large language models\.Nature625,pp\. 468–475\.External Links:[Document](https://dx.doi.org/10.1038/s41586-023-06924-6)Cited by:[NeuralCert: certified computational discovery of extremal mathematical constructions](https://arxiv.org/html/2609.30296#p1.1),[NeuralCert: certified computational discovery of extremal mathematical constructions](https://arxiv.org/html/2609.30296#p4.1)\.
- \[27\]C\. Semay\(2015\)The hellmann–feynman theorem, the comparison theorem, and the envelope theory\.Results in Physics5,pp\. 322–323\.Cited by:[§1\.2](https://arxiv.org/html/2609.30296#S1.SS2.p2.3)\.
- \[28\]J\. Stadlmann\(2025\)On primes in arithmetic progressions and bounded gaps between many primes\.Advances in Mathematics468,pp\. 110190\.Cited by:[Discovery and certification of sieve constants improve prime gap bounds](https://arxiv.org/html/2609.30296#Sx2.p4.2)\.
- \[29\]J\. Stadlmann\(2026\)Bounded gaps between primes\.Note:Submitted 31 August 2026External Links:2608\.31126,[Link](https://arxiv.org/abs/2608.31126)Cited by:[Discovery and certification of sieve constants improve prime gap bounds](https://arxiv.org/html/2609.30296#Sx2.p4.2),[Discovery and certification of sieve constants improve prime gap bounds](https://arxiv.org/html/2609.30296#Sx2.p7.1),[Discussion](https://arxiv.org/html/2609.30296#Sx5.p3.1)\.
- \[30\]Y\. K\. Tan, J\. Yang, M\. Soos, M\. O\. Myreen, and K\. S\. Meel\(2024\)Formally certified approximate model counting\.InInternational Conference on Computer Aided Verification,pp\. 153–177\.Cited by:[Discussion](https://arxiv.org/html/2609.30296#Sx5.p2.1)\.
- \[31\]T\. H\. Trinh, Y\. Wu, Q\. V\. Le, H\. He, and T\. Luong\(2024\)Solving olympiad geometry without human demonstrations\.Nature625\(7995\),pp\. 476–482\.Cited by:[Discussion](https://arxiv.org/html/2609.30296#Sx5.p2.1)\.
- \[32\]S\. Udrescu and M\. Tegmark\(2020\)AI feynman: a physics\-inspired method for symbolic regression\.Science advances6\(16\),pp\. eaay2631\.Cited by:[Discussion](https://arxiv.org/html/2609.30296#Sx5.p1.1)\.

## 5Supplementary Methods: Maynard–Tao sieve

### 5\.1Variational definitions and normalization

Let

ℛk=\{𝐭∈ℝ≥0k:∑i=1kti≤1\}\.\\mathcal\{R\}\_\{k\}=\\left\\\{\\mathbf\{t\}\\in\\mathbb\{R\}\_\{\\geq 0\}^\{k\}:\\sum\_\{i=1\}^\{k\}t\_\{i\}\\leq 1\\right\\\}\.For an admissible square\-integrable trial functionFFsupported onℛk\\mathcal\{R\}\_\{k\}, define

Ik​\(F\)=∫ℛkF​\(𝐭\)2​𝑑𝐭I\_\{k\}\(F\)=\\int\_\{\\mathcal\{R\}\_\{k\}\}F\(\\mathbf\{t\}\)^\{2\}\\,d\\mathbf\{t\}and

Jk\(m\)​\(F\)=∫ℛk−1\(∫01−∑i≠mtiF⁡\(𝐭\)​d​tm\)2​d​𝐭−m\.J\_\{k\}^\{\(m\)\}\(F\)=\\int\_\{\\mathcal\{R\}\_\{k\-1\}\}\\left\(\\int\_\{0\}^\{1\-\\sum\_\{i\\neq m\}t\_\{i\}\}F\(\\mathbf\{t\}\)\\,dt\_\{m\}\\right\)^\{2\}d\\mathbf\{t\}\_\{\-m\}\.The Maynard variational constant is

Mk=supF≠0∑m=1kJk\(m\)​\(F\)Ik​\(F\)\.M\_\{k\}=\\sup\_\{F\\neq 0\}\\frac\{\\sum\_\{m=1\}^\{k\}J\_\{k\}^\{\(m\)\}\(F\)\}\{I\_\{k\}\(F\)\}\.For symmetricFF, all marginal functionals are equal and

Mk=k​supF≠0Jk\(1\)​\(F\)Ik​\(F\)\.M\_\{k\}=k\\sup\_\{F\\neq 0\}\\frac\{J\_\{k\}^\{\(1\)\}\(F\)\}\{I\_\{k\}\(F\)\}\.Throughout the computational pipeline, the reported Rayleigh quotient is normalized according to this convention\. We also use the classical Cauchy–Schwarz\[[23](https://arxiv.org/html/2609.30296#bib.bib27)\]ceiling

Mk≤kk−1​log⁡k,M\_\{k\}\\leq\\frac\{k\}\{k\-1\}\\log k,as a validation constraint on both single\-channel and full numerical quotients\. The enlarged\-support calculations use the corresponding Polymathε\\varepsilon\-variant, with the same discovery–certification separation\. Because the exact normalization of the enlarged polytope and the external conversion from a certified variational threshold to a numericalHmH\_\{m\}bound involve the sieve theorem and an admissiblekk\-tuple rather than the NeuralCert certificate itself, those ingredients are kept distinct from the variational certificate\. In particular, the certificate proves the stated lower bound for the explicit trial function; the finalHmH\_\{m\}conversion additionally uses the external Maynard–Tao/Polymath theorem and a separately supplied admissible\-tuple diameter\.

### 5\.2From exact replication to neural discovery

The Maynard experiments began from an exact finite\-dimensional control rather than from a neural model\. Following Lemmas 8\.1 and 8\.2 of\[[19](https://arxiv.org/html/2609.30296#bib.bib1)\], we used the symmetric polynomial family

F⁡\(𝐭\)=∑ℓ=1daℓ​\(1−P1​\(𝐭\)\)bℓ​P2​\(𝐭\)cℓ,Pj​\(𝐭\)=∑i=1ktij,F\(\\mathbf\{t\}\)=\\sum\_\{\\ell=1\}^\{d\}a\_\{\\ell\}\(1\-P\_\{1\}\(\\mathbf\{t\}\)\)^\{b\_\{\\ell\}\}P\_\{2\}\(\\mathbf\{t\}\)^\{c\_\{\\ell\}\},\\qquad P\_\{j\}\(\\mathbf\{t\}\)=\\sum\_\{i=1\}^\{k\}t\_\{i\}^\{j\},for which bothIkI\_\{k\}and∑mJk\(m\)\\sum\_\{m\}J\_\{k\}^\{\(m\)\}reduce to rational quadratic forms in the coefficient vectoraa\. The finite\-dimensional optimization is therefore a generalized Rayleigh–Ritz problem\. This approach is consistent with Maynard’s observation that an extremizer should satisfy an eigenfunction equation for the underlying integral operator\. The neural stage can therefore be interpreted as a nonlinear numerical search for a dominant variational eigenfunction, rather than as unconstrained function fitting\. As a first control we reproduced Maynard’s explicitk=5k=5trial

F⁡\(𝐭\)=\(1−P1\)​P2\+710​\(1−P1\)2\+114​P2−314​\(1−P1\),F\(\\mathbf\{t\}\)=\(1\-P\_\{1\}\)P\_\{2\}\+\\frac\{7\}\{10\}\(1\-P\_\{1\}\)^\{2\}\+\\frac\{1\}\{14\}P\_\{2\}\-\\frac\{3\}\{14\}\(1\-P\_\{1\}\),for which

Q5​\(F\)=1417255708216\>2\.Q\_\{5\}\(F\)=\\frac\{1417255\}\{708216\}\>2\.We subsequently reproduced the degree\-constrained generalized eigenvalue calculation atk=105k=105\. Exact rational assembly was retained throughout; the floating\-point generalized eigensolve was used only to locate a Ritz vector, which was then rationalized and reevaluated in the exact quadratic forms\.

The exact polynomial controls were next used to calibrate Monte Carlo estimators before any black\-box optimization\. Uniform simplex samples were generated from normalized independent exponential variables\. For the marginal functional, two conditionally independent samplesU,U′U,U^\{\\prime\}were drawn on the same fibre\. This gives the unbiased product estimator

𝔼⁡\[F⁡\(U,𝐖\)​F​\(U′,𝐖\)∣𝐖\]=\(1C​∫0CF⁡\(u,𝐖\)​𝑑u\)2,\\mathbb\{E\}\[F\(U,\\mathbf\{W\}\)F\(U^\{\\prime\},\\mathbf\{W\}\)\\mid\\mathbf\{W\}\]=\\left\(\\frac\{1\}\{C\}\\int\_\{0\}^\{C\}F\(u,\\mathbf\{W\}\)\\,du\\right\)^\{2\},whereC=1−∑WiC=1\-\\sum W\_\{i\}\. Squaring a single inner sample would add its conditional variance and bias the estimate upward\. Agreement of the Monte Carlo estimates with the exact polynomial values, expressed through standardized residuals, was used as the first stochastic validation gate\.

Only after these controls were in place was the fixed polynomial ansatz replaced by a trainable symmetric function\. Early neural models used scaled power\-sum coordinates

ϕj​\(𝐭\)=kj−1​∑itij\\phi\_\{j\}\(\\mathbf\{t\}\)=k^\{j\-1\}\\sum\_\{i\}t\_\{i\}^\{j\}and optimized a stochastic Rayleigh quotient\. These experiments established that the optimizer could recover useful symmetric structure without being given the final polynomial representation, but also exposed the variance, conditioning and representation\-transfer problems that motivated the deterministic separable pipeline used in the main analysis\.

Finally, further control showed that enlargement of the admissible domain does not by itself improve a fixed trial function\. For example, Maynard’s Eq\. \(8\.16\), optimized for the simplex geometry, decreases from approximately2\.002\.00to1\.661\.66when evaluated atε=1/4\\varepsilon=1/4\. The gain from enlarged support therefore requires re\-optimization to the modified geometry rather than simple reuse of a simplex\-optimized extremizer\. This provided an additional motivation for learning the trial function directly in eachε\\varepsilon\-geometry\.

Recent computational work has also explored heuristic modifications of Maynard\-type multidimensional sieves, including dynamically enlarged support regions and random\-matrix\-inspired perturbations\[[10](https://arxiv.org/html/2609.30296#bib.bib18)\], but their predicted prime\-gap improvements remain heuristic\. In contrast, the present framework separates exploratory optimization from certification of an explicit variational trial function and invokes only independently stated sieve\-theoretic inputs for the final number\-theoretic consequence\.

### 5\.3Exploratory representation experiments

Before arriving at the final separable channel representation, five candidate compression mechanisms, most of which were ultimately rejected\. These experiments were diagnostic rather than proof components: their role was to identify which representations preserved the variational structure well enough to justify further development\. All bounds in the paper are certified independently of the outcomes described below\.

#### 5\.3\.1Experiment 1: partitions

Exact Rayleigh–Ritz optimization in the full monomial\-symmetric basis is effective at smallkk, but encounters a symmetric curse of dimensionality: although the partition basis isL2L^\{2\}\-dense, the number of basis elements up to degreeDDgrows as the restricted partition countpk​\(D\)p\_\{k\}\(D\)and becomes rapidly prohibitive\. This motivated attempts to retain only a small, variationally relevant subset of partitions\.

Three selectors were tested: exact one\-vector residual greedy expansion, a global floating\-point Ritz oracle, and neural coefficient extraction through a power\-sum\-product representation\. The residual score

ηλ=\|⟨mλ,\(L−λS\)​FS⟩\|‖mλ‖2\\eta\_\{\\lambda\}=\\frac\{\|\\langle m\_\{\\lambda\},\(L\-\\lambda\_\{S\}\)F\_\{S\}\\rangle\|\}\{\\\|m\_\{\\lambda\}\\\|\_\{2\}\}improved the recorded exploratory bound by only about2\.5×10−32\.5\\times 10^\{\-3\}after adding approximately 30 highest\-scoring partitions, much less than adding a complete degree shell\. The global floating\-point and neural coefficient selectors were additionally destabilized by the ill\-conditioned change of coordinates into the monomial\-symmetric basis\. The experiment therefore rejected sparse single\-partition selection as an effective acceleration mechanism: the missing variational gain was collective and strongly correlated rather than localized in a few coordinates\.

This failure suggested that the relevant structure was not sparse in the partition basis but instead closer to a low\-complexity product or tensor representation, motivating the tensor\-decomposition experiments below\.

#### 5\.3\.2Experiment 2: symmetric tensor decompositions

The failure of sparse partition selection suggested that the optimizer might be compact in a tensor representation\. We therefore fitted the symmetric CP family

FCP​\(𝐭\)=∑q=1rcq​∏i=1kϕq​\(ti\)F\_\{\\rm CP\}\(\\mathbf\{t\}\)=\\sum\_\{q=1\}^\{r\}c\_\{q\}\\prod\_\{i=1\}^\{k\}\\phi\_\{q\}\(t\_\{i\}\)to the certifiedk=5k=5, degree\-1010Ritz optimizer\. The relative empiricalL2L^\{2\}reconstruction errors were

r123456ℰr0\.18870\.15360\.10320\.10280\.10760\.1089\\begin\{array\}\[\]\{c\|rrrrrr\}r&1&2&3&4&5&6\\\\ \\hline\\cr\\mathcal\{E\}\_\{r\}&0\.1887&0\.1536&0\.1032&0\.1028&0\.1076&0\.1089\\end\{array\}and plateaued near10%10\\%beyond rank three\. The corresponding Monte Carlo Rayleigh estimates were statistically compatible with the exact target but noisy and non\-monotone; values above the target were within the quoted Monte Carlo uncertainty and were never treated as improved lower bounds\. Thus, increasing the symmetric CP rank beyondr=3r=3produced essentially no further reduction in reconstruction error over the tested range\. This rejected the simplest low\-rank CP compression as a complete representation, but not the broader tensor\-network hypothesis\.

#### 5\.3\.3Experiment 3: structural signatures of the CP residual

For the rank\-six symmetric CP fit, the relative empirical reconstruction error was

‖F⋆−FCP,6‖2‖F⋆‖2=0\.10903\.\\frac\{\\\|F^\{\\star\}\-F\_\{\\mathrm\{CP\},6\}\\\|\_\{2\}\}\{\\\|F^\{\\star\}\\\|\_\{2\}\}=0\.10903\.We projected the residual onto the full degree\-10 monomial\-symmetric basis\. The largest fitted coordinate signatures occurred in the one\-part directions

\(4\),\(3\),\(2\),\(5\),\(1\)\(4\),\(3\),\(2\),\(5\),\(1\)and the two\-part directions

\(2,1\),\(3,1\),\(1,1\)\.\(2,1\),\(3,1\),\(1,1\)\.Since one\-part monomial\-symmetric functions are power sums, this pointed to low\-order power\-sum structure in the component not captured by the CP fit\.

As a diagnostic, we formed the diagonal coordinate proxy

ζλ=bλ2​‖mλ‖2,emp2\.\\zeta\_\{\\lambda\}=b\_\{\\lambda\}^\{2\}\\\|m\_\{\\lambda\}\\\|\_\{2,\\mathrm\{emp\}\}^\{2\}\.Approximately83%83\\%of this proxy was associated with one\-body partitions and17%17\\%with two\-body partitions, with negligible proxy weight at higher body counts\. This must not be interpreted as an orthogonalL2L^\{2\}\-energy decomposition:

∑λζλ=6\.71×106,\\sum\_\{\\lambda\}\\zeta\_\{\\lambda\}=6\.71\\times 10^\{6\},whereas the empirical squared residual norm was only1\.431\.43, revealing strong cancellation in the non\-orthogonal partition basis\.

The diagnostic nevertheless suggested that the residual was concentrated in structurally simple low\-order directions rather than in a diffuse high\-body tail\. Since such components are in principle representable within a sufficiently flexible CP family, the observed rank saturation pointed more naturally to an optimization or conditioning limitation than to an immediate expressivity obstruction\. One plausible explanation, explored in subsequent experiments, was that these directions were difficult to resolve by gradient\-based fitting against Monte Carlo noise\.

#### 5\.3\.4Experiment 4: explicit power\-sum corrections

The residual signature motivated the augmented model

FCP\+P=FCP\+β0\+∑j=1Jβj​P~j,P~j=Pj/𝔼⁡\[Pj\]\.F\_\{\\rm CP\+P\}=F\_\{\\rm CP\}\+\\beta\_\{0\}\+\\sum\_\{j=1\}^\{J\}\\beta\_\{j\}\\widetilde\{P\}\_\{j\},\\qquad\\widetilde\{P\}\_\{j\}=P\_\{j\}/\\mathbb\{E\}\[P\_\{j\}\]\.This experiment did not yield a reliable acceleration\. The revised implementation failed first to reproduce its own CP\-only baseline, the Monte Carlo estimator forJJbecame high variance for some augmented functions, and the explicit power\-sum branch overlapped strongly with directions already represented by the CP factors\. Consequently, the small observed changes with increasingJJcould not be separated cleanly from optimization and implementation signal\.

#### 5\.3\.5Experiment 5: body\-restricted partition spaces

The final pre\-separable experiment restricted the exact polynomial trial space simultaneously by total degree and partition length, or “body count”\. For

𝒱\(\{Db\}\)=span\{mλ:ℓ\(λ\)=b,\|λ\|≤Db\},\\mathcal\{V\}\(\\\{D\_\{b\}\\\}\)=\\operatorname\{span\}\\\{m\_\{\\lambda\}:\\ell\(\\lambda\)=b,\\ \|\\lambda\|\\leq D\_\{b\}\\\},matrix entries were assembled by enumerating coalescences of the positive parts of two partitions rather than their full permutation orbits\. The restriction is therefore specified by a degree capDbD\_\{b\}for each body countbb; atk=50k=50no run includedb\>5b\>5, so every space reported below is a genuine body\-restricted subspace of the full monomial\-symmetric space of the same total degree\. Atk=5k=5, whereℓ⁡\(λ\)≤k=5\\ell\(\\lambda\)\\leq k=5holds automatically, the unrestricted degree\-1010space contained113113partitions and reproduced

M5≥2\.0071443298,M\_\{5\}\\geq 2\.0071443298,approaching Bogaert’s Krylov value2\.00714514442\.0071451444\. Restricting to partition length at most two reduced the space to3636basis elements and gave a strictly weaker value, confirming the coalescence construction while showing that low\-body truncation alone does not recover the full small\-kkoptimizer\.

Atk=50k=50, exact matrix assembly remained feasible, but the generalized eigensolve became the dominant limitation\. Rank\-revealing whitening retained only a fraction of the nominal Gram modes; for a common degree cap of1515,1818,2020and2222on every body count the reductions were

408→191,769→278,1125→344,1601→421\.408\\to 191,\\qquad 769\\to 278,\\qquad 1125\\to 344,\\qquad 1601\\to 421\.The retained block consistently approached the imposed conditioning threshold,κ∼1012\\kappa\\sim 10^\{12\}, indicating that the effective ceiling was set by float64 numerical resolution rather than by the nominal size of the partition space\. This effect is visible in the non\-monotone certified ladder obtained from these increasing nominal spaces:

3\.36957,3\.36618,3\.36800,3\.36435\.3\.36957,\\qquad 3\.36618,\\qquad 3\.36800,\\qquad 3\.36435\.This does not contradict Rayleigh–Ritz monotonicity: the nominal trial spaces are nested, whereas the independently whitened effective subspaces need not be\. Exact reevaluation still certifies the returned trial vector, but cannot restore directions discarded before the eigensolve\.

The precision dependence is particularly clear in a matched408408\-dimensional calculation at degree cap1515\. Retaining the complete pencil at high precision gave the certified value

M50≥3\.6347,M\_\{50\}\\geq 3\.6347,whereas float64 spectral whitening retained only191191modes and yielded3\.36963\.3696\. Clearly, the lost variational gain was already present in the polynomial trial space but fell below the numerical resolution of the whitened Gram pencil\. The loss is strongly space\-dependent rather than systematic: on the441441\-dimensional one–two\-body space the high\-precision solve gave3\.18620923\.1862092against3\.18620743\.1862074from134134whitened modes, so there the discarded directions carried almost no variational weight\. A larger space admitting three\-body partitions up to degree2020similarly reached3\.39043\.3904with all678678modes retained at high precision, while larger float64\-whitened spaces did not systematically improve the bound\. Supplementary Table[1](https://arxiv.org/html/2609.30296#S5.T1)summarizes representative calculations\.

##### Why float64 whitening was insufficient in the monomial\-symmetric basis\.

One numerical difficulty in the early polynomial calculations was severe ill\-conditioning of the denominator Gram matrix in the monomial\-symmetric basis\. The matrix is positive definite on a linearly independent trial space, but after conversion to double precision its smallest eigenvalues can fall below the numerical resolution of the eigensolver\. Whitening then became problematic because the transformation

W=Ukeepdiag\(σi−1/2\)W=U\_\{\\mathrm\{keep\}\}\\operatorname\{diag\}\(\\sigma\_\{i\}^\{\-1/2\}\)amplified perturbations associated with small Gram eigenvalues, while directions below the chosen threshold were removed entirely\. So, the optimization is no longer performed over the nominal polynomial space, but only over its numerically resolved spectral image\. Exact rational re\-evaluation of the returned Ritz vector still yields a valid lower bound for that particular vector, but it cannot recover variationally relevant directions that were discarded during whitening\. With increasingkk, larger float64\-whitened spaces could produce weaker certified bounds than smaller spaces solved at high precision, demonstrating the numerical precision bottleneck\. This motivated representations with Gram structure without spectral truncation\.

Supplementary Table 1:Representativek=50k=50partition\-space calculations\. The trial space is specified by a degree cap per body countbb\(partition length\); no run includedb\>5b\>5\. Certified values refer to the returned explicit trial function\. High\-precision solves retain the full trial space, whereas float64 whitening removes Gram directions below the numerical resolution threshold\.Note\.Bogaert reportedM50≥3\.9358660346M\_\{50\}\\geq 3\.9358660346\.

The experiment therefore identified two distinct obstacles in the monomial\-symmetric representation: rapid growth of the partition basis and loss of numerically resolvable Gram directions\. The limiting issue atk=50k=50was not expressive power of the polynomial space itself, but the combination of basis growth and near\-degeneracy\. This motivated the move to representations whose complexity is controlled before the generalized eigensolve and whose structure is more directly compatible with exact certification\.

Full symmetric\-polynomial spaceexpressive andL2L^\{2\}\-densePartition growthdim𝒱D∼pk​\(D\)\\dim\\mathcal\{V\}\_\{D\}\\sim p\_\{k\}\(D\)Gram near\-degeneracyloss of float64\-resolved modesPracticalk=50k=50wall:large nominal space, limited effective rankPivot to compact separable /product\-structured representationsFigure 3:Scaling obstruction of the monomial\-symmetric representation\. Increasing polynomial degree expands an expressive symmetric trial space, but atk=50k=50this produces both rapid partition growth and increasingly ill\-conditioned Gram pencils\. The resulting loss of numerically resolved directions motivates the later separable representation\.

#### 5\.3\.6Consequence for the final representation

The abovementioned five experiments ruled out several superficially natural routes: sparse partition support, very\-low\-rank symmetric CP compression, unqualified interpretation of residual coefficients, additive power\-sum patches, and brute\-force growth of ill\-conditioned polynomial spaces\. The common failure mode was not lack of a strong candidate function but loss of structure during representation or optimization\. This motivated the separable\-channel formulation used in the main Methods, in which a small dictionary of one\-dimensional functions is optimized directly, the generalized Ritz coefficients are profiled out, and independent certification operates only after the discovered function has been reduced to an explicit mathematical representation\.

Scripts for the exploratory analyses described above are not included in NeuralCert\. Together with additional exploratory outputs including full residual rankings, body\-count and parity summaries, and repeated fit diagnostics, they are available upon request\.

### 5\.4Neural channel model

The discovery stage uses the symmetric separable representation

Fθ,c​\(t1,…,tk\)=∑j=1mcj​∏i=1kgθ,j​\(ti\)\.F\_\{\\theta,c\}\(t\_\{1\},\\ldots,t\_\{k\}\)=\\sum\_\{j=1\}^\{m\}c\_\{j\}\\prod\_\{i=1\}^\{k\}g\_\{\\theta,j\}\(t\_\{i\}\)\.The one\-dimensional channels are written

gθ,j​\(x\)=rθ,j​\(x\)​e−ρj​x,ρj\>0,g\_\{\\theta,j\}\(x\)=r\_\{\\theta,j\}\(x\)e^\{\-\\rho\_\{j\}x\},\\qquad\\rho\_\{j\}\>0,with trainable rates initialized across a broad geometric range\. The residual factors are produced jointly by a shared multilayer perceptron with channel\-specific outputs and input

For the positive\-channel model,

rθ,j​\(x\)=exp⁡\(fθ,j​\(x,e−k​x\)\),gθ,j​\(x\)=exp⁡\(fθ,j​\(x,e−k​x\)−ρj​x\)\.r\_\{\\theta,j\}\(x\)=\\exp\\\!\\left\(f\_\{\\theta,j\}\(x,e^\{\-kx\}\)\\right\),\\qquad g\_\{\\theta,j\}\(x\)=\\exp\\\!\\left\(f\_\{\\theta,j\}\(x,e^\{\-kx\}\)\-\\rho\_\{j\}x\\right\)\.Positivity permits cancellation\-free log\-domain convolution\. The architecture differs from a conventional DeepSets model: permutation invariance follows from the product structure of the separable expansion, while the neural network supplies a dictionary of one\-dimensional factors\.

Listing 1:Positive neural channel dictionary\. A shared network produces the residual factors, while the exponential envelopes are represented analytically\.PROCEDURECHANNELS\(x,k,theta,eta,rho\_max\)

rho<\-MIN\(SOFTPLUS\(eta\),rho\_max\)

features<\-STACK\(x,EXP\(\-k\*x\)\)

log\_raw<\-SHARED\_MLP\_theta\(features\)

g<\-EXP\(log\_raw\-OUTER\(x,rho\)\)

RETURNg,log\_raw,rho

END

FUNCTIONTRIAL\(t,c,theta\)

RETURNSUM\_jc\[j\]\*PRODUCT\_iCHANNEL\_j\(t\[i\],theta\)

END

### 5\.5Factored convolution powers and Gauss–Jacobi recursion

For a channel pair\(j,ℓ\)\(j,\\ell\), write

hj​ℓ​\(x\)=rj​ℓ​\(x\)​e−aj​ℓ​x,rj​ℓ​\(x\)=rθ,j​\(x\)​rθ,ℓ​\(x\),aj​ℓ=ρj\+ρℓ\.h\_\{j\\ell\}\(x\)=r\_\{j\\ell\}\(x\)e^\{\-a\_\{j\\ell\}x\},\\qquad r\_\{j\\ell\}\(x\)=r\_\{\\theta,j\}\(x\)r\_\{\\theta,\\ell\}\(x\),\\qquad a\_\{j\\ell\}=\\rho\_\{j\}\+\\rho\_\{\\ell\}\.Every convolution power is represented as

νp\(j​ℓ\)​\(s\):=hj​ℓ∗p​\(s\)=sp−1​e−aj​ℓ​s​eσp\(j​ℓ\)​ψp\(j​ℓ\)​\(s\)\.\\nu^\{\(j\\ell\)\}\_\{p\}\(s\):=h\_\{j\\ell\}^\{\*p\}\(s\)=s^\{p\-1\}e^\{\-a\_\{j\\ell\}s\}e^\{\\sigma^\{\(j\\ell\)\}\_\{p\}\}\\psi^\{\(j\\ell\)\}\_\{p\}\(s\)\.The initialization is

σ1\(j​ℓ\)=0,ψ1\(j​ℓ\)​\(s\)=rj​ℓ​\(s\)\.\\sigma^\{\(j\\ell\)\}\_\{1\}=0,\\qquad\\psi^\{\(j\\ell\)\}\_\{1\}\(s\)=r\_\{j\\ell\}\(s\)\.We suppress the channel\-pair superscript below\.

For positive integersppandqq, substitution into

νp\+q​\(s\)=∫0sνp​\(x\)​νq​\(s−x\)​𝑑x\\nu\_\{p\+q\}\(s\)=\\int\_\{0\}^\{s\}\\nu\_\{p\}\(x\)\\nu\_\{q\}\(s\-x\)\\,dxand the change of variablesx=s​ux=sugive

νp\+q​\(s\)=sp\+q−1​e−a​s​eσp\+σq​∫01ψp​\(s​u\)​ψq​\(s⁡\(1−u\)\)​up−1​\(1−u\)q−1​𝑑u\.\\nu\_\{p\+q\}\(s\)=s^\{p\+q\-1\}e^\{\-as\}e^\{\\sigma\_\{p\}\+\\sigma\_\{q\}\}\\int\_\{0\}^\{1\}\\psi\_\{p\}\(su\)\\psi\_\{q\}\(s\(1\-u\)\)u^\{p\-1\}\(1\-u\)^\{q\-1\}\\,du\.Let

B⁡\(p,q\)=∫01up−1​\(1−u\)q−1​𝑑u=Γ⁡\(p\)​Γ​\(q\)Γ⁡\(p\+q\)\.B\(p,q\)=\\int\_\{0\}^\{1\}u^\{p\-1\}\(1\-u\)^\{q\-1\}\\,du=\\frac\{\\Gamma\(p\)\\Gamma\(q\)\}\{\\Gamma\(p\+q\)\}\.We retain the beta factor in the residual recursion and therefore use

σp\+q=σp\+σq,\\sigma\_\{p\+q\}=\\sigma\_\{p\}\+\\sigma\_\{q\},together with

ψp\+q​\(s\)=B⁡\(p,q\)​∫01ψp​\(s​u\)​ψq​\(s⁡\(1−u\)\)​up−1​\(1−u\)q−1B⁡\(p,q\)​𝑑u\.\\psi\_\{p\+q\}\(s\)=B\(p,q\)\\int\_\{0\}^\{1\}\\psi\_\{p\}\(su\)\\psi\_\{q\}\(s\(1\-u\)\)\\frac\{u^\{p\-1\}\(1\-u\)^\{q\-1\}\}\{B\(p,q\)\}\\,du\.Thus the integral is an expectation with respect to the normalizedBeta⁡\(p,q\)\\operatorname\{Beta\}\(p,q\)density\. The factorB⁡\(p,q\)B\(p,q\)occurs once, inψp\+q\\psi\_\{p\+q\}, and is not also included inσp\+q\\sigma\_\{p\+q\}\.

The expectation is evaluated by Gauss–Jacobi quadrature:

ψp\+q​\(s\)≈B⁡\(p,q\)​∑r=1nqw^r\(p,q\)​ψp​\(s​ur\(p,q\)\)​ψq​\(s⁡\(1−ur\(p,q\)\)\),\\psi\_\{p\+q\}\(s\)\\approx B\(p,q\)\\sum\_\{r=1\}^\{n\_\{q\}\}\\widehat\{w\}\_\{r\}^\{\(p,q\)\}\\psi\_\{p\}\(su\_\{r\}^\{\(p,q\)\}\)\\psi\_\{q\}\(s\(1\-u\_\{r\}^\{\(p,q\)\}\)\),where

0<ur\(p,q\)<1,w^r\(p,q\)\>0,∑r=1nqw^r\(p,q\)=1\.0<u\_\{r\}^\{\(p,q\)\}<1,\\qquad\\widehat\{w\}\_\{r\}^\{\(p,q\)\}\>0,\\qquad\\sum\_\{r=1\}^\{n\_\{q\}\}\\widehat\{w\}\_\{r\}^\{\(p,q\)\}=1\.The nodes and normalized weights incorporate the beta\-weight concentration associated with the convolution ordersppandqq\.

An optional numerical normalization can transfer residual amplitude into the scalar scale\. Ifψ~p\+q\\widetilde\{\\psi\}\_\{p\+q\}denotes the residual produced by the recursion above, choose a positive scalarCp\+qC\_\{p\+q\}, independent ofss, and set

ψp\+q​\(s\)=ψ~p\+q​\(s\)Cp\+q,σp\+q=σp\+σq\+log⁡Cp\+q\.\\psi\_\{p\+q\}\(s\)=\\frac\{\\widetilde\{\\psi\}\_\{p\+q\}\(s\)\}\{C\_\{p\+q\}\},\\qquad\\sigma\_\{p\+q\}=\\sigma\_\{p\}\+\\sigma\_\{q\}\+\\log C\_\{p\+q\}\.This leaves the represented convolution power unchanged\. TakingCp\+q=1C\_\{p\+q\}=1recovers the unnormalized recursion\. For positive residuals, the factorB⁡\(p,q\)B\(p,q\)is handled throughlog⁡B⁡\(p,q\)\\log B\(p,q\)in the log\-domain recursion; any additional normalization subtractslog⁡Cp\+q\\log C\_\{p\+q\}from the log residual and adds it toσp\+q\\sigma\_\{p\+q\}\.

The powerh∗\(k−1\)h^\{\*\(k\-1\)\}is assembled by a binary addition chain, reducing convolution depth fromk−2k\-2sequential stages toO⁡\(log⁡k\)O\(\\log k\)\.

### 5\.6Log\-domain recursion

For positive channels defineξp​\(s\)=log⁡ψp​\(s\)\\xi\_\{p\}\(s\)=\\log\\psi\_\{p\}\(s\)\. The recursion becomes

ξp\+q​\(s\)=log⁡B⁡\(p,q\)\+logsumexpr⁡\[log⁡w^r\(p,q\)\+ξp​\(s​ur\(p,q\)\)\+ξq​\(s⁡\(1−ur\(p,q\)\)\)\]\.\\xi\_\{p\+q\}\(s\)=\\log B\(p,q\)\+\\operatorname\{logsumexp\}\_\{r\}\\left\[\\log\\widehat\{w\}\_\{r\}^\{\(p,q\)\}\+\\xi\_\{p\}\(su\_\{r\}^\{\(p,q\)\}\)\+\\xi\_\{q\}\(s\(1\-u\_\{r\}^\{\(p,q\)\}\)\)\\right\]\.Values that would underflow in linear arithmetic remain representable as large negative logarithms, and the positive quadrature sum is evaluated without cancellation\. This is important because loss of a small component during an early convolution can propagate into entire entries of the subsequent Gram matrices\.

Listing 2:Binary\-chain evaluation of the factored convolution power in logarithmic coordinates\.PROCEDURELOG\_CONVOLUTION\_POWER\(log\_raw,j,l,k\)

xi\[1\]\(s\)<\-log\_raw\_j\(s\)\+log\_raw\_l\(s\)

FOREACH\(p,q\)INBINARY\_ADDITION\_CHAIN\(k\-1\)

\(u,w\)<\-NORMALIZED\_GAUSS\_JACOBI\(p,q\)

FOREACHrepresentationnodes

a<\-LOG\_POSITIVE\_INTERPOLATE\(xi\[p\],s\*u\)

IFq=1

b<\-EVALUATE\_BASE\_PAIR\_EXACTLY\(s\*\(1\-u\),j,l\)

ELSE

b<\-LOG\_POSITIVE\_INTERPOLATE\(xi\[q\],s\*\(1\-u\)\)

END

xi\[p\+q\]\(s\)<\-LOG\_BETA\(p,q\)

\+LOGSUMEXP\_r\(LOG\(w\[r\]\)\+a\[r\]\+b\[r\]\)

END

END

RETURNxi\[k\-1\]

END

### 5\.7Representation and integration grids

Residual functions are stored on an endpoint\-inclusive Chebyshev–Lobatto grid

0=s0<s1<⋯<sN−1=L,L=1\+ε\.0=s\_\{0\}<s\_\{1\}<\\cdots<s\_\{N\-1\}=L,\\qquad L=1\+\\varepsilon\.Every convolution targetsi​urs\_\{i\}u\_\{r\}orsi​\(1−ur\)s\_\{i\}\(1\-u\_\{r\}\)lies in\[0,si\]\[0,s\_\{i\}\], so no extrapolation is required\. Four numerical point sets are kept distinct: the Chebyshev–Lobatto representation grid; Gauss–Jacobi rules for the beta\-weighted convolution; Gauss–Legendre rules for the primitives

Gj​\(y\)=∫0ygj​\(x\)​𝑑x,Hj​ℓ​\(y\)=∫0ygj​\(x\)​gℓ​\(x\)​𝑑x;G\_\{j\}\(y\)=\\int\_\{0\}^\{y\}g\_\{j\}\(x\)\\,dx,\\qquad H\_\{j\\ell\}\(y\)=\\int\_\{0\}^\{y\}g\_\{j\}\(x\)g\_\{\\ell\}\(x\)\\,dx;and a geometrically graded outer mesh for the final entries ofAAandBB\. The latter must resolve the pair\-dependent factorxk−2​e−aj​ℓ​xx^\{k\-2\}e^\{\-a\_\{j\\ell\}x\}\.

Listing 3:Construction of the four numerical point sets used by the discovery evaluator\.PROCEDUREBUILD\_GRIDS\(k,epsilon,N,n\_jacobi,n\_inner,n\_outer\)

L<\-1\+epsilon;U<\-1\-epsilon

s\[i\]<\-L\*\(1\-COS\(pi\*i/\(N\-1\)\)\)/2,i=0,\.\.\.,N\-1

FOREACH\(p,q\)INBINARY\_ADDITION\_CHAIN\(k\-1\)

jacobi\[p,q\]<\-NORMALIZED\_GAUSS\_JACOBI\(n\_jacobi,p,q\)

END

inner<\-GAUSS\_LEGENDRE\(n\_inner,interval=\[0,1\]\)

x\_star<\-MIN\(X,\(k\-2\)/a\_ref\)

edges<\-GEOMETRIC\_MESH\(x\_star/32,X\)

outer<\-COMPOSITE\_GAUSS\_LEGENDRE\(edges,n\_outer\)

RETURNs,jacobi,inner,outer

END

### 5\.8Preservation of Gram structure

The exact denominator matrix is a Gram matrix and therefore satisfiesA⪰0A\\succeq 0\. To preserve this structure numerically, a targetz∈\[si,si\+1\]z\\in\[s\_\{i\},s\_\{i\+1\}\]is interpolated using the positive stencil

ψ⁡\(z\)≈ωi​\(z\)​ψ​\(si\)\+ωi\+1​\(z\)​ψ​\(si\+1\),ωi,ωi\+1≥0,ωi\+ωi\+1=1\.\\psi\(z\)\\approx\\omega\_\{i\}\(z\)\\psi\(s\_\{i\}\)\+\\omega\_\{i\+1\}\(z\)\\psi\(s\_\{i\+1\}\),\\qquad\\omega\_\{i\},\\omega\_\{i\+1\}\\geq 0,\\qquad\\omega\_\{i\}\+\\omega\_\{i\+1\}=1\.In log\-space, this interpolation is evaluated by a two\-term log\-sum\-exp\. Interpolatingψ\\psiitself, rather thanlog⁡ψ\\log\\psi, is what makes the stencil multilinear and hence compatible with the Gram form\. When the inner quadrature nodes are common to all channel pairs, the shared positive\-weight outer rule yields the discrete Gram representation

Aj​ℓ\(N\)=∑αwα​Φα​j​Φα​ℓ,wα\>0,A^\{\(N\)\}\_\{j\\ell\}=\\sum\_\{\\alpha\}w\_\{\\alpha\}\\Phi\_\{\\alpha j\}\\Phi\_\{\\alpha\\ell\},\\qquad w\_\{\\alpha\}\>0,and henceA\(N\)⪰0A^\{\(N\)\}\\succeq 0up to floating\-point round\-off\. The shared mesh therefore preserves a common discrete Gram representation rather than merely improving entrywise quadrature accuracy\.

In the saturated regime, where the substitution used for the channel primitives places the inner nodes at pair\-dependent positions, the common discrete Gram representation is preserved only up to inner\-quadrature error\. Numerical adequacy is then assessed using the positive\-semidefiniteness, two\-grid and refinement diagnostics described below; these diagnostics are safeguards rather than rigorous bounds on the quadrature error\. Since the reported bounds come from the exact certifier and not from the discrete pencil, a failure of this structure degrades the discovery objective rather than the validity of any certificate\.

Listing 4:Log\-domain assembly of one pair of Gram entries\. With common inner quadrature nodes, the shared positive\-weight outer rule gives an exact discrete Gram representation; in the saturated regime, pair\-dependent inner quadrature preserves this structure only up to inner\-quadrature error\.FUNCTIONLOG\_INTERPOLATE\(z,xi\)

\(i,i\+1,omega0,omega1\)<\-POSITIVE\_TWO\_POINT\_STENCIL\(z\)

RETURNLOGADDEXP\(LOG\(omega0\)\+xi\[i\],LOG\(omega1\)\+xi\[i\+1\]\)

END

PROCEDUREGRAM\_PAIR\(j,l\)

a<\-rho\[j\]\+rho\[l\]

xi<\-LOG\_CONVOLUTION\_POWER\(log\_raw,j,l,k\)

log\_Ghat,log\_Hhat<\-GAUSS\_LEGENDRE\_PRIMITIVES\(j,l\)

log\_A\[j,l\]<\-LOGSUMEXP\_x\(\(k\-2\)\*LOG\(x\)\-a\*x

\+LOG\_INTERPOLATE\(x,xi\)

\+LOG\(L\-x\)\+LOG\_INTERPOLATE\(L\-x,log\_Hhat\)\+LOG\(w\_x\)\)

log\_B\[j,l\]<\-LOGSUMEXP\_x\(\(k\-2\)\*LOG\(x\)\-a\*x

\+LOG\_INTERPOLATE\(x,xi\)\+2\*LOG\(L\-x\)

\+LOG\_INTERPOLATE\(L\-x,log\_Ghat\_j\)

\+LOG\_INTERPOLATE\(L\-x,log\_Ghat\_l\)\+LOG\(w\_x\)\)

RETURNlog\_A\[j,l\],log\_B\[j,l\]

END

### 5\.9Diagonal equilibration and rank\-revealing Ritz solution

With

D=diag⁡\(A11,…,Am​m\),D=\\operatorname\{diag\}\(A\_\{11\},\\ldots,A\_\{mm\}\),the generalized Rayleigh quotient is invariant under

A↦D−1/2AD−1/2,B↦D−1/2BD−1/2\.A\\mapsto D^\{\-1/2\}AD^\{\-1/2\},\\qquad B\\mapsto D^\{\-1/2\}BD^\{\-1/2\}\.After equilibration, write

A~=U​diag⁡\(α1,…,αm\)​U⊤\.\\widetilde\{A\}=U\\operatorname\{diag\}\(\\alpha\_\{1\},\\ldots,\\alpha\_\{m\}\)U^\{\\top\}\.Directions satisfying

αi≤τrank​maxj​αj\\alpha\_\{i\}\\leq\\tau\_\{\\rm rank\}\\max\_\{j\}\\alpha\_\{j\}are removed before whitening\. On the retained space,

W=Ukeepdiag\(αi−1/2\),W=U\_\{\\rm keep\}\\operatorname\{diag\}\(\\alpha\_\{i\}^\{\-1/2\}\),and the leading eigenpair ofW⊤​B~​WW^\{\\top\}\\widetilde\{B\}Wdetermines the profiled quotient and mixing vector\. No ridge term is introduced; rank truncation is treated as a numerical\-resolution decision and monitored by rank\-stability tests\.

Listing 5:Diagonally equilibrated, rank\-revealing generalized Ritz solve\.PROCEDURETOP\_RITZ\(log\_A,log\_B,tau\_rank\)

d\[j\]<\-log\_A\[j,j\]

Ahat\[j,l\]<\-EXP\(log\_A\[j,l\]\-\(d\[j\]\+d\[l\]\)/2\)

Bhat\[j,l\]<\-EXP\(log\_B\[j,l\]\-\(d\[j\]\+d\[l\]\)/2\)

\(alpha,U\)<\-SYMMETRIC\_EIGENDECOMPOSITION\(Ahat\)

keep<\-\{i:alpha\[i\]\>tau\_rank\*MAX\(alpha\)\}

REQUIREkeepisnonempty

W<\-U\[:,keep\]\*DIAG\(alpha\[keep\]^\(\-1/2\)\)

\(lambda,z\)<\-LEADING\_EIGENPAIR\(TRANSPOSE\(W\)\*Bhat\*W\)

v<\-W\*z

v<\-v/SQRT\(TRANSPOSE\(v\)\*Ahat\*v\)

RETURNk\*lambda,v,CARDINALITY\(keep\),d

END

### 5\.10Hellmann–Feynman gradients and optimization

Numerical solution of the generalized Ritz problem may require removal of directions associated with numerically negligible eigenvalues ofA⁡\(θ\)A\(\\theta\)\. Let

𝒮⁡\(θ\)\\mathcal\{S\}\(\\theta\)denote the retained coefficient subspace after this rank\-revealing truncation, and define the corresponding numerical Ritz value by

R^​\(θ\)=maxv∈𝒮⁡\(θ\)v≠0⁡k​v⊤​B​\(θ\)​vv⊤​A​\(θ\)​v\.\\widehat\{R\}\(\\theta\)=\\max\_\{\\begin\{subarray\}\{c\}v\\in\\mathcal\{S\}\(\\theta\)\\\\ v\\neq 0\\end\{subarray\}\}k\\frac\{v^\{\\top\}B\(\\theta\)v\}\{v^\{\\top\}A\(\\theta\)v\}\.Because𝒮⁡\(θ\)\\mathcal\{S\}\(\\theta\)may change withθ\\theta, the mapR^​\(θ\)\\widehat\{R\}\(\\theta\)need not be differentiable when a singular direction crosses the truncation threshold\. Moreover, the standard Hellmann–Feynman or envelope identity does not automatically apply across such a parameter\-dependent change of retained subspace\.

We therefore freeze the retained subspace during each differentiable optimization block\. At outer iterationss, let

𝒮s:=𝒮⁡\(θs\)\\mathcal\{S\}\_\{s\}:=\\mathcal\{S\}\(\\theta\_\{s\}\)and define the fixed\-subspace profile

Rs​\(θ\)=maxv∈𝒮sv≠0⁡qv​\(θ\),qv​\(θ\)=k​v⊤​B​\(θ\)​vv⊤​A​\(θ\)​v\.R\_\{s\}\(\\theta\)=\\max\_\{\\begin\{subarray\}\{c\}v\\in\\mathcal\{S\}\_\{s\}\\\\ v\\neq 0\\end\{subarray\}\}q\_\{v\}\(\\theta\),\\qquad q\_\{v\}\(\\theta\)=k\\frac\{v^\{\\top\}B\(\\theta\)v\}\{v^\{\\top\}A\(\\theta\)v\}\.Letvs∈𝒮sv\_\{s\}\\in\\mathcal\{S\}\_\{s\}be the normalized leading generalized Ritz vector at the anchor point:

B⁡\(θs\)​vs=λs​A​\(θs\)​vson​𝒮s,vs⊤​A​\(θs\)​vs=1\.B\(\\theta\_\{s\}\)v\_\{s\}=\\lambda\_\{s\}A\(\\theta\_\{s\}\)v\_\{s\}\\quad\\text\{on \}\\mathcal\{S\}\_\{s\},\\qquad v\_\{s\}^\{\\top\}A\(\\theta\_\{s\}\)v\_\{s\}=1\.Provided thatA⁡\(θ\)A\(\\theta\)remains positive definite on𝒮s\\mathcal\{S\}\_\{s\}in a neighbourhood ofθs\\theta\_\{s\}and that the leading restricted generalized eigenvalue is simple, the fixed\-subspace envelope identity gives

∇θRs​\(θs\)=∇θqvs​\(θs\)\.\\nabla\_\{\\theta\}R\_\{s\}\(\\theta\_\{s\}\)=\\nabla\_\{\\theta\}q\_\{v\_\{s\}\}\(\\theta\_\{s\}\)\.Thus the Ritz vector may be frozen when differentiating the quotient, but only while the retained subspace is held fixed\.

For every fixedv∈𝒮sv\\in\\mathcal\{S\}\_\{s\},

qv​\(θ\)≤Rs​\(θ\),q\_\{v\}\(\\theta\)\\leq R\_\{s\}\(\\theta\),and at the anchor point

qvs​\(θs\)=Rs​\(θs\)\.q\_\{v\_\{s\}\}\(\\theta\_\{s\}\)=R\_\{s\}\(\\theta\_\{s\}\)\.Consequently, an inner step satisfying

qvs​\(θs\+1\)≥qvs​\(θs\)q\_\{v\_\{s\}\}\(\\theta\_\{s\+1\}\)\\geq q\_\{v\_\{s\}\}\(\\theta\_\{s\}\)obeys the fixed\-subspace minorize–maximize relation

Rs​\(θs\+1\)≥qvs​\(θs\+1\)≥qvs​\(θs\)=Rs​\(θs\)\.R\_\{s\}\(\\theta\_\{s\+1\}\)\\geq q\_\{v\_\{s\}\}\(\\theta\_\{s\+1\}\)\\geq q\_\{v\_\{s\}\}\(\\theta\_\{s\}\)=R\_\{s\}\(\\theta\_\{s\}\)\.This monotonicity statement concernsRsR\_\{s\}, with𝒮s\\mathcal\{S\}\_\{s\}fixed\. It does not by itself imply monotonicity of the adaptively truncated valueR^​\(θ\)\\widehat\{R\}\(\\theta\)after the numerical rank is recomputed\.

Neural channels are first optimized with Adam using learning\-rate warm\-up, cosine decay, and gradient clipping\. Candidate solutions are subsequently polished by block minorize–maximize L\-BFGS\. Within each such block, the retained subspace, truncation mask, and frozen Ritz vector are held fixed\. The rank\-revealing decomposition is recomputed only between blocks\.

After this recomputation, a candidate outer step is accepted only if the adaptively evaluated Ritz value satisfies

R^​\(θs\+1\)≥R^​\(θs\)−τacc,\\widehat\{R\}\(\\theta\_\{s\+1\}\)\\geq\\widehat\{R\}\(\\theta\_\{s\}\)\-\\tau\_\{\\mathrm\{acc\}\},whereτacc\\tau\_\{\\mathrm\{acc\}\}is a prescribed numerical acceptance tolerance\. If a change in retained rank causes this test to fail, the candidate is rejected and the step size or optimization block is restarted\. This acceptance test is an explicit numerical safeguard; it is not presented as a consequence of the Hellmann–Feynman identity\.

Finite\-difference gradient tests are performed with the retained subspace and truncation mask frozen\. A finite\-difference stencil that changes the retained numerical rank crosses a nonsmooth truncation boundary and is therefore not used as a test of the Hellmann–Feynman gradient\. Such cases are instead flagged as rank\-transition events\. Away from these events, the test checks the complete differentiation path through the neural channels, factored convolution recursion, integration, and matrix assembly\.

Optimization proceeds through increasing channel ranks

m1<⋯<mSm\_\{1\}<\\cdots<m\_\{S\}and representation resolutions

N1<⋯<NS\.N\_\{1\}<\\cdots<N\_\{S\}\.Existing channels are copied when the channel rank is increased, and new channels are initialized as perturbed duplicates\. Because pairwise assembly scales asm⁡\(m\+1\)/2m\(m\+1\)/2, more iterations are allocated to the smaller early models\.

Listing 6:Profiled optimization with a frozen Ritz vector\.FUNCTIONFROZEN\_QUOTIENT\(theta,v\)

\(A\_theta,B\_theta\)<\-ASSEMBLE\_GRAM\(theta\)

RETURNk\*\(TRANSPOSE\(v\)\*B\_theta\*v\)/\(TRANSPOSE\(v\)\*A\_theta\*v\)

END

PROCEDUREOPTIMIZE\_CHANNELS\(theta\)

FOREACHincreasingpair\(channelrankm,gridsizeN\)

GROW\_DICTIONARY\_BY\_PERTURBED\_DUPLICATES\(theta,m\)

FOREACHAdamiteration,withwarm\-upandcosinedecay

IFthisisareportingpoint

VALIDATE\_CURRENT\_STATE\(theta\)

END

\(\_,v\)<\-TOP\_RITZ\(ASSEMBLE\_GRAM\(theta\)\)

theta<\-CLIPPED\_ADAM\_ASCENT\(

FROZEN\_QUOTIENT\(theta,STOP\_GRADIENT\(v\)\)\)

END

END

REPEATforeachblockminorize\-maximizestep

\(\_,v\)<\-TOP\_RITZ\(ASSEMBLE\_GRAM\(theta\)\)

theta<\-LBFGS\_MAXIMIZE\(FROZEN\_QUOTIENT\(theta,STOP\_GRADIENT\(v\)\)\)

VALIDATE\_CURRENT\_STATE\(theta\)

UPDATEincumbentonlyifitstwo\-gridscoreimproves

END

RETURNtheta

END

### 5\.11Validation\-aware model selection

At each reporting point the candidate must satisfy:

1. 1\.agreement between the training objective and an independently recomputed Rayleigh quotient;
2. 2\.numerical positive semidefiniteness of the denominator matrix;
3. 3\.retention of a prescribed minimum fraction of the channel span after rank truncation;
4. 4\.stability of the leading quotient under a more conservative rank threshold;
5. 5\.compliance of every diagonal ratiok​Bj​j/Aj​jkB\_\{jj\}/A\_\{jj\}with the applicable rigorous upper bound;
6. 6\.compliance of the full quotient with the corresponding rigorous upper bound; and
7. 7\.reproduction on a finer representation grid within the prescribed snapshot tolerance\.

When two grids are used during training, candidates are ranked by

Rscore=min⁡\{RN,RNhi\}\.R\_\{\\rm score\}=\\min\\\{R\_\{N\},R\_\{N\_\{\\rm hi\}\}\\\}\.A substantial negative eigenvalue ofAAis treated as failure of the numerical representation rather than repaired by discarding the negative branch\. Rank\-stability tests separately guard against whitening directions below the accuracy of the assembled matrices\.

Listing 7:Validation and two\-grid model selection applied at each reporting point\.FUNCTIONRIGOROUS\_CEILING\(k,epsilon\)

IFepsilon=0

RETURNk\*LOG\(k\)/\(k\-1\)

ELSE

RETURNk\*LOG\(2\*k\-1\)/\(k\-1\)

END

END

PROCEDUREVALIDATE\_CURRENT\_STATE\(theta\)

R\_bound<\-RIGOROUS\_CEILING\(k,epsilon\)

\(R\_train,A,B,v,rank\)<\-INDEPENDENT\_RECOMPUTATION\(theta\)

REQUIREAGREES\(R\_train,FROZEN\_QUOTIENT\(theta,v\)\)

REQUIREMIN\_EIGENVALUE\(A\)\>=\-tau\_psd\*MAX\_EIGENVALUE\(A\)

REQUIRErank\>=required\_fraction\*NUMBER\_OF\_CHANNELS\(theta\)

REQUIRESTABLE\_UNDER\_STRONGER\_RANK\_TRUNCATION\(A,B\)

REQUIREEACHk\*B\[j,j\]/A\[j,j\]<=R\_bound

REQUIRER\_train<=R\_bound

R\_hi<\-RECOMPUTE\_ON\_GRID\(theta,1\.5\*N\)

REQUIRERELATIVE\_DIFFERENCE\(R\_train,R\_hi\)<=tau\_snapshot

score<\-MIN\(R\_train,R\_hi\)

STOREthetaonlyifscoreimprovesthevalidatedincumbent

RETURNscore

END

### 5\.12Closed\-form controls and representation refinement

Every run begins with closed\-form controls\. Forg⁡\(x\)=1g\(x\)=1andε=0\\varepsilon=0,

Rconst=2​kk\+1,R\_\{\\rm const\}=\\frac\{2k\}\{k\+1\},and forg⁡\(x\)=xg\(x\)=x,

Rx=3​k3​k\+1\.R\_\{x\}=\\frac\{3k\}\{3k\+1\}\.The first exercises convolution, outer integration and scaling; the second also probes interpolation of a nonconstant residual convolution shape\.

After training, candidates are evaluated on gridsNN,1\.25​N1\.25Nand1\.5​N1\.5Nand fitted to

RN=R∞−c​N−p\.R\_\{N\}=R\_\{\\infty\}\-cN^\{\-p\}\.The fitted limit is diagnostic only\. the export is marked unconverged when the sequence is noncontracting, violates a rigorous ceiling or fails the refinement tolerance\. No extrapolated discovery value is treated as a certified lower bound\.

The final separable discovery solver emerged through a sequence of numerical and structural corrections\. Supplementary Table[2](https://arxiv.org/html/2609.30296#S5.T2)summarizes the principal failure modes identified during development and the corresponding fixes\. These historical iterations are not separate proof components; they document how the numerical invariants enforced in the final solver were established\.

Supplementary Table[2](https://arxiv.org/html/2609.30296#S5.T2)lists the different development steps for the discovery script\.

Supplementary Table 2:Reconstructed development history of themaynard\_separable\_rayleigh\_v9\_factoredlineage \(neural discovery\)\. Version numbers map to successive saved iterations\.\#CorrectionCause / symptom1Baseline v9 \(factored\):νp=sp−1​e−a​s​eσ​ψp\\nu\_\{p\}=s^\{p\-1\}e^\{\-as\}e^\{\\sigma\}\\psi\_\{p\}per pair; stableinv\_softplus; per\-pair max\-rescale; adaptive log\-space GL outer windows; diagonal preconditioning \(diag⁡A=1\\operatorname\{diag\}A=1\);parts\(\)protocol\.Foundation\.inv\_softplusfixessoftplus−1=log⁡\(expm1\)→\+∞\\mathrm\{softplus\}^\{\-1\}=\\log\(\\mathrm\{expm1\}\)\\to\+\\inftyfor rate≥710\\geq 710\(allk≥355k\\geq 355\); per\-pair rescale replaces a global max that flushed small pairs to round\-off\.2σ\\sigmadouble\-count fix: divide exact base re\-evaluations bye−σ1e^\{\-\\sigma\_\{1\}\}\.Exact base evals carriedeσ1e^\{\\sigma\_\{1\}\}whileσ1\\sigma\_\{1\}was also added innew\_sig; the spurious factor is not of Gram form⇒\\RightarrowindefiniteAA\.3train\(\)validation\-ordering fix: validate beforeopt\.step\(\)\.RtrainR\_\{\\text\{train\}\}read pre\-step,report\(\)post\-step;agreemeasured the Adam step, rejecting nearly all snapshots \(fallback to iter 700\)\.4Log\-space channel protocol:log\_parts;\-\-channel\-sign positive/free\.Prerequisite for the log\-space recursion; needs sign\-definite channels\.5Log\-spaceξ\\xirecursion:ξp\+q=log⁡B⁡\(p,q\)\+logsumexp⁡\(log⁡w^\+ξp\+ξq\)\\xi\_\{p\+q\}=\\log B\(p,q\)\+\\operatorname\{logsumexp\}\(\\log\\hat\{w\}\+\\xi\_\{p\}\+\\xi\_\{q\}\)\.The\(min⁡ψ1/max⁡ψ1\)k−1\(\\min\\psi\_\{1\}/\\max\\psi\_\{1\}\)^\{k\-1\}collapse droveψ\\psipast the float64 floor; round\-off put random signs on diagonal pairs \(ψ≥0\\psi\\geq 0\), makingAAindefinite\.6Shared graded outer mesh \(one geometric Gauss rule for all pairs\)\.Per\-pair windows gave each entry its own rule⇒A\\Rightarrow Athe Gram matrix of nothing;λmin/λmax≈−1\\lambda\_\{\\min\}/\\lambda\_\{\\max\}\\approx\-1and did not improve under refinement\.7Log\-singularity fix:G⁡\(y\)=y​G^​\(y\)G\(y\)=y\\,\\hat\{G\}\(y\); interpolatelog⁡G^\\log\\hat\{G\}, reattachlog⁡y\\log yanalytically\.G⁡\(y\)G\(y\)vanishes linearly⇒log⁡G→−∞\\Rightarrow\\log G\\to\-\\infty\(clamped\); interpolating that oscillated⇒B≈1045\\Rightarrow B\\approx 10^\{45\}\.8Gram\-preserving positive interpolation: two\-point stencil, weights≥0\\geq 0, vialogaddexp;\-\-interp positive/spectral\.Interpolating the log gives∏ψiMi\\prod\\psi\_\{i\}^\{M\_\{i\}\}\(geometric\), not∑Mi​ψi\\sum M\_\{i\}\\psi\_\{i\}\(linear\)⇒\\Rightarrownot multilinear⇒\\Rightarrownot Gram⇒\\RightarrowindefiniteAA\.9Hard gates \(PSD,R≤kR\\leq k, gradient\) \+ removed false “PSD by construction” \+ parameter\-independent mesh\.R=118,979,5\.2×108R=118,\\,979,\\,5\.2\\times 10^\{8\}were optimised toward; positive\-entrywise≠\\neqPSD; moving mesh gave a spurious HF\-gradient gap\.10Resolution diagnostics \(ξ\\xispan/jump\),g=xg=xpreflight \(3​k/\(3​k\+1\)3k/\(3k\{\+\}1\)\), interp\-dependentnrepn\_\{\\text\{rep\}\}, sharp boundkk−1​ln⁡k\\tfrac\{k\}\{k\-1\}\\ln k, Richardson gate,NNprovenance\.g≡1g\\equiv 1is exact for anyNNin positive mode \(blind to under\-resolution\);R≤kR\\leq ktoo loose to catch a6→496\\to 49climb\.11Refinement\-stable snapshots \+ strikes \+1\.5×1\.5\\timesvalidation twin\.Optimiser maximisesRtrue\+error⁡\(θ\)R\_\{\\text\{true\}\}\+\\mathrm\{error\}\(\\theta\); refinement\-instability is the exploitation signature and must gate snapshots\.12Two\-tier snapshots \+min⁡\(RN,R1\.5​N\)\\min\(R\_\{N\},R\_\{1\.5N\}\)ranking \+ fitted\-order Richardson \+ grid continuation \+ result reorder\.One tolerance did two jobs and rejected every honest state; trained\-channel orderp≈1\.6≠2p\\approx 1\.6\\neq 2, so fixed\-order Richardson mis\-extrapolated\.13ε\>0\\varepsilon\>0bound fix:Rbound=kR\_\{\\text\{bound\}\}=kforε\>0\\varepsilon\>0; tight bound only forε=0\\varepsilon=0\.Mk,ε,1/2M\_\{k,\\varepsilon,1/2\}legitimately exceeds the vanilla bound \(4\.02\>3\.994\.02\>3\.99atk=50k=50\), so the tight bound falsely aborted\. Later aligned this with proposition 6\.5 from Polymath8b toRbound=kk−1​log⁡\(2​k−1\)\(ε\>0\)\.R\_\{\\mathrm\{bound\}\}=\\frac\{k\}\{k\-1\}\\log\(2k\-1\)\\hskip 18\.49988pt\(\\varepsilon\>0\)\.There is a more stringent bound in Polymath8b Remark 6\.6 but this suffices as hard\-coded discovery threshold / safeguard\.14Per\-pair diagnostic \+\-\-dump\-violation\+\-\-patience;kk\-aware panels; excursion guard;polish\(\)incumbent seeding\.Fixed 8\-panel mesh under\-resolvedxk−2x^\{k\-2\}atk≥250k\\geq 250\(e152e^\{152\}vs2​nout=1282n\_\{\\text\{out\}\}\{=\}128\);polish\(\)started at−∞\-\\inftyso step 1 installedR=70R=70as best\.15\-\-trunc\-floorprevention \+ rank\-stability detection \+ strike\-counter fix\.R=56R=56lived inAA\-directions atλ/λmax∼10−6\\lambda/\\lambda\_\{\\max\}\\sim 10^\{\-6\}, below assembly accuracy10−310^\{\-3\}; whitening amplified error \(again\!\)∼103\\sim 10^\{3\}\.strikesreset before the check⇒\\Rightarrownever accrued\.16Stability recalibration \(25%, floored\) \+ absoluteR≤kR\\leq kbackstop \+ excursion baseline anchor \+ grow\-step guard \+ Richardson already\-converged path\.5%5\\%flagged honest wobble;best=−∞\\text\{best\}=\-\\inftylet a runaway pass at iter 0; a0\.3750\.375baseline made3\.663\.66look9\.7×9\.7\\times; a post\-grow state crashed the run; a convergedk=500k\{=\}500run got a false FAIL\.Listing 8:Closed\-form preflights and post\-training representation refinement\.PROCEDUREPREFLIGHT\_AND\_REFINE\(theta,k,epsilon,N\)

REQUIRENUMERICAL\_QUOTIENT\(g=1\)AGREESWITHCLOSED\_FORM\_CONSTANT\(k,epsilon\)

IFepsilon=0

REQUIRENUMERICAL\_QUOTIENT\(g=x\)AGREESWITH3\*k/\(3\*k\+1\)

END

FORNqIN\{N,CEIL\(1\.25\*N\),CEIL\(1\.5\*N\)\}

R\[Nq\]<\-VALIDATED\_QUOTIENT\(theta,Nq\)

END

IFsuccessivedifferencescontractwithacommonsign

FITR\[Nq\]=R\_infinity\-c\*Nq^\(\-p\)

REQUIRE0\.8<=p<=4

REQUIREABS\(R\[1\.5\*N\]\-R\_infinity\)/ABS\(R\_infinity\)<=tau\_refine

REQUIRER\_infinity<=RIGOROUS\_CEILING\(k,epsilon\)

ELSE

MARKcandidateasnumericallyunresolved

END

RETURNR,R\_infinity,p

END

### 5\.13Aggregate modular evaluation

After projection, letqj,qℓq\_\{j\},q\_\{\\ell\}be rational channel polynomials and

hj​ℓ​\(u\)=qj​\(u\)​qℓ​\(u\)\.h\_\{j\\ell\}\(u\)=q\_\{j\}\(u\)q\_\{\\ell\}\(u\)\.Rather than reconstructing every exact entry of the Gram matrices, the aggregate certifier accumulates the weighted pair contributions modulo machine\-word primes:

SA​\(p\)=∑j≤ℓκj​ℓ​cj​cℓ​Aj​ℓint\(modp\),SB​\(p\)=∑j≤ℓκj​ℓ​cj​cℓ​Bj​ℓint\(modp\)\.S\_\{A\}\(p\)=\\sum\_\{j\\leq\\ell\}\\kappa\_\{j\\ell\}c\_\{j\}c\_\{\\ell\}A^\{\\rm int\}\_\{j\\ell\}\\pmod\{p\},\\qquad S\_\{B\}\(p\)=\\sum\_\{j\\leq\\ell\}\\kappa\_\{j\\ell\}c\_\{j\}c\_\{\\ell\}B^\{\\rm int\}\_\{j\\ell\}\\pmod\{p\}\.All polynomial powers, coefficient extractions, pair weights and channel sums are evaluated in𝔽p​\[x\]\\mathbb\{F\}\_\{p\}\[x\]\. The machine\-word primes have no denominator occurring in the modular formulas with vanishing modulopp, ensuring that every required inverse in𝔽p\\mathbb\{F\}\_\{p\}is defined\.

Polynomial arithmetic is performed with exact compiled integer and modular kernels\. Earlier exact\-certification implementations used Kronecker substitution: integer polynomial coefficients were packed into large integers, multiplied using GMP\-backed arbitrary\-precision arithmetic, and unpacked exactly; polynomial powers were then formed by binary exponentiation\. In the later FLINT\-backed implementations, polynomial multiplication, powering and modular reduction are delegated to FLINT\[[14](https://arxiv.org/html/2609.30296#bib.bib17)\], accessed throughpython\-flint\. FLINT selects its internal multiplication strategy according to operand size and representation\. These implementation choices affect runtime and memory use but not the mathematical certificate, which depends only on the exact resulting elements of𝔽p​\[x\]\\mathbb\{F\}\_\{p\}\[x\]and the subsequent CRT reconstruction\.

Only the two scalar aggregate integersSA,SBS\_\{A\},S\_\{B\}are reconstructed\. Their exact quadratic forms are

c⊤​A​c=SADA,c⊤​B​c=SBDB,c^\{\\top\}Ac=\\frac\{S\_\{A\}\}\{D\_\{A\}\},\\qquad c^\{\\top\}Bc=\\frac\{S\_\{B\}\}\{D\_\{B\}\},with global denominators fixed before the modular sweep\.

Let

NCRT=k​smax,rmax=\(k−1\)​smax,E=k−1\+rmax\+mmax,N\_\{\\mathrm\{CRT\}\}=ks\_\{\\max\},\\qquad r\_\{\\max\}=\(k\-1\)s\_\{\\max\},\\qquad E=k\-1\+r\_\{\\max\}\+m\_\{\\max\},and

Lcm=lcm⁡\{k−1,…,k−1\+rmax\+mmax\}\.L\_\{\\mathrm\{cm\}\}=\\operatorname\{lcm\}\\\{k\-1,\\ldots,k\-1\+r\_\{\\max\}\+m\_\{\\max\}\\\}\.The resulting global denominators are

DA=\(k\+NCRT\)\!​dhk​dc2,D\_\{A\}=\(k\+N\_\{\\mathrm\{CRT\}\}\)\!\\,d\_\{h\}^\{k\}d\_\{c\}^\{2\},and

DB=\(k−2\+rmax\)\!​Lcm​ρdE​Tden​dhk−1​dc2\.D\_\{B\}=\(k\-2\+r\_\{\\max\}\)\!\\,L\_\{\\mathrm\{cm\}\}\\,\\rho\_\{d\}^\{E\}T\_\{\\rm den\}\\,d\_\{h\}^\{k\-1\}d\_\{c\}^\{2\}\.All factors are therefore known before reconstruction\.

The numerator\-side coefficient sum can be expressed as a polynomial correlation\. IfTj​ℓ​\(x\)T\_\{j\\ell\}\(x\)denotes the pair\-dependent integrated\-channel polynomial and

Vp​\(x\)=∑v=0rmax\+mmax\(k−1\+v\)−1​xv∈𝔽p​\[x\],V\_\{p\}\(x\)=\\sum\_\{v=0\}^\{r\_\{\\max\}\+m\_\{\\max\}\}\(k\-1\+v\)^\{\-1\}x^\{v\}\\in\\mathbb\{F\}\_\{p\}\[x\],then the required shifted inner products are obtained from multiplication with a reversed form ofTj​ℓT\_\{j\\ell\}, implementing the transposition principle\.

Polynomial powers are formed deterministically from exact arithmetic\. Classical sparse\-polynomial powering algorithms provide related powering recurrences\[[24](https://arxiv.org/html/2609.30296#bib.bib20)\], but in the present application the factorially weighted pair polynomials are typically dense\. The principal computational gain therefore comes from aggregate modular evaluation, exact coefficient extraction and CRT reconstruction rather than from sparse polynomial storage itself\.

### 5\.14Known denominators and deterministic CRT reconstruction

The modular stage reconstructs signed integer numerators rather than generic rational numbers\. Rigorous bounds

\|SA\|≤ℬA,\|SB\|≤ℬB\|S\_\{A\}\|\\leq\\mathcal\{B\}\_\{A\},\\qquad\|S\_\{B\}\|\\leq\\mathcal\{B\}\_\{B\}are obtained using

‖hr‖1≤‖h‖1r\\\|h^\{r\}\\\|\_\{1\}\\leq\\\|h\\\|\_\{1\}^\{r\}together with explicit bounds for factorial ratios, powers of the rational support ratio, integrated channel coefficients and rational mixing weights\. If

MP=∏p∈PpM\_\{P\}=\\prod\_\{p\\in P\}pand

MP\>2​max⁡\(ℬA,ℬB\),M\_\{P\}\>2\\max\(\\mathcal\{B\}\_\{A\},\\mathcal\{B\}\_\{B\}\),the centered CRT representative in\(−MP/2,MP/2\]\(\-M\_\{P\}/2,M\_\{P\}/2\]is unique\. The implementation uses an additional safety margin, selecting primes until the product exceeds four times the larger bound\. Held\-out primes provide an implementation\-level check but are not needed for mathematical uniqueness\.

The certified quotient is

Rcert=k⁡\(1\+ε\)​SB/DBSA/DA\.R\_\{\\rm cert\}=k\(1\+\\varepsilon\)\\frac\{S\_\{B\}/D\_\{B\}\}\{S\_\{A\}/D\_\{A\}\}\.Once the rational channels and mixing vector are fixed, this value is independent of the neural discovery stage\.

### 5\.15Development and validation history of the CRT certifier

The exact certifier was developed through a sequence of increasingly large stress tests in which numerical agreement between the floating\-point proposal stage and the exact reconstruction stage was treated as a falsification criterion\. Supplementary Table[3](https://arxiv.org/html/2609.30296#S5.T3)records the principal failure modes and the corresponding corrections\. These historical iterations are not independent proof components; they document how the final implementation acquired the numerical invariants and trust boundaries used throughout the reported certifications\.

Supplementary Table 3:Development history of the exact multimodular \(CRT\) certifier forMk,ε,1/2M\_\{k,\\varepsilon,1/2\}\. Each version records the change and the failure mode that motivated it\. With one exception \(the FFT accuracy floor, v4\), every defect was a silent disagreement between the floating\-point pass and the exact pass at a specific weight vectorcc, invisible at smallkkand exposed only by scale, conditioning, or polynomial degree\.VerChangeFailure mode / motivationv1Initial multimodular backend\.h^k−1\\hat\{h\}^\{\\,k\-1\}is never formed overℤ\\mathbb\{Z\}; per6262\-bit prime, onenmodpowering plus two weight\-table multiplies execute both linear functionals as single coefficient extractions:\[xnmax\]​\(b​hk⋅Wrev\)\[x^\{n\_\{\\max\}\}\]\(bh^\{k\}\\cdot W\_\{\\mathrm\{rev\}\}\)forAA, and forBBthe pair\-dependenttt\-weights combined as∼mmax\{\\sim\}m\_\{\\max\}scalar–poly ops against per\-prime tablesG\(m\)G^\{\(m\)\}folding inΛ=lcm⁡\(k−1,…,k−1\+rmax\+mmax\)\\Lambda=\\operatorname\{lcm\}\(k\{\-\}1,\\dots,k\{\-\}1\{\+\}r\_\{\\max\}\{\+\}m\_\{\\max\}\), theρ\\rho\-power clearing, and batched harmonic inverses\. Reconstruction is aggregate\-only—two integersSA,SBS\_\{A\},S\_\{B\}with known\-denominator scalings—via deterministic centered CRT against exact a priori bounds, plus held\-out verification primes\. Float pass decoupled from the exact engine \(A=∫01h∗kA=\\int\_\{0\}^\{1\}h^\{\*k\},B=∫0ρh∗\(k−1\)​Qj​Ql​\(1−s\)B=\\int\_\{0\}^\{\\rho\}h^\{\*\(k\-1\)\}Q\_\{j\}Q\_\{l\}\(1\{\-\}s\)by renormalized binary powering\)\. Greedycc\-weighted pruning on the float pencil \(\-\-prune\-tol\); only surviving pairs reach the CRT engine\.Establishes the architecture: avoid the one unavoidable giant integer, reconstruct only two aggregate scalars, keep the float discovery pass cheap and separate from exact arithmetic\.v2\_ritzno longer calls the Cholesky\-based generalized solver: it equilibrates by the diagonal, eigendecomposes the metricAA, truncates to eigenvalues above10−1210^\{\-12\}of the top one, and solves the standard symmetric problem in whitened coordinates, embedding zeros for dead channels\. Self\-test adds exact agreement withscipy\.eigh\(B,A\)on a PD case and a finiteλ\\lambdaon a singular metric\.scipy\.eigh\(B,A\)requiresAApositive definite; atk=500k=500the float Gram is numerically semidefinite and Cholesky fails \(LinAlgError: leading minor not positive definite\)\. Restricting torange⁡\(A\)\\operatorname\{range\}\(A\)is the correct regularization: null\-space components contribute≈0\{\\approx\}0toc𝖳​A​cc^\{\\mathsf\{T\}\}Ac\.v3B\-side of the float DP rescaled byu=s/ρu=s/\\rho: run the\(k−1\)\(k\{\-\}1\)\-fold convolution power onhρ​\(t\)=h​\(ρ​t\)h\_\{\\rho\}\(t\)=h\(\\rho t\), soB=ρk−1​∫01hρ∗\(k−1\)​\(u\)​T​\(ρ​u\)​𝑑uB=\\rho^\{\\,k\-1\}\\\!\\int\_\{0\}^\{1\}h\_\{\\rho\}^\{\*\(k\-1\)\}\(u\)\\,T\(\\rho u\)\\,du, withρk−1\\rho^\{k\-1\}carried in the log\-scale \(ε=0\\varepsilon=0degenerates to the old path\)\. Driver interlock: float\-predictedRRchecked against discoveryRNNR\_\{\\mathrm\{NN\}\}before any exact work\.Atk=500k=500the certified value collapsed to∼10−12\{\\sim\}10^\{\-12\}\(valid but worthless\)\.h∗\(k−1\)h^\{\*\(k\-1\)\}concentrates likesk−2s^\{k\-2\}nears=1s=1, so the whole B\-regions≤ρ=1213s\\leq\\rho=\\tfrac\{12\}\{13\}satρk∼e−40\\rho^\{k\}\\sim e^\{\-40\}below the renormalized peak—beneath float64 accuracy\.BBwas noise; the pruner dropped4848channels “for free” and the exact engine certified a garbage frozencc\.v4Truncated convolution\_tconvforced to*direct*np\.convolve, removing the FFT path\. Large\-kkfloat\-DP regression test against exactg≡1g\\equiv 1closed forms atk=200,500k=200,500\.FFT convolution carries an*absolute*error∼10−16\{\\sim\}10^\{\-16\}relative to the vector peak, whileh∗jh^\{\*j\}spanssj−1s^\{j\-1\}internally; entries sit up to2−j2^\{\-j\}below peak and are buried by noise byj∼55j\\sim 55, then amplified through each squaring \(closed\-form check failed bye\+1000e^\{\+1000\}\)\. Direct summation keeps per\-entry*relative*accuracy; the self\-convolution integrand peaks at the interiort=s/2t=s/2, representable toj∼511j\\sim 511\(float64 ceiling neark∼1000k\\sim 1000\)\.v5Product\-exact prime budgeting \(gen\_primes\_for\_bound\): accumulate the actual prime product until it exceeds4⋅bound4\\cdot\\text\{bound\}rather than assuming6161bits/prime\. Miller–Rabin citation corrected \(least strong pseudoprime to bases2​…​372\\ldots 37is3\.19×10233\.19\\times 10^\{23\}; Sorenson–Webster 2017, OEIS A014233\)\. Self\-test cross\-checks generated primes against FLINT’s provenfmpz\.is\_prime; new strong\-cancellation test \(c=\(1,−1\)c=\(1,\-1\), near\-identical channels\)\.Collaborator review round: tighten certification claims \(exact prime count, correct primality bound, FLINT belt\-and\-braces\), cover the small\-signed\-aggregate CRT branch, and*measure*which proposed speedups were real \(mul\_lowbenched slower than the full FFT multiply; native modular dot kernel deferred\)\.v6Added\-\-float\-onlysurvey mode: projection\+\+float DP\+\+pruning per degree, stopping before any exact/CRT work\.Cheap degree scouting at largekk: projectionL∞L\_\{\\infty\}and float\-predictedRRanswer “is thiskkreachable at this basis degree” in minutes per degree, without spending the prime budget\.v7Stable Legendre\-basis float evaluation\.project\_channels\_v3returns both monomial integers \(exact engine\) and dyadic Legendre numerators \(float engine\); the float engine evaluates channels and the B\-weight antiderivative via the shifted\-Legendre recurrence, with a non\-finite guard that raises with a diagnosis\.Atd=160d=160the float pass produced NaN: channels were evaluated bynp\.polyvalon monomial coefficients of size∼8d∼10150\{\\sim\}8^\{d\}\\sim 10^\{150\}, so Horner formed anO⁡\(1\)O\(1\)value as an alternating sum of1015010^\{150\}\-scale terms \(∼150\{\\sim\}150orders of cancellation\), poisoning the convolutions to±∞\\pm\\inftyand henceoffA=max⁡\(−∞\)=−∞\\text\{off\}\_\{A\}=\\max\(\-\\infty\)=\-\\infty,\(−∞\)−\(−∞\)=NaN\(\-\\infty\)\-\(\-\\infty\)=\\text\{NaN\}\. The Legendre recurrence \(\|P~n\|≤1\|\\tilde\{P\}\_\{n\}\|\\leq 1on\[0,1\]\[0,1\]\) is well\-conditioned at any degree\.v8Auto\-escalating rank truncation in the whitened Ritz, anchored toRNNR\_\{\\mathrm\{NN\}\}:\_ritzgains arank\_tolthreaded throughgreedy\_prune; the driver escalates10−10→10−8→10−6→10−410^\{\-10\}\\\!\\to\\\!10^\{\-8\}\\\!\\to\\\!10^\{\-6\}\\\!\\to\\\!10^\{\-4\}until the predictedRRis physical, or uses a fixed\-\-rank\-tol\.Atd=40d=40,k=500k=500the floatRRcame out7\.4×10217\.4\\times 10^\{21\}\. Gram entries were cancellation\-limited to∼10−8\{\\sim\}10^\{\-8\}but the rank cut was10−12​λmax10^\{\-12\}\\lambda\_\{\\max\}; directions above the cut yet below the accuracy floor are noise, and whitening by1/noise1/\\sqrt\{\\text\{noise\}\}makesλ=Bnoise/Anoise\\lambda=B\_\{\\text\{noise\}\}/A\_\{\\text\{noise\}\}explode\. The true pencil is bounded byMk,εM\_\{k,\\varepsilon\}\(single digits\)\.v9Worker refactor: the B\-functional inner sumSR=∑mt​wm/\(k−1\+R\+m\)S\_\{R\}=\\sum\_\{m\}tw\_\{m\}/\(k\{\-\}1\{\+\}R\{\+\}m\)computed for allRRby*one*correlation multiplyI​T=INV⋅T​WrevIT=\\mathrm\{INV\}\\cdot TW\_\{\\mathrm\{rev\}\}against a single global per\-prime inverse polynomial, then combined in a short loop against the base ladder\. The\(mmax\+1\)\(m\_\{\\max\}\{\+\}1\)\-polynomialGrevG\_\{\\mathrm\{rev\}\}tables are eliminated; per\-prime tables become three ladders regardless of degree\. Added\-\-resume: append\-only\(p,SA,SB\)\(p,S\_\{A\},S\_\{B\}\)checkpoint\.Atd=160d=160,mmax=2​\(d\+1\)=322m\_\{\\max\}=2\(d\{\+\}1\)=322, so the v1GrevG\_\{\\mathrm\{rev\}\}tables were323×160​k≈413323\\times 160\\text\{k\}\\approx 413MB per worker and∼5×107\{\\sim\}5\\times 10^\{7\}Python ops per prime for table construction alone—dominating the sweep across∼28\{\\sim\}28k primes\. The correlation form is degree\-oblivious; resume makes the multi\-hour sweep preemption\-tolerant \(spot instances\)\.v10Exact mantissa\-preservingccrationalization plus a frozen\-ccgate\.cq=\[Fraction⁡\(float⁡\(v\)\)\]c\_\{q\}=\[\\,\\mathrm\{Fraction\}\(\\mathrm\{float\}\(v\)\)\\,\]\(full5353\-bit mantissa per component\) replaces uniform\-absolutedyadic⁡\(v,bits\)\\mathrm\{dyadic\}\(v,\\text\{bits\}\)rounding; the gate re\-evaluates the float quotient at the rationalizedccand aborts if it moved by\>1%\>1\\%\. Resume caches gain a SHA\-256 payload fingerprint and refuse stale residues\.After an1111\-hourk=500k=500sweep the certified value was8\.9×10−58\.9\\times 10^\{\-5\}\. Uniform absolute2−482^\{\-48\}rounding annihilated small\-but\-essential components of the balanced Ritzcc, whose entries span dozens of orders \(cj∼Aj​j−1/2c\_\{j\}\\sim A\_\{jj\}^\{\-1/2\},Aj​j∼\(sup\)2​kA\_\{jj\}\\sim\(\\sup\)^\{2k\}, so channels differing by1\.11\.1in sup\-norm have diagonals differing by1\.11000∼10411\.1^\{1000\}\\sim 10^\{41\}\)\. The float pass evaluated the truecc\(≈4\.19\{\\approx\}4\.19\); the exact engine certified the mutilated one\.v11Fixed the Gauss–Jacobi overflow in the shared quadrature builder: replacescipy\.roots\_jacobiwith an overflow\-free Golub–Welsch construction \(Jacobi tridiagonal eigenproblem; weights==squared first\-eigenvector components normalized to sum11, so the divergingμ0\\mu\_\{0\}cancels\)\.*\(Shared\-infrastructure fix; gates whether large\-kkdiscovery npz files can be produced at all\.\)*roots\_jacobicomputesμ0=2a\+b\+1​B​\(a\+1,b\+1\)\\mu\_\{0\}=2^\{a\+b\+1\}B\(a\{\+\}1,b\{\+\}1\)in linear space; withp−1∼2499p\-1\\sim 2499atk=2500k=2500the Beta/Gamma factors overflow float64 \(Γ⁡\(171\)\\Gamma\(171\)already overflows\), returning NaN weights and crashing theg≡1g\\equiv 1preflight\. Golub–Welsch needs only recurrence coefficients \(noΓ\\Gamma\); validated to∼10−14\{\\sim\}10^\{\-14\}vsroots\_jacobiforp≤60p\\leq 60and finite/correct top=6000p=6000\.v12Diagonal preconditioning of the float Gram:matrices\(\)subtracts\(log⁡Aj​j\+log⁡Al​l\)/2\(\\log A\_\{jj\}\+\\log A\_\{ll\}\)/2from each log\-entry \(soA^\\hat\{A\}’s diagonal is exactly11, off\-diagonalsO⁡\(1\)O\(1\)by Cauchy–Schwarz\), applies the same congruence toBB, and returnslog⁡Adiag\\log A\_\{\\mathrm\{diag\}\}\. Regression: a195195\-nat diagonal spread must keep all channels alive with float\-vs\-CRT agreement10−510^\{\-5\}\.A single globalexp⁡\(log⁡A−off\)\\exp\(\\log A\-\\text\{off\}\)offset underflowed entries<off−745<\\text\{off\}\-745to exactly00; atk=1000k=1000the diagonal spans∼6900\{\\sim\}6900nats, so the Ritz solve collapsed to rank22\(observed:3030of3232channels dropped “for free” at0\.00\.0loss\)\. Per\-channel log\-diagonal preconditioning keeps every channel alive regardless ofkk\.v13Frame conversion of the imported discoveryccon load\. When the npz carrieslog⁡Adiag\\log A\_\{\\mathrm\{diag\}\}, the loader maps discovery’s exported \(preconditioned\-frame\)ccto the true channel frame viactrue=c⋅exp⁡\(−12​log⁡Adiag\)⋅channel\_normsc\_\{\\mathrm\{true\}\}=c\\cdot\\exp\(\-\\tfrac\{1\}\{2\}\\log A\_\{\\mathrm\{diag\}\}\)\\cdot\\text\{channel\\\_norms\}\(both factors were then thought necessary: thekk\-th\-power diagonal scale is independent of theL2L^\{2\}scale\)\. Later corrected: discovery assembles its Gram pencil on theL2L^\{2\}\-normalized channelsgj/‖gj‖2g\_\{j\}/\\\|g\_\{j\}\\\|\_\{2\}, and the certifier rebuilds exactly these channels by projectingg\_fine/channel\_norms\\texttt\{g\\\_fine\}/\\text\{channel\\\_norms\}\. The coefficient on the normalized channels is thereforectrue=exp⁡\(−12​log⁡Adiag\)​c^c\_\{\\mathrm\{true\}\}=\\exp\(\-\\tfrac\{1\}\{2\}\\log A\_\{\\mathrm\{diag\}\}\)\\,\\hat\{c\}, with no channel\-norm factor\. A conversion to the raw channels would require‖gj‖2−k\\\|g\_\{j\}\\\|\_\{2\}^\{\-k\}, not‖gj‖2\\\|g\_\{j\}\\\|\_\{2\}\. The conversion is now shared by all exact backends \(certification/frames\.py\)\. Validity was never affected, since any explicit rationalccyields a valid lower bound, and the re\-Ritz path \(v14\) never used the channel\-norm factor\.Atk=53k=53discovery reported3\.9943\.994but the certifier gave3\.5943\.594\. The pencil’s frame\-invariant top eigenvalue was3\.99483\.9948; discovery exportsccin the frame of its*preconditioned*Gram \(c^=D1/2​ctrue\\hat\{c\}=D^\{1/2\}c\_\{\\mathrm\{true\}\}\), while the certifier evaluatesccon the*true*channels\. The deflation grows with thelog⁡Adiag\\log A\_\{\\mathrm\{diag\}\}spread \(308308nats atk=53k=53,∼6900\{\\sim\}6900atk=1000k=1000\)\.v14Frame conversion of the re\-Ritzcc\(second instance of the same bug\)\. The driver captureslog⁡Adiag\\log A\_\{\\mathrm\{diag\}\}from the float engine and maps\_ritz’s output back to the true frame \(ctrue=exp⁡\(−12​log⁡Adiag\)​c^c\_\{\\mathrm\{true\}\}=\\exp\(\-\\tfrac\{1\}\{2\}\\log A\_\{\\mathrm\{diag\}\}\)\\,\\hat\{c\}\) on the active channels before rationalizing; the frozen\-ccgate is made frame\-consistent\.After v13,\-\-use\-nn\-cstill gave3\.5943\.594\.\_ritzwhitens the preconditioned Gram internally, so its eigenvector lives in the preconditioned frame—and the driver certified it as if true\-frame\. Certifyingc^\\hat\{c\}as true\-frame gives3\.59383\.5938\(the observed value\); the mappedccgives3\.99393\.9939, and the end\-to\-end run then certified3\.99443\.9944\(gap\+4×10−4\+4\\times 10^\{\-4\}to discovery\)\. Affected*every*re\-Ritz\-path certification against a preconditioned\-export npz\.
### 5\.16Representation\-preserving transfer

The discovery solver uses diagonal preconditioning\. If

Adisc=D−1/2AtrueD−1/2A\_\{\\rm disc\}=D^\{\-1/2\}A\_\{\\rm true\}D^\{\-1/2\}andc^\\widehat\{c\}is the mixing vector in discovery coordinates, then

ctrue=D−1/2c^\.c\_\{\\rm true\}=D^\{\-1/2\}\\widehat\{c\}\.Any additional channel normalization is likewise restored before certification\.

During projection and floating\-point evaluation the channels are represented in a shifted\-Legendre basis,

qj​\(u\)=∑r=0daj​r​P~r​\(u\),q\_\{j\}\(u\)=\\sum\_\{r=0\}^\{d\}a\_\{jr\}\\widetilde\{P\}\_\{r\}\(u\),and evaluated through the three\-term recurrence\. For exact certification the same rational expansion is transformed symbolically into integer monomial coefficients with a common denominator\. If

h⁡\(u\)=∑r=0shr​ur,h\(u\)=\\sum\_\{r=0\}^\{s\}h\_\{r\}u^\{r\},define its factorially weighted coefficient polynomial by

h^​\(x\)=∑r=0sr\!​hr​xr\.\\widehat\{h\}\(x\)=\\sum\_\{r=0\}^\{s\}r\!h\_\{r\}x^\{r\}\.
For enlarged support,

u=t1\+ε,rε=1−ε1\+ε,u=\\frac\{t\}\{1\+\\varepsilon\},\\qquad r\_\{\\varepsilon\}=\\frac\{1\-\\varepsilon\}\{1\+\\varepsilon\},and the associated Jacobian and powers of1\+ε1\+\\varepsilonandrεr\_\{\\varepsilon\}are retained explicitly\. The mixing vector is embedded into the rationals by preserving each binary floating\-point component as an exact dyadic rational, avoiding loss of small but cancellation\-relevant entries\. Before modular evaluation, the frozen rational vector is re\-evaluated in the floating\-point quadratic forms and required to reproduce the Ritz candidate within tolerance\. This is a transfer check, not part of the proof\.

### 5\.17Loss\-controlled channel reduction

Exact certification cost is quadratic in the number of surviving channels\. Starting from the full floating\-point Gram pencil, channels are removed by backward elimination\. For each candidate removal, the generalized Rayleigh–Ritz problem is re\-solved\. Ifλ0\\lambda\_\{0\}is the original value andλS\\lambda\_\{S\}the value on the active subset,

λ0−λSλ0≤τprune\\frac\{\\lambda\_\{0\}\-\\lambda\_\{S\}\}\{\\lambda\_\{0\}\}\\leq\\tau\_\{\\rm prune\}is required\. At each step the channel whose removal leaves the largest quotient is discarded\. This differs from coefficient thresholding because importance is measured directly in the reoptimized variational objective\. The reduced rational trial function is then certified exactly; pruning changes only cost and attainable bound, not mathematical validity\.

![Refer to caption](https://arxiv.org/html/2609.30296v1/Section1/images/k3000-g.png)Supplementary Figure 1:Diagnostic large\-kkfailure of the early discovery pipeline\.Fork=3600k=3600, the discovery stage reported the spurious valueR=8\.1375743146R=8\.1375743146\. Nearly all Ritz weight collapsed onto a single channel, shown here together with its degree\-120 polynomial projection, residual structure, logarithmic\-scale amplitude and dyadic shifted\-Legendre coefficient magnitudes\. Although the discovery objective was stable under grid refinement and remained below the Cauchy–Schwarz ceiling, the certification failed independent reconstruction and exposed the representation inconsistency; rejecting this candidate\.![Refer to caption](https://arxiv.org/html/2609.30296v1/Section1/images/k3000-onechannelsweep.png)Supplementary Figure 2:Polynomial projection convergence for the single Ritz weight holding channel ofk=3600k=3600\.RelativeL2L^\{2\}andL∞L^\{\\infty\}projection errors decrease with polynomial degree, illustrating that the transfer error can be reduced systematically once a numerically resolved channel has been obtained\.
### 5\.18Diagnostic large\-kkfailure and rational distillation

The large\-kkexperiments exposed a failure mode that was not detected by the original discovery diagnostics\. In an early run atk=3600k=3600, the multi\-channel optimizer collapsed onto a single sharply localized channel and reported a Rayleigh value above88\(See Supplementary Figures[1](https://arxiv.org/html/2609.30296#S5.F1)and[2](https://arxiv.org/html/2609.30296#S5.F2)\)\. The value was stable under refinement of the representation grid and remained below the aforementioned ceiling \(Rk≤kk−1​log⁡kR\_\{k\}\\leq\\frac\{k\}\{k\-1\}\\log k\) so the candidate initially appeared numerically credible\. Also, across the validated separable runs, increasingkkwas accompanied by a progressive reduction in the number of channels needed to represent the best surviving candidates\. This trend was empirical\.

Upon certification, it appeared difficult to adequately represent the one\-channel discovery solution\. For a single exported channelqq, the scale\-invariant quantity

R1​D=k​\(∫q\)2∫q2R\_\{\\mathrm\{1D\}\}=k\\frac\{\\left\(\\int q\\right\)^\{2\}\}\{\\int q^\{2\}\}disagreed by orders of magnitude with the stored diagonal ratiok​Bj​j/Aj​jkB\_\{jj\}/A\_\{jj\}\. Since multiplication ofqqby an arbitrary scalar cannot change either ratio, this discrepancy cannot be explained by a normalization or frame convention\. For the sharpest channel, the support of the exported function also gave the Cauchy–Schwarz boundR1​D≤k​σR\_\{\\mathrm\{1D\}\}\\leq k\\,\\sigma, whereσ\\sigmais the effective support measure; the stored discovery value violated the corresponding scale by more than two orders of magnitude\.

The discrepancy was caused by a scale\-law inconsistency in the discovery evaluation of \(increasingly\) narrow channels\. The convolution recursion operated in a normalized variable, whereas the exported channel lived in the physical coordinate\. In the intermediate\-rate regime the error followed approximately the missing\-Jacobian law

Rstored≈ρk​Rdirect,R\_\{\\mathrm\{stored\}\}\\approx\\frac\{\\rho\}\{k\}R\_\{\\mathrm\{direct\}\},withρ\\rhothe channel rate\. At the highest rates a second resolution effect appeared, and the sharpest exported channel was itself truncated\. The optimizer therefore concentrated its Ritz weight on the channel for which the numerical artifact was largest\. The apparent rank\-one structure was consequently not, by itself, evidence of a genuine extremizer\.

The comparison is exact when the channel is genuinely compactly supported: ifsupp⁡\(g\)⊆\[0,σ\]\\operatorname\{supp\}\(g\)\\subseteq\[0,\\sigma\]withk​σ≤1k\\sigma\\leq 1, then\[0,σ\]k⊆ℛk\[0,\\sigma\]^\{k\}\\subseteq\\mathcal\{R\}\_\{k\}, soIk​\(F\)=\(∫g2\)kI\_\{k\}\(F\)=\(\\int g^\{2\}\)^\{k\}andJk\(1\)​\(F\)=\(∫g\)2​\(∫g2\)k−1J^\{\(1\)\}\_\{k\}\(F\)=\(\\int g\)^\{2\}\(\\int g^\{2\}\)^\{k\-1\}, givingR=k​\(∫g\)2/∫g2R=k\(\\int g\)^\{2\}/\\int g^\{2\}identically\. For a channel with only an*effective*support the identity acquires an uncontrolled tail term, so the scalar test is applied as an empirical safeguard against gross inflation rather than as an exact identity\.

This observation is important because the artifact was close to the true large\-kkscale\. The sharp\-channel artifact approached a value proportional to the first inner quadrature weight, and for the quadrature order used in the run it crossed88in the samekk\-range in which the true variational constant approaches88\. Grid stability and upper bounds were therefore insufficient to identify artifact from valid\.

The final discovery solver adds an independent scalar\-channel gate\. For each sufficiently localized channel it comparesk​Bj​j/Aj​jkB\_\{jj\}/A\_\{jj\}with a direct one\-dimensional quadrature of

k​\(∫gj\)2∫gj2,k\\frac\{\(\\int g\_\{j\}\)^\{2\}\}\{\\int g\_\{j\}^\{2\}\},using a quadrature path independent of the convolution assembly\. Wide channels are additionally checked against support\-aware Cauchy–Schwarz bounds\. Narrow channels get the hardμ×inflation\\mu\\times\\text\{inflation\}rejection\. Candidates failing these tests are rejected before export\.

Conditional on candidates that survive these discovery checks, rational distillation provides a second and logically independent filter\. Richer multi\-cluster rational fits can reduce pointwise approximation error while failing probability\-mass, conditioning or Fourier\-stability checks\. Because small one\-dimensional transform errors are raised to the\(k−1\)\(k\-1\)st power in the large\-kkevaluation, pointwise agreement alone does not preserve the Maynard functional\.

In the stable distilled representation, the surviving family reduces to the single rational pole

g⁡\(t\)=1c\+\(k−1\)​t\.g\(t\)=\\frac\{1\}\{c\+\(k\-1\)t\}\.Its Rayleigh quotient is then optimized and evaluated independently of the neural model\. Thus the rank\-one collapse observed in the failed neural run and the rank\-one structure of the certified rational family have opposite logical status: the former was an optimization response to a numerical artifact, whereas the latter survives reconstruction and independent certification\. Neither the failed discovery value nor the neural rank collapse is used as evidence for the final bound\.

Listing 9:Checks used to diagnose unresolved or spuriously inflated neural channels before transitioning to the rational large\-kkroute\. The all\-width moment comparison is a conservative empirical discovery safeguard and is not part of the exact certificate\.FUNCTIONLOG\_INNER\_MEAN\(log\_residual,rate,Y,z\_cap\)

IFrate\*Y<z\_cap

u<\-Y\*x\_GL

RETURNLOGSUMEXP\_r\(LOG\(w\_GL\[r\]\)\+log\_residual\(u\[r\]\)\-rate\*u\[r\]\)

ELSE

z<\-z\_cap\*x\_GL

RETURNLOG\(z\_cap/\(rate\*Y\)\)

\+LOGSUMEXP\_r\(LOG\(w\_GL\[r\]\)\+log\_residual\(z\[r\]/rate\)\-z\[r\]\)

END

END

PROCEDURELARGE\_K\_CHANNEL\_CHECKS\(g,A,B\)

rho<\-SOFTPLUS\(eta\)

\(I1\[j\],I2\[j\],sigma\[j\]\)<\-FINE\_GRADED\_ONE\_DIMENSIONAL\_QUADRATURE\(g\_j\)

R\_direct\[j\]<\-k\*I1\[j\]^2/I2\[j\]

R\_conv\[j\]<\-k\*B\[j,j\]/A\[j,j\]

REQUIRER\_conv\[j\]<=\(1\+tau\_moment\)\*R\_direct\[j\]FOREVERYj

IFk\*sigma\[j\]<=1ANDR\_conv\[j\]\>mu\*R\_direct\[j\]

REJECTunresolvedsingle\-channelinflation

END

IFk\*sigma\[j\]\>1ANDR\_conv\[j\]\>\(1\+tau\_support\)\*k\*sigma\[j\]

FLAGthechannelassupport\-inconsistent

END

curvature<\-MEAN\_SQUARED\_CURVATURE\(LOG\_RAW,coordinate=rho\*x\)

RETURNcurvature

END

### 5\.19Neural\-to\-rational structural distillation

At largekk, the neural candidate is compressed into the clustered rational family

grat​\(t\)=∑a=1r∑j=1pmaxua,j​\(ca\+n​t\)−j,n=k−1\.g\_\{\\rm rat\}\(t\)=\\sum\_\{a=1\}^\{r\}\\sum\_\{j=1\}^\{p\_\{\\max\}\}u\_\{a,j\}\(c\_\{a\}\+nt\)^\{\-j\},\\qquad n=k\-1\.\(39\)For fixed cluster locationsc=\(c1,…,cr\)c=\(c\_\{1\},\\ldots,c\_\{r\}\), the coefficients enter linearly and are determined by variable projection,

u⋆​\(c\)=arg⁡min⁡∑iu⁡ωi​\[gNN​\(ti\)−∑a,jua,j​\(ca\+n​ti\)−j\]2\.u^\{\\star\}\(c\)=\\arg\\min\_\{u\}\\sum\_\{i\}\\omega\_\{i\}\\left\[g\_\{\\rm NN\}\(t\_\{i\}\)\-\\sum\_\{a,j\}u\_\{a,j\}\(c\_\{a\}\+nt\_\{i\}\)^\{\-j\}\\right\]^\{2\}\.\(40\)The outer optimizer therefore varies only the nonlinear pole locations\. Negligible coefficients are removed and the highest retained power at each pole determines an inferred multiplicity\. Higher powers are confluent directions, since

∂q∂cq​\(c\+n​t\)−1=\(−1\)q​q\!​\(c\+n​t\)−\(q\+1\)\.\\frac\{\\partial^\{q\}\}\{\\partial c^\{q\}\}\(c\+nt\)^\{\-1\}=\(\-1\)^\{q\}q\!\(c\+nt\)^\{\-\(q\+1\)\}\.
Each distilled candidate is subsequently reconstructed in an independent deterministic evaluator\. Fourier\-domain cross terms are reduced to

Fp​\(ξ\)=∫01ei​ξ​t​\(c\+n​t\)−p​𝑑t,F\_\{p\}\(\\xi\)=\\int\_\{0\}^\{1\}e^\{i\\xi t\}\(c\+nt\)^\{\-p\}\\,dt,\(41\)with, forp≥2p\\geq 2,

Fp​\(ξ\)=−ei​ξ​\(c\+n\)1−p−c1−p−i​ξ​Fp−1​\(ξ\)n⁡\(p−1\)\.F\_\{p\}\(\\xi\)=\-\\frac\{e^\{i\\xi\}\(c\+n\)^\{1\-p\}\-c^\{1\-p\}\-i\\xi F\_\{p\-1\}\(\\xi\)\}\{n\(p\-1\)\}\.\(42\)Because upward recursion can amplify round\-off at large\|ξ\|\|\\xi\|, the Fourier computation is restricted to a stability window and requires the characteristic function to decay below a prescribed threshold before truncation\. Products involving distinct clusters are reduced by partial fractions; cross\-cluster cancellation is monitored and ill\-conditioned configurations are rejected\.

Further checks include conservation of probability mass, agreement of the first moment with its analytic value, control of the negative Fourier\-inversion noise floor, rank diagnostics for the Gram matrix, and the analytic ceiling

Rk≤kk−1​log⁡k\.R\_\{k\}\\leq\\frac\{k\}\{k\-1\}\\log k\.A candidate that violates any check is rejected rather than assigned a numerical objective\. A surviving structure is locally refined in the deterministic evaluator, so the final rational value is a newly optimized Rayleigh quotient rather than the neural objective or the pointwise fit\.

Listing 10:Weighted variable\-projection distillation of a neural channel, followed by independent Rayleigh optimization and evaluation in the resulting clustered rational family\.PROCEDUREDISTILL\_NEURAL\_CHANNEL\(t,g\_NN,omega,initialc,multiplicitiesmu,k\)

n<\-k\-1

PARAMETERIZEcaspositiveandstrictlyordered

FUNCTIONPROFILED\_FIT\_ERROR\(c\)

X\[:,a,j\]<\-\(c\[a\]\+n\*t\)^\(\-j\),j=1,\.\.\.,mu\[a\]

Xw<\-DIAG\(SQRT\(omega\)\)\*X

yw<\-SQRT\(omega\)\*g\_NN

scaleeachcolumnofXwtounitnorm

u<\-LEAST\_SQUARES\(Xw,yw\)

undothecolumnscalinginu

RETURNSUM\_iomega\[i\]\*\(g\_NN\[i\]\-\(X\*u\)\[i\]\)^2

END

c<\-LBFGS\_MINIMIZE\(PROFILED\_FIT\_ERROR,c\)

u<\-PROFILE\_LINEAR\_COEFFICIENTS\(c\)

PRUNEtermswithnegligiblescaledcontribution

INFERclustermultiplicitiesfromtheretainedpowers

RETURNc,u,inferredmultiplicities,relativefiterror

END

PROCEDUREREFINE\_RATIONAL\_FAMILY\(k,cluster\_locationsc,multiplicitiesmu\)

n<\-k\-1;L<\-1\+epsilon

basis<\-\{\(c\[a\]\+n\*t\)^\(\-j\):j=1,\.\.\.,mu\[a\]\}

PARAMETERIZEcaspositiveandstrictlyordered

FOREACHchannelpair

IFbothchannelsbelongtoonecluster

F\[1\]\(theta\)<\-SINE\_COSINE\_INTEGRAL\_FORMULA\(theta,c,n,L\)

FORp=2,\.\.\.,p\_max

F\[p\]<\-\-\(EXP\(i\*theta\*L\)\*\(c\+n\*L\)^\(1\-p\)\-c^\(1\-p\)

\-i\*theta\*F\[p\-1\]\)/\(n\*\(p\-1\)\)

END

ELSE

REDUCEtheproductbypartialfractions

REJECTexcessivecross\-clustercancellation

END

theta\_max<\-n\*\(amplification\_budget\*\(p\_max\-1\)\!\)^\(1/\(p\_max\-1\)\)

REQUIREthecharacteristicfunctionisnegligibleattheta\_max

phi\(theta\)<\-\(F\_pair\(theta\)/F\_pair\(0\)\)^nFOR\|theta\|<=theta\_max

density<\-REAL\(FFT\(phi\)\)/period

REQUIREconservationofmass,firstmoment,andnoise\-floorcontrol

INTEGRATEA\[j,l\]andB\[j,l\]againstdensity

END

PROFILEthelinearcoefficientsbyTOP\_RITZ\(A,B\)

OPTIMIZEcbyBrent\(onecluster\)orfinite\-differenceLBFGS\(multiple\)

RETURNthebestcandidatere\-evaluatedonafinergrid

END

### 5\.20Probabilistic reduction for the rational trial

Fork≥3k\\geq 3, letn=k−1n=k\-1, choose an exact rationalc\>0c\>0, and define

F⁡\(t1,…,tk\)=∏i=1kg⁡\(ti\)​𝟏ℛk​\(t\),g⁡\(t\)=1c\+n​t,F\(t\_\{1\},\\ldots,t\_\{k\}\)=\\prod\_\{i=1\}^\{k\}g\(t\_\{i\}\)\\mathbf\{1\}\_\{\\mathcal\{R\}\_\{k\}\}\(t\),\\qquad g\(t\)=\\frac\{1\}\{c\+nt\},\(43\)whereℛk=\{ti≥0:∑iti≤1\}\\mathcal\{R\}\_\{k\}=\\\{t\_\{i\}\\geq 0:\\sum\_\{i\}t\_\{i\}\\leq 1\\\}\. Putw=g2w=g^\{2\}and

m0=∫01w⁡\(t\)​𝑑t=1c⁡\(c\+n\)\.m\_\{0\}=\\int\_\{0\}^\{1\}w\(t\)\\,dt=\\frac\{1\}\{c\(c\+n\)\}\.LetS=SnS=S\_\{n\}denote the sum ofnnindependent variables with densityw/m0w/m\_\{0\}on\[0,1\]\[0,1\]\. Then

G⁡\(ρ\)=∫0ρg⁡\(t\)​𝑑t=log⁡\(1\+n​ρ/c\)n,H⁡\(ρ\)=∫0ρw⁡\(t\)​𝑑t=1n​\(1c−1c\+n​ρ\)\.G\(\\rho\)=\\int\_\{0\}^\{\\rho\}g\(t\)\\,dt=\\frac\{\\log\(1\+n\\rho/c\)\}\{n\},\\qquad H\(\\rho\)=\\int\_\{0\}^\{\\rho\}w\(t\)\\,dt=\\frac\{1\}\{n\}\\left\(\\frac\{1\}\{c\}\-\\frac\{1\}\{c\+n\\rho\}\\right\)\.\(44\)Peeling one coordinate from the Maynard integrals gives

Ik​\(F\)=m0n​D,Jk\(1\)​\(F\)=m0n​N,I\_\{k\}\(F\)=m\_\{0\}^\{n\}D,\\qquad J\_\{k\}^\{\(1\)\}\(F\)=m\_\{0\}^\{n\}N,\(45\)where

N=𝔼⁡\[G​\(1−S\)2​𝟏S<1\],D=𝔼⁡\[H⁡\(1−S\)​𝟏S<1\],N=\\mathbb\{E\}\[G\(1\-S\)^\{2\}\\mathbf\{1\}\_\{S<1\}\],\\qquad D=\\mathbb\{E\}\[H\(1\-S\)\\mathbf\{1\}\_\{S<1\}\],\(46\)and consequentlyMk≥k​N/DM\_\{k\}\\geq kN/D\.

Let

w^​\(θ\)=∫01ei​θ​t​w​\(t\)​𝑑t,φ⁡\(θ\)=\(w^​\(θ\)m0\)n,\\widehat\{w\}\(\\theta\)=\\int\_\{0\}^\{1\}e^\{i\\theta t\}w\(t\)\\,dt,\\qquad\\varphi\(\\theta\)=\\left\(\\frac\{\\widehat\{w\}\(\\theta\)\}\{m\_\{0\}\}\\right\)^\{n\},and, forX∈\{G2,H\}X\\in\\\{G^\{2\},H\\\},

ΨX​\(θ\)=∫01X⁡\(ρ\)​ei​θ​ρ​𝑑ρ\.\\Psi\_\{X\}\(\\theta\)=\\int\_\{0\}^\{1\}X\(\\rho\)e^\{i\\theta\\rho\}\\,d\\rho\.The required expectation has the Fourier representation

IX=12​π​∫ℝφ⁡\(θ\)​e−i​θ​ΨX​\(θ\)​𝑑θ\.I\_\{X\}=\\frac\{1\}\{2\\pi\}\\int\_\{\\mathbb\{R\}\}\\varphi\(\\theta\)e^\{\-i\\theta\}\\Psi\_\{X\}\(\\theta\)\\,d\\theta\.\(47\)The integrability needed for inversion follows from integration\-by\-parts decay of both factors\.

### 5\.21Aliasing bound

FixP\>2P\>2andΔ​θ=2​π/P\\Delta\\theta=2\\pi/P\. For the full trapezoidal sumTXT\_\{X\}, Poisson summation and compact support imply

0≤TX−IX≤Xmax​ℙ​\(S\>P−1\)\.0\\leq T\_\{X\}\-I\_\{X\}\\leq X\_\{\\max\}\\mathbb\{P\}\(S\>P\-1\)\.\(48\)The one\-sided sign follows from nonnegativity of the aliased density\. The residual probability is bounded by certified moments\. WithM~0=1\\widetilde\{M\}\_\{0\}=1and

M~r=∑j=1r\(r−1j−1\)​n​𝔼​\[tj\]​M~r−j,\\widetilde\{M\}\_\{r\}=\\sum\_\{j=1\}^\{r\}\\binom\{r\-1\}\{j\-1\}n\\,\\mathbb\{E\}\[t^\{j\}\]\\,\\widetilde\{M\}\_\{r\-j\},\(49\)the supplied compound\-Poisson domination gives

ℙ⁡\(S\>y\)≤min1≤r≤rmax⁡M~ryr\.\\mathbb\{P\}\(S\>y\)\\leq\\min\_\{1\\leq r\\leq r\_\{\\max\}\}\\frac\{\\widetilde\{M\}\_\{r\}\}\{y^\{r\}\}\.\(50\)E⁡\[Sr\]E\[S^\{r\}\]expands over set partitions with falling\-factorial coefficientsnb¯≤nbn^\{\\underline\{b\}\}\\leq n^\{b\}, and the compound\-Poisson moments have exactlynbn^\{b\}with all terms nonnegative\.

### 5\.22Fourier truncation bound

Integration by parts yields, forθ≠0\\theta\\neq 0,

\|w^​\(θ\)m0\|≤A\|θ\|,A=2​\(c\+n\)c,\\left\|\\frac\{\\widehat\{w\}\(\\theta\)\}\{m\_\{0\}\}\\right\|\\leq\\frac\{A\}\{\|\\theta\|\},\\qquad A=\\frac\{2\(c\+n\)\}\{c\},\(51\)and hence\|φ⁡\(θ\)\|≤\(A/\|θ\|\)n\|\\varphi\(\\theta\)\|\\leq\(A/\|\\theta\|\)^\{n\}\. The omitted positive frequencies are divided into a finite band and an analytic tail\. In the finite band, indices are partitioned into blocks\[aj,bj\]\[a\_\{j\},b\_\{j\}\]and Arb wide\-ball evaluation supplies certified supremaσj\\sigma\_\{j\}for\|w^/m0\|\|\\widehat\{w\}/m\_\{0\}\|on each block\. Thus

\|TXband\|≤Δ​θπ​Xint​∑j\(bj−aj\+1\)​σjn\.\|T\_\{X\}^\{\\rm band\}\|\\leq\\frac\{\\Delta\\theta\}\{\\pi\}X\_\{\\rm int\}\\sum\_\{j\}\(b\_\{j\}\-a\_\{j\}\+1\)\\sigma\_\{j\}^\{n\}\.\(52\)For

M0=max⁡\(⌊Θball/Δ​θ⌋,nnear\),Θball=2​A,M\_\{0\}=\\max\\\!\\left\(\\left\\lfloor\\Theta\_\{\\rm ball\}/\\Delta\\theta\\right\\rfloor,n\_\{\\rm near\}\\right\),\\qquad\\Theta\_\{\\rm ball\}=2A,the remaining tail satisfies

\|TXtail\|≤Δ​θπ​Xint​\(AΔ​θ\)n​M01−nn−1\.\|T\_\{X\}^\{\\rm tail\}\|\\leq\\frac\{\\Delta\\theta\}\{\\pi\}X\_\{\\rm int\}\\left\(\\frac\{A\}\{\\Delta\\theta\}\\right\)^\{n\}\\frac\{M\_\{0\}^\{1\-n\}\}\{n\-1\}\.\(53\)TakingM0≥nnearM\_\{0\}\\geq n\_\{\\rm near\}ensures that the analytic tail begins where the computed sum ends\.

### 5\.23Arb evaluation and directed rounding

Every scalar entering the computation is produced by Arb ball arithmetic\. The transformw^/m0\\widehat\{w\}/m\_\{0\}is evaluated from a closed form involving sine and cosine integrals away from the origin and by direct rigorous integration in the region where the special\-function representation is ill\-conditioned\. The transformsΨX\\Psi\_\{X\}and the moments entering the aliasing bound are likewise evaluated by rigorous integration\. The pole and logarithmic branch point lie att=−c/n<0t=\-c/n<0, outside the real integration interval; subdivision is used whenever an evaluation enclosure approaches a singularity or branch cut\.

All error terms remain in ball arithmetic until the final quotient is assembled\. Floating\-point exports are explicitly nudged until Arb comparisons prove outward rounding\. IfNballN\_\{\\rm ball\}andDballD\_\{\\rm ball\}denote the truncated\-sum enclosures andEN,EDE\_\{N\},E\_\{D\}the certified totals of the aliasing and truncation errors, then

Rlow=lb⁡\(k​lb⁡\(Nball−EN\)ub⁡\(Dball\+ED\)\)\.R\_\{\\rm low\}=\\operatorname\{lb\}\\\!\\left\(\\frac\{k\\,\\operatorname\{lb\}\(N\_\{\\rm ball\}\-E\_\{N\}\)\}\{\\operatorname\{ub\}\(D\_\{\\rm ball\}\+E\_\{D\}\)\}\\right\)\.\(54\)The stored claim is additionally rounded downward so that independent recomputation with different valid Arb subdivisions or enclosure radii cannot strengthen the stated result\.

### 5\.24Certified rank\-one ladder and uniform finite\-range bound

Forn=k−1n=k\-1andc\>0c\>0, let

gc​\(t\)=1c\+n​t,g\_\{c\}\(t\)=\\frac\{1\}\{c\+nt\},and denote byRk\(1\)R\_\{k\}^\{\(1\)\}the value attained by this rank\-one rational family\. Every such value is a lower bound forMkM\_\{k\}\.

The fixed headroom value used below was selected from the independently optimized rank\-one sweep rather than imposed a priori\. Writing

η⁡\(c,k\)=1−𝔼⁡\[S\]sd⁡\(S\),\\eta\(c,k\)=\\frac\{1\-\\mathbb\{E\}\[S\]\}\{\\operatorname\{sd\}\(S\)\},the median optimized value ofη\\etamoves steadily toward approximately−0\.475\-0\.475askkincreases:

range of​kmedian⁡ηk<103−0\.4256103≤k<105−0\.4593105≤k<107−0\.4715107≤k<109−0\.4747\\begin\{array\}\[\]\{c\|c\}\\text\{range of \}k&\\operatorname\{median\}\\eta\\\\ \\hline\\cr k<10^\{3\}&\-0\.4256\\\\ 10^\{3\}\\leq k<10^\{5\}&\-0\.4593\\\\ 10^\{5\}\\leq k<10^\{7\}&\-0\.4715\\\\ 10^\{7\}\\leq k<10^\{9\}&\-0\.4747\\end\{array\}This convergence motivated the fixed choice

η∗=−0\.4746,\\eta^\{\\ast\}=\-0\.4746,which agrees most closely with the independently optimized scale in the large\-kkregime containing the principal threshold results\.

Fixingη\\etaincurs very little loss for the bulk of the optimized sweep\. Interpolating the fixed\-headroom ladder to thekk\-values of the direct rank\-one sweep, the optimized certified value exceeds the corresponding fixed\-η\\etavalue by a median of only

4\.2×10−5,4\.2\\times 10^\{\-5\},with maximum positive difference

4\.2×10−3\.4\.2\\times 10^\{\-3\}\.The comparison is not uniformly positive: the minimum difference is

−4\.2×10−2,\-4\.2\\times 10^\{\-2\},so at least one nominally optimized sweep row is outperformed by the fixed\-headroom construction at the samekk\. This is consistent with the isolated high\-kkoutlier nearη=−0\.32\\eta=\-0\.32, and is one reason the direct sweep is interpreted as a record of the numerical optimization landscape rather than as a certified sequence of exact one\-parameter optima\.

To obtain a uniform finite\-range statement rather than isolated certified dimensions, we constructed an approximately log\-uniform grid beginning atk=100k=100, with target spacing

Δ​log⁡k=0\.02\.\\Delta\\log k=0\.02\.At each grid point,ccwas determined by solving

η⁡\(c,k\)=η∗=−0\.4746\\eta\(c,k\)=\\eta^\{\\ast\}=\-0\.4746using the closed\-form moments of the rank\-one probability measure\. The resulting explicit trial function was then certified independently using the ball\-arithmetic procedure of Section[5\.23](https://arxiv.org/html/2609.30296#S5.SS23)\. All retained rungs passed certification\.

For consecutive certified grid pointski<ki\+1k\_\{i\}<k\_\{i\+1\}, monotonicity ofMkM\_\{k\}implies that every integerk∈\[ki,ki\+1\]k\\in\[k\_\{i\},k\_\{i\+1\}\]satisfies

Mk≥Mki≥Rcert\(1\)​\(ki\)≥log⁡k−\[log⁡ki\+1−Rcert\(1\)​\(ki\)\]\.M\_\{k\}\\geq M\_\{k\_\{i\}\}\\geq R\_\{\\rm cert\}^\{\(1\)\}\(k\_\{i\}\)\\geq\\log k\-\\left\[\\log k\_\{i\+1\}\-R\_\{\\rm cert\}^\{\(1\)\}\(k\_\{i\}\)\\right\]\.Thus the relevant covering deficit on the interval\[ki,ki\+1\]\[k\_\{i\},k\_\{i\+1\}\]is

δi=log⁡ki\+1−Rcert\(1\)​\(ki\)\.\\delta\_\{i\}=\\log k\_\{i\+1\}\-R\_\{\\rm cert\}^\{\(1\)\}\(k\_\{i\}\)\.
The largest interval deficit over the certified ladder is

maxi⁡δi=0\.306754936,\\max\_\{i\}\\delta\_\{i\}=0\.306754936,attained on the penultimate full\-size interval

ki=1,834,543,653,ki\+1=1,880,000,000\.k\_\{i\}=1\{,\}834\{,\}543\{,\}653,\\qquad k\_\{i\+1\}=1\{,\}880\{,\}000\{,\}000\.The maximum does not occur on the final interval because the last grid step was deliberately shortened to terminate at the round endpoint

k=1\.88×109\.k=1\.88\\times 10^\{9\}\.Its logarithmic width is only

Δ​log⁡k=0\.004476,\\Delta\\log k=0\.004476,substantially smaller than the nominal spacing0\.020\.02\. The final truncated interval therefore contributes a smaller covering deficit than the preceding full\-size rung\.

The distinction between pointwise and interval\-covering deficits is also important\. On the certified grid points themselves, the largest observed pointwise deficit is

maxi⁡\[log⁡ki−Rcert\(1\)​\(ki\)\]=0\.286812601\.\\max\_\{i\}\\left\[\\log k\_\{i\}\-R\_\{\\rm cert\}^\{\(1\)\}\(k\_\{i\}\)\\right\]=0\.286812601\.Hence the certified rungs satisfy the sharper pointwise statement

Rcert\(1\)​\(ki\)≥log⁡ki−0\.287\.R\_\{\\rm cert\}^\{\(1\)\}\(k\_\{i\}\)\\geq\\log k\_\{i\}\-0\.287\.The larger constant0\.3070\.307is required only to cover every intermediate integerkkbetween successive certified rungs\.

##### Consistency of the threshold certificates with the ladder\.

The four threshold certificates reported in the Results lie inside the range covered by this ladder, and were produced independently of it: the pole parameterccwas optimised at each threshold rather than selected from the fixed\-headroom relation, and each was certified as an isolated trial function\. They therefore provide an internal check on the finite\-range behaviour rather than additional coverage\.

Supplementary Table[4](https://arxiv.org/html/2609.30296#S5.T4)records the certified deficitlog⁡k−Rcert\(1\)\\log k\-R^\{\(1\)\}\_\{\\mathrm\{cert\}\}at each threshold, together with the value obtained from the empirical finite\-size relation

log⁡k−Rk\(1\)≈𝒞∗−1\.047log⁡k,𝒞∗≃0\.3343,\\log k\-R^\{\(1\)\}\_\{k\}\\;\\approx\\;\\mathcal\{C\}\_\{\*\}\-\\frac\{1\.047\}\{\\log k\},\\qquad\\mathcal\{C\}\_\{\*\}\\simeq 0\.3343,\(55\)which was introduced in Section[5\.25](https://arxiv.org/html/2609.30296#S5.SS25)from the optimised rank\-one sweep\. Agreement is within3×10−33\\times 10^\{\-3\}at every point, across more than five orders of magnitude inkk, and the final row shows that the upper end of the certified ladder follows the same relation\.

Supplementary Table 4:Certified deficits at the four thresholds and at the upper end of the fixed\-headroom ladder, compared with the finite\-size relation Eq\. \([55](https://arxiv.org/html/2609.30296#S5.E55)\)\. Thresholds were optimised and certified independently of the ladder; the final row is the pointwise deficit at the last ladder rung\. The relation was not fitted to these values\.Two remarks on the status of this comparison\. First, it is a consistency check and not a proof component: the constant𝒞∗\\mathcal\{C\}\_\{\*\}is a numerical evaluation of the limiting formula, and the coefficient1\.0471\.047is an empirical finite\-size fit for which Section[5\.26](https://arxiv.org/html/2609.30296#S5.SS26)establishes no quantitative rate\. Neither enters any certified inequality\. Second, the agreement is nevertheless informative precisely because nothing was tuned to produce it: the thresholds were optimised for a different purpose \(crossing the integer values88,1212,1616and2020\), the ladder was generated from a fixed headroom target, and the relation \([55](https://arxiv.org/html/2609.30296#S5.E55)\) was extracted from a third sweep\. That three independently constructed families of certified rank\-one trials trace the same𝒞∗−b/log⁡k\\mathcal\{C\}\_\{\*\}\-b/\\log kcurve supports the interpretation of the finite\-kkbehaviour as the pre\-asymptotic regime of the stable\-law limit rather than as a numerical coincidence\. The single deviation of appreciable size, atk=3655k=3655, is in the conservative direction and is consistent with the relation being fitted at largerkk\.

The interval constant is close to saturated at the upper end of the computed range\. The final four full\-size interval deficits are

0\.306613,0\.306660,0\.306708,0\.306755,0\.306613,\\qquad 0\.306660,\\qquad 0\.306708,\\qquad 0\.306755,increasing by approximately

4\.7×10−54\.7\\times 10^\{\-5\}per rung\. Continuing the same fixed\-headroom ladder for only several further full\-size steps would therefore push the covering deficit above0\.3070\.307\. Accordingly, rounding the observed maximum upward gives the finite\-range statement

Mk≥logk−0\.307,100≤k≤1,880,000,000\.\\boxed\{M\_\{k\}\\geq\\log k\-0\.307,\\qquad 100\\leq k\\leq 1\{,\}880\{,\}000\{,\}000\.\}No extension of the constant0\.3070\.307beyond this explicitly certified range is claimed\.

### 5\.25Rank\-one scale selection and the headroom invariant

Under the probability measure proportional togc​\(t\)2​d​tg\_\{c\}\(t\)^\{2\}\\,dt, letS=∑i<ktiS=\\sum\_\{i<k\}t\_\{i\}and write

c−1=log⁡n\+a,n=k−1\.c^\{\-1\}=\\log n\+a,\\qquad n=k\-1\.Uniformly fora=O⁡\(log⁡n\)a=O\(\\sqrt\{\\log n\}\), the closed\-form moments of this measure imply

1−𝔼⁡\[S\]=c⁡\(a\+1−log⁡c−1\)\+O⁡\(c2​log⁡nn\),sd⁡\(S\)=c​\(1\+O⁡\(log⁡nn\)\)\.1\-\\mathbb\{E\}\[S\]=c\\left\(a\+1\-\\log c^\{\-1\}\\right\)\+O\\\!\\left\(\\frac\{c^\{2\}\\log n\}\{n\}\\right\),\\qquad\\operatorname\{sd\}\(S\)=\\sqrt\{c\}\\left\(1\+O\\\!\\left\(\\frac\{\\log n\}\{n\}\\right\)\\right\)\.\(56\)For fixedaa,log⁡c−1=log⁡log⁡n\+O⁡\(1/log⁡n\)\\log c^\{\-1\}=\\log\\log n\+O\(1/\\log n\)\. Consequently,

η⁡\(c,k\)=1−𝔼⁡\[S\]sd⁡\(S\)=c​\(a\+1−log⁡c−1\)\+O⁡\(log⁡nn\)\.\\eta\(c,k\)=\\frac\{1\-\\mathbb\{E\}\[S\]\}\{\\operatorname\{sd\}\(S\)\}=\\sqrt\{c\}\\left\(a\+1\-\\log c^\{\-1\}\\right\)\+O\\\!\\left\(\\frac\{\\log n\}\{n\}\\right\)\.\(57\)Fixingη=−γh\\eta=\-\\gamma\_\{\\mathrm\{h\}\}, multiplying Eq\. \([57](https://arxiv.org/html/2609.30296#S5.E57)\) byc−1/2c^\{\-1/2\}and substitutinga=c−1−log⁡na=c^\{\-1\}\-\\log ngives the implicit scale equation

c−1\+γhc−1/2−logc−1=logn−1\+O\(\(log⁡n\)3/2n\)\.c^\{\-1\}\+\\gamma\_\{\\mathrm\{h\}\}\\,c^\{\-1/2\}\-\\log c^\{\-1\}=\\log n\-1\+O\\\!\\left\(\\frac\{\(\\log n\)^\{3/2\}\}\{n\}\\right\)\.\(58\)The left\-hand side is strictly increasing inc−1c^\{\-1\}forc−1≥1c^\{\-1\}\\geq 1, so Eq\. \([58](https://arxiv.org/html/2609.30296#S5.E58)\) determines the fixed\-headroom scale to the stated accuracy\. Writingx=c−1x=c^\{\-1\}, it givesx=log⁡n\+O⁡\(log⁡n\)x=\\log n\+O\(\\sqrt\{\\log n\}\), hencelogx=loglogn\+O\(\(logn\)−1/2\)\\log x=\\log\\log n\+O\\big\(\(\\log n\)^\{\-1/2\}\\big\)andx=log⁡n−γh/2\+O⁡\(log⁡log⁡n/log⁡n\)\\sqrt\{x\}=\\sqrt\{\\log n\}\-\\gamma\_\{\\mathrm\{h\}\}/2\+O\\big\(\\log\\log n/\\sqrt\{\\log n\}\\big\)\. Substituting these expansions yields

c−1=log⁡n−γh​log⁡n\+log⁡log⁡n−1\+γh22\+O⁡\(log⁡log⁡nlog⁡n\)\.c^\{\-1\}=\\log n\-\\gamma\_\{\\mathrm\{h\}\}\\sqrt\{\\log n\}\+\\log\\log n\-1\+\\frac\{\\gamma\_\{\\mathrm\{h\}\}^\{2\}\}\{2\}\+O\\\!\\left\(\\frac\{\\log\\log n\}\{\\sqrt\{\\log n\}\}\\right\)\.\(59\)The constantγh2/2\\gamma\_\{\\mathrm\{h\}\}^\{2\}/2arises from the shift ofc−1/2c^\{\-1/2\}relative tolog⁡n\\sqrt\{\\log n\}\. The remainder decays only likelog⁡log⁡n/log⁡n\\log\\log n/\\sqrt\{\\log n\}and is not small over the certified range \(log⁡log⁡n/log⁡n≈0\.66\\log\\log n/\\sqrt\{\\log n\}\\approx 0\.66atlog⁡n≈21\.35\\log n\\approx 21\.35\)\. The ladder was therefore computed from the exact closed\-form moments, which agree with Eq\. \([58](https://arxiv.org/html/2609.30296#S5.E58)\) to the stated accuracy, and not from the truncated expansion \([59](https://arxiv.org/html/2609.30296#S5.E59)\)\. Thus the empirically selected coefficientγh≃0\.4746\\gamma\_\{\\mathrm\{h\}\}\\simeq 0\.4746is not an independent asymptotic fit: it is the same quantity as−η∗\-\\eta^\{\\ast\}and follows directly from the fixed\-headroom parameterization\.

Equations \([58](https://arxiv.org/html/2609.30296#S5.E58)\) and \([59](https://arxiv.org/html/2609.30296#S5.E59)\) describe the fixed\-headroom parameterization of the certified ladder; they are not a sharper asymptotic optimizer law\. Indeed, fixedη=−γh\\eta=\-\\gamma\_\{\\mathrm\{h\}\}corresponds to the shift

a⁡\(L\)=c−1−L=log⁡L−1−γh​L\+γh22\+O⁡\(log⁡LL\),L=log⁡n,a\(L\)=c^\{\-1\}\-L=\\log L\-1\-\\gamma\_\{\\mathrm\{h\}\}\\sqrt\{L\}\+\\frac\{\\gamma\_\{\\mathrm\{h\}\}^\{2\}\}\{2\}\+O\\\!\\left\(\\frac\{\\log L\}\{\\sqrt\{L\}\}\\right\),\\qquad L=\\log n,soa⁡\(L\)→−∞a\(L\)\\to\-\\inftyand fixedη\\etaeventually leaves every fixed\-shift family\. The leading\-order part ofa⁡\(L\)a\(L\)is not monotone: its derivative1/L−γh/\(2​L\)1/L\-\\gamma\_\{\\mathrm\{h\}\}/\(2\\sqrt\{L\}\)vanishes at

L=2γh,\\sqrt\{L\}=\\frac\{2\}\{\\gamma\_\{\\mathrm\{h\}\}\},where it attains its maximum, and it decreases thereafter\. Over the certified range, however, the induced shift obtained from Eq\. \([58](https://arxiv.org/html/2609.30296#S5.E58)\) remains close to the small negative values where the numerical limiting defect𝒞⁡\(a\)\\mathcal\{C\}\(a\)is minimized\. This explains the effectiveness of the headroom surrogate over the finite range studied here: it selects rational trials lying close to the same near\-optimal fixed\-shift regime, even though the two parameterizations diverge asymptotically\.

At the upper end of the certified ladder, the measured pointwise deficit is approximately

log⁡k−R⁡\(Fk,ck\)≃0\.2868,\\log k\-R\(F\_\{k,c\_\{k\}\}\)\\simeq 0\.2868,whereas numerical evaluation of the limiting fixed\-shift formula gives

𝒞∗≃0\.3343\.\\mathcal\{C\}\_\{\*\}\\simeq 0\.3343\.The difference is approximately

0\.3343−0\.2868=0\.0475\.0\.3343\-0\.2868=0\.0475\.Over the finite range considered in the numerical sweep, this is close to the magnitude of the empirically fitted correction

1\.047log⁡k≃0\.049\(log⁡k≃21\.35\)\.\\frac\{1\.047\}\{\\log k\}\\simeq 0\.049\\qquad\(\\log k\\simeq 21\.35\)\.Equivalently, the observed deficits are numerically consistent with the finite\-range relation

log⁡k−R⁡\(Fk,ck\)≈𝒞∗−1\.047log⁡k\.\\log k\-R\(F\_\{k,c\_\{k\}\}\)\\approx\\mathcal\{C\}\_\{\*\}\-\\frac\{1\.047\}\{\\log k\}\.This relation is an empirical consistency check only\. In particular, the coefficient1\.0471\.047has not been derived analytically and is not identified with the coefficient of a proved second\-order expansion\.

The distinction is important because the stable\-law theorem concerns the fixed\-shift family

c−1=log⁡n\+a,a∈ℝ​fixed,c^\{\-1\}=\\log n\+a,\\qquad a\\in\\mathbb\{R\}\\ \\text\{fixed\},whereas the fixed\-headroom ladder corresponds to a shifta=a⁡\(log⁡n\)a=a\(\\log n\)that eventually tends to−∞\-\\infty\. Therefore the finite\-size expansion for fixedaacannot be applied directly to the optimized headroom ladder\. Even within the fixed\-shift family, the proof establishes only

log⁡k−R⁡\(Fk,ck\)=𝒞⁡\(a\)\+o⁡\(1\),\\log k\-R\(F\_\{k,c\_\{k\}\}\)=\\mathcal\{C\}\(a\)\+o\(1\),because no sufficiently sharp convergence rate has been proved for

LnPn−LaPa\.\\frac\{L\_\{n\}\}\{P\_\{n\}\}\-\\frac\{L\_\{a\}\}\{P\_\{a\}\}\.Consequently, neither the numerical agreement above nor the fixed\-headroom parameterization establishes a rigorous1/log⁡k1/\\log kcorrection\. The coefficient1\.0471\.047should therefore be reported solely as a numerical finite\-size observation\.

### 5\.26Spectrally positive11\-stable limit

Section[6](https://arxiv.org/html/2609.30296#S6)gives the complete asymptotic analysis of the single\-channel rational family\. For

n=k−1,ck−1=log⁡n\+a,n=k\-1,\\qquad c\_\{k\}^\{\-1\}=\\log n\+a,with fixeda∈ℝa\\in\\mathbb\{R\}, the rescaled boundary variable converges in distribution to

Xa=\(a\+1\)−Λ,X\_\{a\}=\(a\+1\)\-\\Lambda,whereΛ\\Lambdais the spectrally positive11\-stable random variable defined by

𝔼ei​s​Λ=exp\{∫0∞\(ei​s​x−1−isx𝟏\{x≤1\}\)d​xx2\}\.\\mathbb\{E\}e^\{is\\Lambda\}=\\exp\\left\\\{\\int\_\{0\}^\{\\infty\}\\left\(e^\{isx\}\-1\-isx\\mathbf\{1\}\_\{\\\{x\\leq 1\\\}\}\\right\)\\frac\{dx\}\{x^\{2\}\}\\right\\\}\.The proof supplements weak convergence with uniform density, small\-ball, tail, and logarithmic moment estimates, which are needed because the Rayleigh quotient contains logarithmic observables at the simplex boundary\.

Theorem[9](https://arxiv.org/html/2609.30296#Thmtheorem9)consequently gives

R⁡\(Fk,ck\)=log⁡k−𝒞⁡\(a\)\+o⁡\(1\),R\(F\_\{k,c\_\{k\}\}\)=\\log k\-\\mathcal\{C\}\(a\)\+o\(1\),where

𝒞⁡\(a\)=a−2​𝔼​\[log⁡Xa∣Xa\>0\]\.\\mathcal\{C\}\(a\)=a\-2\\,\\mathbb\{E\}\[\\log X\_\{a\}\\mid X\_\{a\}\>0\]\.Thus, with

𝒞∗:=infa∈ℝ𝒞⁡\(a\),\\mathcal\{C\}\_\{\*\}:=\\inf\_\{a\\in\\mathbb\{R\}\}\\mathcal\{C\}\(a\),Corollary[10](https://arxiv.org/html/2609.30296#Thmtheorem10)yields

Mk≥log⁡k−𝒞∗−o⁡\(1\)\.M\_\{k\}\\geq\\log k\-\\mathcal\{C\}\_\{\*\}\-o\(1\)\.
Here𝒞∗\\mathcal\{C\}\_\{\*\}is the optimal asymptotic defect only within the fixed\-shift family

ck−1=log⁡\(k−1\)\+a,a∈ℝ\.c\_\{k\}^\{\-1\}=\\log\(k\-1\)\+a,\\qquad a\\in\\mathbb\{R\}\.The proof does not establish optimality over arbitrary sequencesckc\_\{k\}, nor does it require the infimum to be attained\. Numerical evaluation suggests

𝒞∗≈0\.3343,\\mathcal\{C\}\_\{\*\}\\approx 0\.3343,but this decimal value should be regarded as a numerical estimate unless the global infimum is separately certified\. A rigorous enclosure of𝒞⁡\(a0\)\\mathcal\{C\}\(a\_\{0\}\)at one fixed shifta0a\_\{0\}already suffices for an explicit asymptotic lower bound\.

Finally, although the finite\-nnidentity in Section[6](https://arxiv.org/html/2609.30296#S6)isolates a possible1/log⁡k1/\\log kcorrection, the required convergence rate for

LnPn−LaPa\\frac\{L\_\{n\}\}\{P\_\{n\}\}\-\\frac\{L\_\{a\}\}\{P\_\{a\}\}has not been proved\. Accordingly, no second\-order asymptotic coefficient is claimed here\.

##### Relation to the optimized rank\-one sweep\.

The optimized rank\-one sweep explores a different asymptotic parameter regime, including an empirically observedlog⁡n\\sqrt\{\\log n\}correction\. Those numerical observations are not used in the fixed\-shift theorem above\.

![Refer to caption](https://arxiv.org/html/2609.30296v1/Section1/images/figure_crossover.png)Supplementary Figure 3:Enlarged\-support crossover\.\(a\)The certified differenceΔcert​\(k,ε\)=Lk,ε−Lk,0\\Delta\_\{\\mathrm\{cert\}\}\(k,\\varepsilon\)=L\_\{k,\\varepsilon\}\-L\_\{k,0\}, a difference of certified lower bounds rather than of the exact constants, againstkkfor four values ofε\\varepsilon\. It is positive at smallkk, decreases withkk, and changes sign at a finite scale forε=1/12\\varepsilon=1/12and1/161/16; no sign change is reached forε=1/50\\varepsilon=1/50within the computed range\. Grey crosses mark certified but non\-competitive runs excluded from crossover inference \(Section[5\.27](https://arxiv.org/html/2609.30296#S5.SS27)\)\.\(b\)The localised crossoversk∗​\(ε\)k^\{\*\}\(\\varepsilon\)against1/ε1/\\varepsilonon logarithmic axes\. Theε=1/12\\varepsilon=1/12and1/161/16points are linear interpolations between the certified brackets; theε=1/25\\varepsilon=1/25marker is the midpoint of the extrapolative range\[155,192\]\[155,192\], with the bar spanning that range\. The points lie above the linear referencek∗∝1/εk^\{\*\}\\propto 1/\\varepsilon\(dotted\) and pull away from it, givingε​k∗=4\.48,5\.04,6\.94\\varepsilon k^\{\*\}=4\.48,\\,5\.04,\\,6\.94and motivatingε​k∗​\(ε\)→∞\\varepsilon k^\{\*\}\(\\varepsilon\)\\to\\infty\. The band shows the familyk∗/\(log⁡k∗\)p∝1/εk^\{\*\}/\(\\log k^\{\*\}\)^\{p\}\\propto 1/\\varepsilonover the rangep∈\[1\.4,2\.2\]p\\in\[1\.4,2\.2\]admitted by the data, with the constant refitted at eachpp; the exponent is not determined by these points\.

### 5\.27Enlarged\-support crossover analysis

Across the ladder50≤k≤22050\\leq k\\leq 220, the advantage conferred byε\\varepsilonfluctuates at smallkk, then decreases and changes sign at a finite crossoverk∗​\(ε\)k\_\{\\ast\}\(\\varepsilon\), which we locate atk∗​\(1/12\)≈53\.7k\_\{\\ast\}\(1/12\)\\approx 53\.7,k∗​\(1/16\)≈80\.7k\_\{\\ast\}\(1/16\)\\approx 80\.7andk∗​\(1/25\)∈\[155,192\]k\_\{\\ast\}\(1/25\)\\in\[155,192\]\. We conjecture that such a crossover exists for everyε\>0\\varepsilon\>0and thatk∗​\(ε\)k\_\{\\ast\}\(\\varepsilon\)grows superlinearly in1/ε1/\\varepsilon, so thatε​k∗​\(ε\)→∞\\varepsilon\\,k\_\{\\ast\}\(\\varepsilon\)\\to\\infty\(Supplementary Table[9](https://arxiv.org/html/2609.30296#S11.T9)\)\.

WriteLk,εL\_\{k,\\varepsilon\}for the certified lower bound produced for the enlarged\-domain construction andLk,0L\_\{k,0\}for the corresponding vanilla lower bound, and define

Δcert​\(k,ε\)=Lk,ε−Lk,0\.\\Delta\_\{\\rm cert\}\(k,\\varepsilon\)=L\_\{k,\\varepsilon\}\-L\_\{k,0\}\.For the reliable computed configurations, this difference is positive at smallerkk, decreases withkk, and changes sign at a finite scale\. The certified pairs used to localize the observed crossovers include

εkleftkright1/1250601/168090\\begin\{array\}\[\]\{c\|cc\}\\varepsilon&k\_\{\\rm left\}&k\_\{\\rm right\}\\\\ \\hline\\cr 1/12&50&60\\\\ 1/16&80&90\\end\{array\}with certified differences\+0\.016961\+0\.016961and−0\.028948\-0\.028948forε=1/12\\varepsilon=1/12, and\+0\.003771\+0\.003771and−0\.049093\-0\.049093forε=1/16\\varepsilon=1/16\. Linear interpolation gives

k∗​\(1/12\)≈53\.7,k∗​\(1/16\)≈80\.7\.k^\{\*\}\(1/12\)\\approx 53\.7,\\qquad k^\{\*\}\(1/16\)\\approx 80\.7\.Forε=1/25\\varepsilon=1/25, the last reliable positive certified points arek=130k=130and150150, with differences\+0\.035443\+0\.035443and\+0\.006581\+0\.006581; the supplied analysis reports the broader extrapolative range

k∗​\(1/25\)∈\[155,192\]k^\{\*\}\(1/25\)\\in\[155,192\]rather than treating a degradedk=200k=200pair as a direct bracket\. No sign change was reached forε=1/50\\varepsilon=1/50within the reliable computed range\.

All certified runs are in Supplementary Table[9](https://arxiv.org/html/2609.30296#S11.T9)\. Several runs are excluded from crossover inference because they satisfy numerical certification of the particular trial function but are demonstrably noncompetitive discovery outcomes\. These include a reproducible local failure at\(k,ε\)=\(100,1/25\)\(k,\\varepsilon\)=\(100,1/25\), degraded paired runs atk=200k=200, and incomplete or nonmonotone optimization atk=140k=140and220220\. The exclusion criterion is therefore not failure of the certificate: it is failure of the discovery stage to provide a competitive construction\. Monotonicity inkk, the discovery\-to\-certification gap and retained effective rank provide inexpensive diagnostics for these cases\.

The observed crossover locations give increasing values ofε​k∗​\(ε\)\\varepsilon k^\{\*\}\(\\varepsilon\)asε\\varepsilondecreases \(see Supplementary Figure[3](https://arxiv.org/html/2609.30296#S5.F3a)\), motivating the conjecture that for each fixedε\>0\\varepsilon\>0a finite crossover exists and

ε​k∗​\(ε\)→∞\(ε→0\)\.\\varepsilon k^\{\*\}\(\\varepsilon\)\\to\\infty\\qquad\(\\varepsilon\\to 0\)\.The available points are compatible with a scaling variable

x=ε​k\(log⁡k\)p,x=\\frac\{\\varepsilon k\}\{\(\\log k\)^\{p\}\},but do not determinepp: the supplied fits admit approximatelyp∈\[1\.4,2\.2\]p\\in\[1\.4,2\.2\], and the naturalp=1p=1andp=2p=2alternatives cannot be distinguished by the current data\. We therefore do not assign a specific exponent\.

##### Scope of the crossover statement\.

A crucial distinction is thatΔcert\\Delta\_\{\\rm cert\}is a difference of certified lower bounds\. ThereforeΔcert​\(k,ε\)<0\\Delta\_\{\\rm cert\}\(k,\\varepsilon\)<0does not proveMk,ε<MkM\_\{k,\\varepsilon\}<M\_\{k\}; it proves only that the certified vanilla construction found by the pipeline outperforms the certified enlarged\-domain construction found at thatkk\. Extending the observed crossover to the exact variational constants is consequently conjectural\. Likewise, interpolated or extrapolated values ofk∗​\(ε\)k^\{\*\}\(\\varepsilon\)are descriptive summaries of the computed ladder, not independently certified roots\.

### 5\.28Supplementary data: optimized rank\-one sweep

Supplementary Datamaynard\_R\_sweep\_eta\.csvcontains the direct rank\-one optimization sweep for

gc​\(t\)=1c\+\(k−1\)​t,g\_\{c\}\(t\)=\\frac\{1\}\{c\+\(k\-1\)t\},with one optimized pole parameterccfor each reportedkk\. The file records the discovery valueRR, the independently certified lower bound, the discovery–certification gap, the Cauchy–Schwarz ceiling, the optimized scalecc, the standardized simplex headroom

η⁡\(c,k\)=1−𝔼⁡\[S\]sd⁡\(S\),\\eta\(c,k\)=\\frac\{1\-\\mathbb\{E\}\[S\]\}\{\\operatorname\{sd\}\(S\)\},and numerical diagnostics used during optimization\. All retained solutions have rank one in both numerator and denominator Gram representations\.

The sweep contains729729attempted configurations over100≤k≤2×109100\\leq k\\leq 2\\times 10^\{9\}\. Of these,727727pass certification and two are explicitly markedFAILED; failed rows are retained in the data for provenance but are excluded from mathematical claims\. The certified discovery gap is small throughout the successful runs, with median9\.9×10−49\.9\\times 10^\{\-4\}and maximum2\.3×10−32\.3\\times 10^\{\-3\}\. The optimized headroom is not fixed: its variation across the sweep is used diagnostically to identify the approximately constant headroom regime that motivates the geometric ladder below\. This file is therefore a record of the numerical optimization landscape, not itself the basis for the uniform finite\-range theorem\.

### 5\.29Supplementary data: fixed\-headroom geometric ladder

Supplementary Datamaynard\_geometric\_eta\_sweep\.csvcontains the certified fixed\-headroom ladder used to convert isolated rank\-one trial functions into a uniform bound over a continuous range of integerkk\. At each rung,ccis chosen from

η⁡\(c,k\)=−0\.4746,\\eta\(c,k\)=\-0\.4746,and the resulting explicit rank\-one trial is certified independently\. The file contains828828certified rungs fromk=100k=100tok=1\.88×109k=1\.88\\times 10^\{9\}; every row has certification statusPASS\. The target headroom is reproduced to numerical precision throughout the sweep\.

The ladder is approximately uniform inlog⁡k\\log k, with asymptotic spacingΔ​log⁡k≃0\.02\\Delta\\log k\\simeq 0\.02; the small deviations at the lower end and final endpoint arise from integer rounding\. For consecutive certified rungski<ki\+1k\_\{i\}<k\_\{i\+1\}, monotonicity ofMkM\_\{k\}gives

Mk≥Rcert​\(ki\)\(ki≤k≤ki\+1\),M\_\{k\}\\geq R\_\{\\mathrm\{cert\}\}\(k\_\{i\}\)\\qquad\(k\_\{i\}\\leq k\\leq k\_\{i\+1\}\),and therefore

Mk≥log⁡k−\[log⁡ki\+1−Rcert​\(ki\)\]\.M\_\{k\}\\geq\\log k\-\\left\[\\log k\_\{i\+1\}\-R\_\{\\mathrm\{cert\}\}\(k\_\{i\}\)\\right\]\.The largest interval deficit in the supplied data is

0\.3067549357,0\.3067549357,attained between

ki=1 834 543 653,ki\+1=1 871 603 894\.k\_\{i\}=1\\,834\\,543\\,653,\\qquad k\_\{i\+1\}=1\\,871\\,603\\,894\.Rounding this deficit upward yields the reported uniform statement

Mk≥log⁡k−0\.307M\_\{k\}\\geq\\log k\-0\.307throughout the certified range\. The file additionally records the pointwise deficitlog⁡k−Rcert\\log k\-R\_\{\\mathrm\{cert\}\}, the discovery–certification gap, the Cauchy–Schwarz ceiling and the scale diagnostics used to relate the finite ladder to the fixed\-shift asymptotic analysis\.

### 5\.30Conversion to admissible\-tuple bounds

A certified inequalityMk\>4​mM\_\{k\}\>4msupplies the variational input to the Maynard–Tao bounded\-gap theorem\. To obtain a numerical bound onHmH\_\{m\}, one additionally requires an admissiblekk\-tuple and its diameter\. We therefore distinguish the variational certificate from the admissible\-tuple construction and writeDkD\_\{k\}for the diameter of the explicit admissiblekk\-tuple used in the conversion\.

For the first threshold,M3655\>8M\_\{3655\}\>8, a stronger admissible\-tuple diameter is already recorded in the Engelsma–Sutherland database \([https://math\.mit\.edu/~primegaps/](https://math.mit.edu/~primegaps/)\)\. The entry fork=3655k=3655, submitted by Sutherland on 27 June 2013, gives

D3655≤33118\.D\_\{3655\}\\leq 33118\.Our independently constructed hybrid sieve gives the slightly weaker control

D3655≤33270,D\_\{3655\}\\leq 33270,and is retained as a reproducibility check\.

Fork=3655k=3655andk=208910k=208910, admissible tuples were constructed with a hybrid shifted\-Schinzel/greedy sieve\. The algorithm first applies a Schinzel\-type presieve, then greedily removes a least\-populated residue class for successive primes, extracts the narrowest block ofkksurvivors, and finally applies local diameter\-reducing moves\. Crucially, the discovery trajectory is not trusted: the final tuple is rechecked from scratch for every primep≤kp\\leq k\.

Each constructed tuple is exported as anadmissible\-tuple\-certificate/1record containing the ordered gap sequence and, for every primep≤kp\\leq k, an explicitly missing residue class\. A standalone verifier shares no construction code path\. It regenerates the primes up tokk, reconstructs the tuple from the gap list, checks its cardinality and diameter, verifies that the witness list contains exactly the primesp≤kp\\leq k, and confirms by direct enumeration that every stated residue class is absent\. Since akk\-element set cannot occupy all residue classes modulo any primep\>kp\>k, these checks establish admissibility\.

The resulting independently verified hybrid\-sieve bounds are

D3655≤33270,D208910≤2718108\.D\_\{3655\}\\leq 33270,\\qquad D\_\{208910\}\\leq 2718108\.Fork=3655k=3655, the stronger database valueD3655≤33118D\_\{3655\}\\leq 33118is used in the final bounded\-gap comparison, while our independently constructed tuple serves as a reproducibility control\.

For the larger thresholds, explicit window sieving is unnecessary\. Let

ℋk=\{pπ⁡\(k\)\+1,…,pπ⁡\(k\)\+k\},\\mathcal\{H\}\_\{k\}=\\\{p\_\{\\pi\(k\)\+1\},\\ldots,p\_\{\\pi\(k\)\+k\}\\\},the firstkkprimes strictly exceedingkk\. Every element ofℋk\\mathcal\{H\}\_\{k\}is prime and greater thankk, so residue class0modp0\\bmod pis empty for every primep≤kp\\leq k; forp\>kp\>k, akk\-element set cannot meet allppresidue classes\. Henceℋk\\mathcal\{H\}\_\{k\}is admissible by construction\. Segmented prime enumeration gives

D11655069≤214099720,D644589002≤14541349288\.D\_\{11655069\}\\leq 214099720,\\qquad D\_\{644589002\}\\leq 14541349288\.The code additionally evaluates explicit prime\-counting and prime\-location bounds as an independent large\-kkconsistency check\.

Combining these tuple diameters with the certified variational thresholdsMk\>4​mM\_\{k\}\>4mgives

and hence the numerical bounded\-gap statements reported in the Results\. The tuple\-construction code, certificate files, and independent verifier are supplied as Supplementary Software/Data\.

### 5\.31Standalone verification and certificate schema

The large\-kkcertificate is a self\-contained JSON record containing the exact rational trial parametercc, Fourier discretization parameters, intermediate summaries and the claimed lower bound\. Stored intermediate numerical enclosures are not trusted\. The verifier reconstructs the analytic Fourier\-tail constant, regenerates the far\-field block partition, recomputes wide\-ball suprema, moment bounds, Fourier nodes, trapezoidal sums, aliasing and truncation errors, and independently assemblesRlowR\_\{\\rm low\}in ball arithmetic\.

The certificate proves only the Rayleigh quotient of the fixed admissible trial function encoded by the exact rationalcc\. It does not prove optimality ofcc, re\-prove the external sieve criterion or general upper bounds onMkM\_\{k\}, or certify admissible\-tuple diameters\. Monte–Carlo comparisons and external small\-kkvalues are falsification checks rather than proof components\.

## 6Asymptotics of the single\-channel rational family

We now explain analytically why the single\-channel rational trial functions used in the large\-kkcomputations remain within a bounded distance of the logarithmic scalelog⁡k\\log k\.

Fork≥2k\\geq 2, write

and consider the symmetric trial function

Fk,c\(t1,…,tk\)=𝟏\{ti≥0,∑i=1kti≤1\}∏i=1k1c\+n​ti,c\>0\.F\_\{k,c\}\(t\_\{1\},\\ldots,t\_\{k\}\)=\\mathbf\{1\}\_\{\\\{t\_\{i\}\\geq 0,\\ \\sum\_\{i=1\}^\{k\}t\_\{i\}\\leq 1\\\}\}\\prod\_\{i=1\}^\{k\}\\frac\{1\}\{c\+nt\_\{i\}\},\\qquad c\>0\.\(60\)Its overall scalar normalization is immaterial\.

The exact probabilistic reduction established in Section[5\.20](https://arxiv.org/html/2609.30296#S5.SS20)gives

R⁡\(Fk,c\)=k​Nk,cDk,c,R\(F\_\{k,c\}\)=k\\frac\{N\_\{k,c\}\}\{D\_\{k,c\}\},\(61\)where

w⁡\(t\)=1\(c\+n​t\)2,m0=∫01w⁡\(t\)​𝑑t=1c⁡\(c\+n\),w\(t\)=\\frac\{1\}\{\(c\+nt\)^\{2\}\},\\qquad m\_\{0\}=\\int\_\{0\}^\{1\}w\(t\)\\,dt=\\frac\{1\}\{c\(c\+n\)\},and, ifT1,…,TnT\_\{1\},\\ldots,T\_\{n\}are independent with density

w⁡\(t\)m0=c⁡\(c\+n\)\(c\+n​t\)2,0≤t≤1,\\frac\{w\(t\)\}\{m\_\{0\}\}=\\frac\{c\(c\+n\)\}\{\(c\+nt\)^\{2\}\},\\qquad 0\\leq t\\leq 1,and

Sn:=T1\+⋯\+Tn,S\_\{n\}:=T\_\{1\}\+\\cdots\+T\_\{n\},then

Nk,c\\displaystyle N\_\{k,c\}=𝔼⁡\[G​\(1−Sn\)2;Sn<1\],\\displaystyle=\\mathbb\{E\}\\\!\\left\[G\(1\-S\_\{n\}\)^\{2\}\\,;\\,S\_\{n\}<1\\right\],\(62\)Dk,c\\displaystyle D\_\{k,c\}=𝔼⁡\[H⁡\(1−Sn\);Sn<1\],\\displaystyle=\\mathbb\{E\}\\\!\\left\[H\(1\-S\_\{n\}\)\\,;\\,S\_\{n\}<1\\right\],\(63\)with

G⁡\(ρ\)\\displaystyle G\(\\rho\)=1n​log⁡\(1\+n​ρc\),\\displaystyle=\\frac\{1\}\{n\}\\log\\left\(1\+\\frac\{n\\rho\}\{c\}\\right\),\(64\)H⁡\(ρ\)\\displaystyle H\(\\rho\)=1n​\(1c−1c\+n​ρ\)\.\\displaystyle=\\frac\{1\}\{n\}\\left\(\\frac\{1\}\{c\}\-\\frac\{1\}\{c\+n\\rho\}\\right\)\.\(65\)
The relevant scale forccis

1ck=log\(k−1\)\+a,a∈ℝfixed\.\\frac\{1\}\{c\_\{k\}\}=\\log\(k\-1\)\+a,\\qquad a\\in\\mathbb\{R\}\\quad\\text\{fixed\}\.\(66\)The numerical optimization of the rational family independently exhibits thisck≍1/log⁡kc\_\{k\}\\asymp 1/\\log kscaling\. The result below identifies the limitingO⁡\(1\)O\(1\)defect fromlog⁡k\\log kfor every fixed shiftaa\.

### 6\.1The limiting stable law

We first isolate the random variable governing the boundary of the simplex\. Define

Xn:=1−Snc\.X\_\{n\}:=\\frac\{1\-S\_\{n\}\}\{c\}\.\(67\)Thus

Sn<1⟺Xn\>0\.S\_\{n\}<1\\quad\\Longleftrightarrow\\quad X\_\{n\}\>0\.
It is convenient to rescale a single summand by writing

A direct change of variables gives the density

qn​\(v\)=c\+n\(1\+n​v\)2,0≤v≤1c\.q\_\{n\}\(v\)=\\frac\{c\+n\}\{\(1\+nv\)^\{2\}\},\\qquad 0\\leq v\\leq\\frac\{1\}\{c\}\.\(68\)Hence

Xn=1c−∑j=1nVj,X\_\{n\}=\\frac\{1\}\{c\}\-\\sum\_\{j=1\}^\{n\}V\_\{j\},\(69\)whereV1,…,VnV\_\{1\},\\ldots,V\_\{n\}are independent with density \([68](https://arxiv.org/html/2609.30296#S6.E68)\)\.

We use the following normalization of the totally right\-skewed11\-stable law\.

###### Definition 1\(The limiting11\-stable variable\)\.

LetΛ\\Lambdabe the infinitely divisible random variable with characteristic function

𝔼ei​s​Λ=exp\{∫0∞\(ei​s​x−1−isx𝟏\{x≤1\}\)d​xx2\}\.\\mathbb\{E\}e^\{is\\Lambda\}=\\exp\\left\\\{\\int\_\{0\}^\{\\infty\}\\left\(e^\{isx\}\-1\-isx\\mathbf\{1\}\_\{\\\{x\\leq 1\\\}\}\\right\)\\frac\{dx\}\{x^\{2\}\}\\right\\\}\.\(70\)This is a spectrally positive stable law of index11, in the Lévy–Khintchine normalization used below\. Its law has a continuous, everywhere positive density onℝ\\mathbb\{R\}\.

BecauseΛ\\Lambdahas only positive jumps, its Laplace transform is finite in the damping direction\. The following elementary computation is used repeatedly: it controls the left tail ofΛ\\Lambda, supplies the integrability needed in Lemma[11](https://arxiv.org/html/2609.30296#Thmtheorem11), and provides a one\-dimensional quadrature route to the constantsPaP\_\{a\},LaL\_\{a\},QaQ\_\{a\}introduced below\.

###### Lemma 2\(Laplace transform and left tail ofΛ\\Lambda\)\.

For everyθ\>0\\theta\>0,

𝔼​e−θ​Λ=exp⁡\{θ​log⁡θ−θ\+γE​θ\},\\mathbb\{E\}\\,e^\{\-\\theta\\Lambda\}=\\exp\\bigl\\\{\\theta\\log\\theta\-\\theta\+\\gamma\_\{\\mathrm\{E\}\}\\theta\\bigr\\\},\(71\)whereγE\\gamma\_\{\\mathrm\{E\}\}is the Euler–Mascheroni constant\. To avoid collision we useγE\\gamma\_\{\\mathrm\{E\}\}instead ofγh\\gamma\_\{\\mathrm\{h\}\}from Section[5\.25](https://arxiv.org/html/2609.30296#S5.SS25)\. Consequently, for ally≥0y\\geq 0,

ℙ⁡\(Λ≤−y\)≤exp⁡\{−ey−γE\},\\mathbb\{P\}\(\\Lambda\\leq\-y\)\\leq\\exp\\left\\\{\-e^\{\\,y\-\\gamma\_\{\\mathrm\{E\}\}\}\\right\\\},\(72\)and in particular

κ:=𝔼⁡\[\(−Λ\)\+\]<∞\.\\kappa:=\\mathbb\{E\}\\bigl\[\(\-\\Lambda\)^\{\+\}\\bigr\]<\\infty\.\(73\)

###### Proof\.

Analytic continuation of \([70](https://arxiv.org/html/2609.30296#S6.E70)\) tos=i​θs=i\\thetagives

J\(θ\):=log𝔼e−θ​Λ=∫0∞\(e−θ​x−1\+θx𝟏\{x≤1\}\)d​xx2,J\(\\theta\):=\\log\\mathbb\{E\}\\,e^\{\-\\theta\\Lambda\}=\\int\_\{0\}^\{\\infty\}\\left\(e^\{\-\\theta x\}\-1\+\\theta x\\mathbf\{1\}\_\{\\\{x\\leq 1\\\}\}\\right\)\\frac\{dx\}\{x^\{2\}\},the integral being absolutely convergent forθ\>0\\theta\>0\. Differentiating under the integral sign,

J′​\(θ\)=∫011−e−θ​xx​𝑑x−∫1∞e−θ​xx​𝑑x=\(log⁡θ\+γE\+E1​\(θ\)\)−E1​\(θ\)=log⁡θ\+γE,J^\{\\prime\}\(\\theta\)=\\int\_\{0\}^\{1\}\\frac\{1\-e^\{\-\\theta x\}\}\{x\}\\,dx\-\\int\_\{1\}^\{\\infty\}\\frac\{e^\{\-\\theta x\}\}\{x\}\\,dx=\\bigl\(\\log\\theta\+\\gamma\_\{\\mathrm\{E\}\}\+E\_\{1\}\(\\theta\)\\bigr\)\-E\_\{1\}\(\\theta\)=\\log\\theta\+\\gamma\_\{\\mathrm\{E\}\},whereE1E\_\{1\}is the exponential integral and we used the standard identity∫0z\(1−e−t\)​t−1​𝑑t=γE\+log⁡z\+E1​\(z\)\\int\_\{0\}^\{z\}\(1\-e^\{\-t\}\)t^\{\-1\}dt=\\gamma\_\{\\mathrm\{E\}\}\+\\log z\+E\_\{1\}\(z\)\. SinceJ⁡\(0\+\)=0J\(0^\{\+\}\)=0, integration gives \([71](https://arxiv.org/html/2609.30296#S6.E71)\)\.

For \([72](https://arxiv.org/html/2609.30296#S6.E72)\), Markov’s inequality applied toe−θ​Λe^\{\-\\theta\\Lambda\}yields, for everyθ\>0\\theta\>0,

ℙ⁡\(Λ≤−y\)≤e−θ​y​𝔼​e−θ​Λ=exp⁡\{θ⁡\(log⁡θ−1\+γE−y\)\}\.\\mathbb\{P\}\(\\Lambda\\leq\-y\)\\leq e^\{\-\\theta y\}\\,\\mathbb\{E\}\\,e^\{\-\\theta\\Lambda\}=\\exp\\left\\\{\\theta\(\\log\\theta\-1\+\\gamma\_\{\\mathrm\{E\}\}\-y\)\\right\\\}\.The exponent is minimized atθ=ey−γE\\theta=e^\{\\,y\-\\gamma\_\{\\mathrm\{E\}\}\}, where it equals−θ\-\\theta, giving \([72](https://arxiv.org/html/2609.30296#S6.E72)\)\. Finally𝔼⁡\[\(−Λ\)\+\]=∫0∞ℙ⁡\(Λ≤−y\)​𝑑y≤∫0∞exp⁡\{−ey−γE\}​𝑑y<∞\\mathbb\{E\}\[\(\-\\Lambda\)^\{\+\}\]=\\int\_\{0\}^\{\\infty\}\\mathbb\{P\}\(\\Lambda\\leq\-y\)\\,dy\\leq\\int\_\{0\}^\{\\infty\}\\exp\\\{\-e^\{y\-\\gamma\_\{\\mathrm\{E\}\}\}\\\}\\,dy<\\infty\. ∎

###### Lemma 3\(Stable limit\)\.

Suppose

1c=log⁡n\+a\\frac\{1\}\{c\}=\\log n\+awith fixeda∈ℝa\\in\\mathbb\{R\}\. Then

Xn→𝑑Xa:=\(a\+1\)−Λ\.X\_\{n\}\\ \\xrightarrow\{\\ d\\ \}\\ X\_\{a\}:=\(a\+1\)\-\\Lambda\.\(74\)

###### Proof\.

The characteristic function ofXnX\_\{n\}is

ϕn​\(s\)=ei​s/c​\(𝔼​e−i​s​V\)n\.\\phi\_\{n\}\(s\)=e^\{is/c\}\\left\(\\mathbb\{E\}e^\{\-isV\}\\right\)^\{n\}\.Introduce the finite Lévy measure

μn\(dx\):=nqn\(x\)dx=n⁡\(c\+n\)\(1\+n​x\)2𝟏\{0≤x≤1/c\}dx,\\mu\_\{n\}\(dx\):=nq\_\{n\}\(x\)\\,dx=\\frac\{n\(c\+n\)\}\{\(1\+nx\)^\{2\}\}\\mathbf\{1\}\_\{\\\{0\\leq x\\leq 1/c\\\}\}\\,dx,of total massμn​\(\[0,1/c\]\)=n\\mu\_\{n\}\(\[0,1/c\]\)=n\. For every fixedx\>0x\>0,

n⁡\(c\+n\)\(1\+n​x\)2⟶1x2\.\\frac\{n\(c\+n\)\}\{\(1\+nx\)^\{2\}\}\\longrightarrow\\frac\{1\}\{x^\{2\}\}\.
For fixedssand sufficiently largenn, choose the logarithm of1\+zn​\(s\)/n1\+z\_\{n\}\(s\)/nnear11and use the corresponding characteristic exponent forϕn​\(s\)\\phi\_\{n\}\(s\)\. With this convention, we write

log⁡ϕn​\(s\)\\displaystyle\\log\\phi\_\{n\}\(s\)=i​sc\+n​log⁡\(1\+zn​\(s\)n\),zn​\(s\):=∫\(e−i​s​x−1\)​μn​\(𝑑x\)\.\\displaystyle=\\frac\{is\}\{c\}\+n\\log\\left\(1\+\\frac\{z\_\{n\}\(s\)\}\{n\}\\right\),\\qquad z\_\{n\}\(s\):=\\int\(e^\{\-isx\}\-1\)\\,\\mu\_\{n\}\(dx\)\.\(75\)Fixss\. Using\|e−i​s​x−1\|≤min⁡\(2,\|s\|​x\)\|e^\{\-isx\}\-1\|\\leq\\min\(2,\|s\|x\)and splitting the integral atx=1/\|s\|x=1/\|s\|,

\|zn​\(s\)\|≤\|s\|​∫01/\|s\|x​μn​\(𝑑x\)\+2​μn​\(\(1/\|s\|,1/c\]\)=Os​\(log⁡n\),\|z\_\{n\}\(s\)\|\\ \\leq\\ \|s\|\\int\_\{0\}^\{1/\|s\|\}x\\,\\mu\_\{n\}\(dx\)\+2\\,\\mu\_\{n\}\\bigl\(\(1/\|s\|,1/c\]\\bigr\)\\ =\\ O\_\{s\}\(\\log n\),\(76\)by the elementary evaluation \([78](https://arxiv.org/html/2609.30296#S6.E78)\) below\. Thuszn​\(s\)/n→0z\_\{n\}\(s\)/n\\to 0, and the principal branch of the logarithm satisfies

n​log⁡\(1\+znn\)=zn\+O⁡\(\|zn\|2n\)=zn\+Os​\(log2⁡nn\)=zn\+o⁡\(1\)\.n\\log\\left\(1\+\\frac\{z\_\{n\}\}\{n\}\\right\)=z\_\{n\}\+O\\\!\\left\(\\frac\{\|z\_\{n\}\|^\{2\}\}\{n\}\\right\)=z\_\{n\}\+O\_\{s\}\\\!\\left\(\\frac\{\\log^\{2\}n\}\{n\}\\right\)=z\_\{n\}\+o\(1\)\.\(The hypothesis needed here iszn=o⁡\(n\)z\_\{n\}=o\(\\sqrt\{n\}\), not boundedness ofznz\_\{n\}; \([76](https://arxiv.org/html/2609.30296#S6.E76)\) supplies it\.\) Hence

log⁡ϕn​\(s\)\\displaystyle\\log\\phi\_\{n\}\(s\)=i​sc\+∫01/c\(e−i​s​x−1\)​μn​\(𝑑x\)\+o⁡\(1\)\\displaystyle=\\frac\{is\}\{c\}\+\\int\_\{0\}^\{1/c\}\(e^\{\-isx\}\-1\)\\,\\mu\_\{n\}\(dx\)\+o\(1\)=i​s​\[1c−∫01x​μn​\(𝑑x\)\]\\displaystyle=is\\left\[\\frac\{1\}\{c\}\-\\int\_\{0\}^\{1\}x\\,\\mu\_\{n\}\(dx\)\\right\]\+∫01/c\(e−i​s​x−1\+isx𝟏\{x≤1\}\)μn\(dx\)\+o\(1\)\.\\displaystyle\\quad\+\\int\_\{0\}^\{1/c\}\\left\(e^\{\-isx\}\-1\+isx\\mathbf\{1\}\_\{\\\{x\\leq 1\\\}\}\\right\)\\mu\_\{n\}\(dx\)\+o\(1\)\.\(77\)
The compensated integral converges, by dominated convergence on compact subintervals and elementary control at00and∞\\infty, to

∫0∞\(e−i​s​x−1\+isx𝟏\{x≤1\}\)d​xx2\.\\int\_\{0\}^\{\\infty\}\\left\(e^\{\-isx\}\-1\+isx\\mathbf\{1\}\_\{\\\{x\\leq 1\\\}\}\\right\)\\frac\{dx\}\{x^\{2\}\}\.
It remains to identify the drift\. Substitutingu=n​xu=nx,

∫01x​μn​\(𝑑x\)\\displaystyle\\int\_\{0\}^\{1\}x\\,\\mu\_\{n\}\(dx\)=\(1\+cn\)​∫0nu​d​u\(1\+u\)2=\(1\+cn\)​\[log⁡\(1\+n\)−1\+11\+n\]\.\\displaystyle=\\left\(1\+\\frac\{c\}\{n\}\\right\)\\int\_\{0\}^\{n\}\\frac\{u\\,du\}\{\(1\+u\)^\{2\}\}=\\left\(1\+\\frac\{c\}\{n\}\\right\)\\left\[\\log\(1\+n\)\-1\+\\frac\{1\}\{1\+n\}\\right\]\.\(78\)Sincec=1/\(log⁡n\+a\)c=1/\(\\log n\+a\),

1c−∫01x​μn​\(𝑑x\)⟶a\+1\.\\frac\{1\}\{c\}\-\\int\_\{0\}^\{1\}x\\,\\mu\_\{n\}\(dx\)\\longrightarrow a\+1\.Consequently

logϕn\(s\)⟶is\(a\+1\)\+∫0∞\(e−i​s​x−1\+isx𝟏\{x≤1\}\)d​xx2\.\\displaystyle\\log\\phi\_\{n\}\(s\)\\longrightarrow is\(a\+1\)\+\\int\_\{0\}^\{\\infty\}\\left\(e^\{\-isx\}\-1\+isx\\mathbf\{1\}\_\{\\\{x\\leq 1\\\}\}\\right\)\\frac\{dx\}\{x^\{2\}\}\.\(79\)By Definition[1](https://arxiv.org/html/2609.30296#Thmtheorem1), the limiting characteristic function is that of\(a\+1\)−Λ\(a\+1\)\-\\Lambda\. Lévy’s continuity theorem gives \([74](https://arxiv.org/html/2609.30296#S6.E74)\)\. ∎

### 6\.2Uniform density and logarithmic moment control

Weak convergence alone is insufficient for the logarithmic observables appearing in the Rayleigh quotient, becauselog⁡x\\log xis singular atx=0x=0\. We therefore record two uniform estimates\.

###### Lemma 4\(Uniform Fourier integrability and bounded densities\)\.

For fixedaa, withc−1=log⁡n\+ac^\{\-1\}=\\log n\+a, there exists a constantCa<∞C\_\{a\}<\\inftysuch that, for all sufficiently largenn,

∫ℝ\|ϕn​\(s\)\|​𝑑s≤Ca\.\\int\_\{\\mathbb\{R\}\}\|\\phi\_\{n\}\(s\)\|\\,ds\\leq C\_\{a\}\.\(80\)ConsequentlyXnX\_\{n\}possesses a densityfnf\_\{n\}satisfying

supn≥n0‖fn‖∞<∞\.\\sup\_\{n\\geq n\_\{0\}\}\\\|f\_\{n\}\\\|\_\{\\infty\}<\\infty\.\(81\)In particular, uniformly for0<δ≤10<\\delta\\leq 1,

ℙ\(0<Xn≤δ\)≪aδ\.\\mathbb\{P\}\(0<X\_\{n\}\\leq\\delta\)\\ll\_\{a\}\\delta\.\(82\)

###### Proof\.

Since translation does not affect the modulus of the characteristic function,

\|ϕn​\(s\)\|=\|ψn​\(s\)\|n,ψn​\(s\):=𝔼​e−i​s​V,\|\\phi\_\{n\}\(s\)\|=\|\\psi\_\{n\}\(s\)\|^\{n\},\\qquad\\psi\_\{n\}\(s\):=\\mathbb\{E\}e^\{\-isV\},and\|ψn\|\|\\psi\_\{n\}\|is even inss, so we may assumes\>0s\>0\.

*An explicit two\-interval bound\.*Forβ\>0\\beta\>0let

I1​\(β\):=\[0,π4​β\],I2​\(β\):=\[3​π4​β,5​π4​β\]\.I\_\{1\}\(\\beta\):=\\Bigl\[0,\\tfrac\{\\pi\}\{4\}\\beta\\Bigr\],\\qquad I\_\{2\}\(\\beta\):=\\Bigl\[\\tfrac\{3\\pi\}\{4\}\\beta,\\tfrac\{5\\pi\}\{4\}\\beta\\Bigr\]\.These are disjoint, and forv∈I1​\(β\)v\\in I\_\{1\}\(\\beta\),v′∈I2​\(β\)v^\{\\prime\}\\in I\_\{2\}\(\\beta\)we havev′−v∈\[π2​β,5​π4​β\]v^\{\\prime\}\-v\\in\[\\tfrac\{\\pi\}\{2\}\\beta,\\tfrac\{5\\pi\}\{4\}\\beta\]\. Takingβ=1/s\\beta=1/stherefore givess⁡\(v′−v\)∈\[π/2,5​π/4\]s\(v^\{\\prime\}\-v\)\\in\[\\pi/2,5\\pi/4\], on whichcos≤0\\cos\\leq 0and hence1−cos⁡\(s⁡\(v−v′\)\)≥11\-\\cos\\bigl\(s\(v\-v^\{\\prime\}\)\\bigr\)\\geq 1\. WritingV′V^\{\\prime\}for an independent copy ofVVand discarding all other configurations,

1−\|ψn​\(s\)\|2=𝔼⁡\[1−cos⁡\(s⁡\(V−V′\)\)\]≥2​ℙ​\(V∈I1​\(1/s\)\)​ℙ​\(V∈I2​\(1/s\)\)\.1\-\|\\psi\_\{n\}\(s\)\|^\{2\}=\\mathbb\{E\}\\bigl\[1\-\\cos\\bigl\(s\(V\-V^\{\\prime\}\)\\bigr\)\\bigr\]\\ \\geq\\ 2\\,\\mathbb\{P\}\\bigl\(V\\in I\_\{1\}\(1/s\)\\bigr\)\\,\\mathbb\{P\}\\bigl\(V\\in I\_\{2\}\(1/s\)\\bigr\)\.\(83\)Both intervals lie in\[0,1/c\]\[0,1/c\]once5​π/\(4​s\)≤1/c5\\pi/\(4s\)\\leq 1/c, which holds for alls≥1s\\geq 1andnnlarge, since1/c=log⁡n\+a→∞1/c=\\log n\+a\\to\\infty\.

From \([68](https://arxiv.org/html/2609.30296#S6.E68)\), for0≤α<β≤1/c0\\leq\\alpha<\\beta\\leq 1/c,

ℙ⁡\(V∈\[α,β\]\)=c\+nn​\[11\+n​α−11\+n​β\]\.\\mathbb\{P\}\(V\\in\[\\alpha,\\beta\]\)=\\frac\{c\+n\}\{n\}\\left\[\\frac\{1\}\{1\+n\\alpha\}\-\\frac\{1\}\{1\+n\\beta\}\\right\]\.\(84\)Putb:=π​n/\(4​s\)b:=\\pi n/\(4s\)\. Then \([84](https://arxiv.org/html/2609.30296#S6.E84)\) gives

ℙ⁡\(V∈I1​\(1/s\)\)=c\+nn⋅b1\+b,ℙ⁡\(V∈I2​\(1/s\)\)=c\+nn⋅2​b\(1\+3​b\)​\(1\+5​b\)\.\\mathbb\{P\}\\bigl\(V\\in I\_\{1\}\(1/s\)\\bigr\)=\\frac\{c\+n\}\{n\}\\cdot\\frac\{b\}\{1\+b\},\\qquad\\mathbb\{P\}\\bigl\(V\\in I\_\{2\}\(1/s\)\\bigr\)=\\frac\{c\+n\}\{n\}\\cdot\\frac\{2b\}\{\(1\+3b\)\(1\+5b\)\}\.
*Low frequencies1≤s≤n1\\leq s\\leq n\.*Hereb≥π/4b\\geq\\pi/4, sob/\(1\+b\)≥\(π/4\)/\(1\+π/4\)\>13b/\(1\+b\)\\geq\(\\pi/4\)/\(1\+\\pi/4\)\>\\tfrac\{1\}\{3\}and2​b/\{\(1\+3​b\)​\(1\+5​b\)\}≥2​b/\{\(3\+4/π\)​b⋅\(5\+4/π\)​b\}≫1/b≫s/n2b/\\\{\(1\+3b\)\(1\+5b\)\\\}\\geq 2b/\\\{\(3\+4/\\pi\)b\\cdot\(5\+4/\\pi\)b\\\}\\gg 1/b\\gg s/n\. With \([83](https://arxiv.org/html/2609.30296#S6.E83)\),

1−\|ψn​\(s\)\|2≫sn,1≤s≤n,1\-\|\\psi\_\{n\}\(s\)\|^\{2\}\\gg\\frac\{s\}\{n\},\\qquad 1\\leq s\\leq n,\(85\)and therefore, for someκ\>0\\kappa\>0independent ofnn,

\|ϕn\(s\)\|=\(\|ψn\(s\)\|2\)n/2≤\(1−κ​sn\)n/2≤e−κs/2\(1≤s≤n\)\.\|\\phi\_\{n\}\(s\)\|=\\bigl\(\|\\psi\_\{n\}\(s\)\|^\{2\}\\bigr\)^\{n/2\}\\leq\\left\(1\-\\frac\{\\kappa s\}\{n\}\\right\)^\{n/2\}\\leq e^\{\-\\kappa s/2\}\\qquad\(1\\leq s\\leq n\)\.\(86\)
*Intermediate frequenciesn≤s≤8​\(c\+n\)n\\leq s\\leq 8\(c\+n\)\.*Nowb≤π/4b\\leq\\pi/4, sob/\(1\+b\)≫bb/\(1\+b\)\\gg band2​b/\{\(1\+3​b\)​\(1\+5​b\)\}≫b2b/\\\{\(1\+3b\)\(1\+5b\)\\\}\\gg b, whence1−\|ψn​\(s\)\|2≫b2≫n2/s21\-\|\\psi\_\{n\}\(s\)\|^\{2\}\\gg b^\{2\}\\gg n^\{2\}/s^\{2\}\. Sinces=O⁡\(n\)s=O\(n\)throughout this range,

\|ϕn​\(s\)\|≤e−κ′​n\.\|\\phi\_\{n\}\(s\)\|\\leq e^\{\-\\kappa^\{\\prime\}n\}\.\(87\)
*High frequencies\.*Integration by parts in

ψn​\(s\)=∫01/ce−i​s​v​c\+n\(1\+n​v\)2​𝑑v\\psi\_\{n\}\(s\)=\\int\_\{0\}^\{1/c\}e^\{\-isv\}\\frac\{c\+n\}\{\(1\+nv\)^\{2\}\}\\,dvgives

\|ψn​\(s\)\|≤2​\(c\+n\)s\.\|\\psi\_\{n\}\(s\)\|\\leq\\frac\{2\(c\+n\)\}\{s\}\.\(88\)Thus, fors≥8​\(c\+n\)s\\geq 8\(c\+n\),

\|ϕn​\(s\)\|≤\(14\)n​\(8​\(c\+n\)s\)n≤\(14\)n,\|\\phi\_\{n\}\(s\)\|\\leq\\left\(\\frac\{1\}\{4\}\\right\)^\{n\}\\left\(\\frac\{8\(c\+n\)\}\{s\}\\right\)^\{n\}\\leq\\left\(\\frac\{1\}\{4\}\\right\)^\{n\},\(89\)and more precisely\|ϕn​\(s\)\|≤\(2​\(c\+n\)/s\)n\|\\phi\_\{n\}\(s\)\|\\leq\(2\(c\+n\)/s\)^\{n\}, whose integral overs≥8​\(c\+n\)s\\geq 8\(c\+n\)isO⁡\(\(c\+n\)​4−n\)O\\bigl\(\(c\+n\)4^\{\-n\}\\bigr\), hence uniformly bounded\.

Combining \([86](https://arxiv.org/html/2609.30296#S6.E86)\), \([87](https://arxiv.org/html/2609.30296#S6.E87)\), and \([89](https://arxiv.org/html/2609.30296#S6.E89)\), together with the trivial bound\|ϕn​\(s\)\|≤1\|\\phi\_\{n\}\(s\)\|\\leq 1for\|s\|≤1\|s\|\\leq 1, proves \([80](https://arxiv.org/html/2609.30296#S6.E80)\)\.

Fourier inversion then gives

‖fn‖∞≤12​π​‖ϕn‖L1,\\\|f\_\{n\}\\\|\_\{\\infty\}\\leq\\frac\{1\}\{2\\pi\}\\\|\\phi\_\{n\}\\\|\_\{L^\{1\}\},which proves \([81](https://arxiv.org/html/2609.30296#S6.E81)\); \([82](https://arxiv.org/html/2609.30296#S6.E82)\) follows immediately\. ∎

###### Lemma 5\(Right\-tail bound and logarithmic uniform integrability\)\.

For fixedaa, the random variables

logXn1\{Xn\>0\}and\(logXn\)21\{Xn\>0\}\\log X\_\{n\}\\,\\mathbf\{1\}\_\{\\\{X\_\{n\}\>0\\\}\}\\quad\\text\{and\}\\quad\(\\log X\_\{n\}\)^\{2\}\\,\\mathbf\{1\}\_\{\\\{X\_\{n\}\>0\\\}\}are uniformly integrable\. Consequently, writing

Pn\\displaystyle P\_\{n\}:=ℙ⁡\(Xn\>0\),\\displaystyle:=\\mathbb\{P\}\(X\_\{n\}\>0\),\(90\)Ln\\displaystyle L\_\{n\}:=𝔼⁡\[log⁡Xn;Xn\>0\],\\displaystyle:=\\mathbb\{E\}\[\\log X\_\{n\}\\,;\\,X\_\{n\}\>0\],\(91\)Qn\\displaystyle Q\_\{n\}:=𝔼⁡\[\(log⁡Xn\)2;Xn\>0\],\\displaystyle:=\\mathbb\{E\}\[\(\\log X\_\{n\}\)^\{2\}\\,;\\,X\_\{n\}\>0\],\(92\)we have

Pn\\displaystyle P\_\{n\}⟶Pa:=ℙ⁡\(Xa\>0\)\>0,\\displaystyle\\longrightarrow P\_\{a\}:=\\mathbb\{P\}\(X\_\{a\}\>0\)\>0,\(93\)Ln\\displaystyle L\_\{n\}⟶La:=𝔼⁡\[log⁡Xa;Xa\>0\],\\displaystyle\\longrightarrow L\_\{a\}:=\\mathbb\{E\}\[\\log X\_\{a\}\\,;\\,X\_\{a\}\>0\],\(94\)Qn\\displaystyle Q\_\{n\}⟶Qa:=𝔼⁡\[\(log⁡Xa\)2;Xa\>0\]\.\\displaystyle\\longrightarrow Q\_\{a\}:=\\mathbb\{E\}\[\(\\log X\_\{a\}\)^\{2\}\\,;\\,X\_\{a\}\>0\]\.\(95\)

###### Proof\.

The negative logarithmic tail is controlled by Lemma[4](https://arxiv.org/html/2609.30296#Thmtheorem4)\. Fory≥0y\\geq 0,

ℙ⁡\(0<Xn<e−y\)≪e−y\.\\mathbb\{P\}\(0<X\_\{n\}<e^\{\-y\}\)\\ll e^\{\-y\}\.Hence the negative parts oflog⁡Xn\\log X\_\{n\}have uniformly bounded moments of every fixed order\.

For the positive tail, note from \([69](https://arxiv.org/html/2609.30296#S6.E69)\) that

Xn\>x⟺∑j=1nVj<1c−x\.X\_\{n\}\>x\\quad\\Longleftrightarrow\\quad\\sum\_\{j=1\}^\{n\}V\_\{j\}<\\frac\{1\}\{c\}\-x\.Forλ\>0\\lambda\>0, Chernoff’s inequality gives

ℙ⁡\(Xn\>x\)≤exp⁡\(λ⁡\(1c−x\)\)​\(𝔼​e−λ​V\)n\.\\mathbb\{P\}\(X\_\{n\}\>x\)\\leq\\exp\\\!\\left\(\\lambda\\left\(\\frac\{1\}\{c\}\-x\\right\)\\right\)\\left\(\\mathbb\{E\}e^\{\-\\lambda V\}\\right\)^\{n\}\.\(96\)
We claim that there is a constantCaC\_\{a\}, independent ofnnandλ\\lambda, such that

n⁡\(1−𝔼​e−λ​V\)≥λ⁡\(log⁡nλ−Ca\),1≤λ≤n\.n\\bigl\(1\-\\mathbb\{E\}e^\{\-\\lambda V\}\\bigr\)\\ \\geq\\ \\lambda\\left\(\\log\\frac\{n\}\{\\lambda\}\-C\_\{a\}\\right\),\\qquad 1\\leq\\lambda\\leq n\.\(97\)Substitutingu=n​vu=nvin \([68](https://arxiv.org/html/2609.30296#S6.E68)\) and writingβ:=λ/n∈\(0,1\]\\beta:=\\lambda/n\\in\(0,1\],

n⁡\(1−𝔼​e−λ​V\)\\displaystyle n\\bigl\(1\-\\mathbb\{E\}e^\{\-\\lambda V\}\\bigr\)=\(c\+n\)​∫0n/c\(1−e−β​u\)​d​u\(1\+u\)2\.\\displaystyle=\(c\+n\)\\int\_\{0\}^\{n/c\}\\bigl\(1\-e^\{\-\\beta u\}\\bigr\)\\frac\{du\}\{\(1\+u\)^\{2\}\}\.\(98\)Rather than splitting the range, evaluate the integral exactly\. Integration by parts gives, for the untruncated integral,

∫0∞\(1−e−β​u\)​d​u\(1\+u\)2=β​∫0∞e−β​u1\+u​𝑑u=β​eβ​E1​\(β\)\.\\int\_\{0\}^\{\\infty\}\\bigl\(1\-e^\{\-\\beta u\}\\bigr\)\\frac\{du\}\{\(1\+u\)^\{2\}\}=\\beta\\int\_\{0\}^\{\\infty\}\\frac\{e^\{\-\\beta u\}\}\{1\+u\}\\,du=\\beta\\,e^\{\\beta\}E\_\{1\}\(\\beta\)\.To make the estimate uniform over the entire range0<β≤10<\\beta\\leq 1, note thateβ​E1​\(β\)−log⁡\(1/β\)e^\{\\beta\}E\_\{1\}\(\\beta\)\-\\log\(1/\\beta\)is continuous on\(0,1\]\(0,1\]and tends to−γE\-\\gamma\_\{\\mathrm\{E\}\}asβ↓0\\beta\\downarrow 0\. Consequently,

sup0<β≤1\|eβ​E1​\(β\)−log⁡\(1/β\)\|<∞\.\\sup\_\{0<\\beta\\leq 1\}\\left\|e^\{\\beta\}E\_\{1\}\(\\beta\)\-\\log\(1/\\beta\)\\right\|<\\infty\.The discarded tail satisfies∫n/c∞\(1\+u\)−2​𝑑u≤c/n\\int\_\{n/c\}^\{\\infty\}\(1\+u\)^\{\-2\}du\\leq c/n, contributing at most\(c\+n\)​c/n=O⁡\(1\)≤O⁡\(λ\)\(c\+n\)c/n=O\(1\)\\leq O\(\\lambda\)sinceλ≥1\\lambda\\geq 1\. Multiplying by\(c\+n\)=n⁡\(1\+c/n\)\(c\+n\)=n\(1\+c/n\)and usingβ=λ/n\\beta=\\lambda/n, with\(c/n\)​log⁡\(n/λ\)≤\(c/n\)​log⁡n=O⁡\(n−1\)\(c/n\)\\log\(n/\\lambda\)\\leq\(c/n\)\\log n=O\(n^\{\-1\}\), yields

n⁡\(1−𝔼​e−λ​V\)=λ⁡\(log⁡nλ−γE\+O⁡\(1\)\),n\\bigl\(1\-\\mathbb\{E\}e^\{\-\\lambda V\}\\bigr\)=\\lambda\\left\(\\log\\frac\{n\}\{\\lambda\}\-\\gamma\_\{\\mathrm\{E\}\}\+O\(1\)\\right\),which is \([97](https://arxiv.org/html/2609.30296#S6.E97)\)\. \(The coefficient11in front oflog⁡\(n/λ\)\\log\(n/\\lambda\)is essential: it is what cancels1/c=log⁡n\+a1/c=\\log n\+ain \([96](https://arxiv.org/html/2609.30296#S6.E96)\)\. A cruder split using1−e−y≥y/21\-e^\{\-y\}\\geq y/2produces the coefficient1/21/2and the argument fails\.\)

Sincelog⁡z≤z−1\\log z\\leq z\-1forz\>0z\>0,

n​log⁡𝔼​e−λ​V≤−n⁡\(1−𝔼​e−λ​V\)\.n\\log\\mathbb\{E\}e^\{\-\\lambda V\}\\leq\-n\\bigl\(1\-\\mathbb\{E\}e^\{\-\\lambda V\}\\bigr\)\.Usingc−1=log⁡n\+ac^\{\-1\}=\\log n\+ain \([96](https://arxiv.org/html/2609.30296#S6.E96)\) therefore gives

ℙ⁡\(Xn\>x\)≤exp⁡\{λ⁡\(log⁡λ\+Aa−x\)\},1≤λ≤n,\\mathbb\{P\}\(X\_\{n\}\>x\)\\leq\\exp\\left\\\{\\lambda\(\\log\\lambda\+A\_\{a\}\-x\)\\right\\\},\\qquad 1\\leq\\lambda\\leq n,\(99\)withAa:=a\+CaA\_\{a\}:=a\+C\_\{a\}, enlargingCaC\_\{a\}if necessary so thatCa≥0C\_\{a\}\\geq 0\.

Choosing

λ=ex−Aa−1\\lambda=e^\{x\-A\_\{a\}\-1\}whenever this lies in\[1,n\]\[1,n\]makes the exponent equal to−λ\-\\lambda, giving the double\-exponential bound

ℙ⁡\(Xn\>x\)≤exp⁡\{−ex−Aa−1\},Aa\+1≤x≤log⁡n\+Aa\+1\.\\mathbb\{P\}\(X\_\{n\}\>x\)\\leq\\exp\\left\\\{\-e^\{\\,x\-A\_\{a\}\-1\}\\right\\\},\\qquad A\_\{a\}\+1\\leq x\\leq\\log n\+A\_\{a\}\+1\.\(100\)Forx<Aa\+1x<A\_\{a\}\+1the trivial boundℙ≤1\\mathbb\{P\}\\leq 1suffices, and forx\>log⁡n\+Aa\+1x\>\\log n\+A\_\{a\}\+1we havex\>1/cx\>1/c\(asCa\>−1C\_\{a\}\>\-1\), soℙ⁡\(Xn\>x\)=0\\mathbb\{P\}\(X\_\{n\}\>x\)=0becauseXn≤1/cX\_\{n\}\\leq 1/calmost surely\.

Thus both the positive and negative logarithmic tails are uniformly integrable\. Combining this with Lemma[3](https://arxiv.org/html/2609.30296#Thmtheorem3), and noting that the limiting stable law has a continuous positive density, so thatℙ⁡\(Xa=0\)=0\\mathbb\{P\}\(X\_\{a\}=0\)=0andPa\>0P\_\{a\}\>0, gives \([93](https://arxiv.org/html/2609.30296#S6.E93)\)–\([95](https://arxiv.org/html/2609.30296#S6.E95)\)\. ∎

### 6\.3Asymptotics of the Rayleigh quotient

We now return to the exact Rayleigh quotient\.

Since1−Sn=c​Xn1\-S\_\{n\}=cX\_\{n\}, equations \([64](https://arxiv.org/html/2609.30296#S6.E64)\) and \([65](https://arxiv.org/html/2609.30296#S6.E65)\) give, on\{Xn\>0\}\\\{X\_\{n\}\>0\\\},

G⁡\(1−Sn\)\\displaystyle G\(1\-S\_\{n\}\)=1n​log⁡\(1\+n​Xn\),\\displaystyle=\\frac\{1\}\{n\}\\log\(1\+nX\_\{n\}\),\(101\)H⁡\(1−Sn\)m0\\displaystyle\\frac\{H\(1\-S\_\{n\}\)\}\{m\_\{0\}\}=\(c\+n\)​Xn1\+n​Xn\.\\displaystyle=\\frac\{\(c\+n\)X\_\{n\}\}\{1\+nX\_\{n\}\}\.\(102\)Therefore

R⁡\(Fk,c\)=c⁡\(1\+1n\)​\(1\+cn\)​𝔼⁡\[log2⁡\(1\+n​Xn\);Xn\>0\]𝔼⁡\[\(c\+n\)​Xn1\+n​Xn;Xn\>0\]\.R\(F\_\{k,c\}\)=c\\left\(1\+\\frac\{1\}\{n\}\\right\)\\left\(1\+\\frac\{c\}\{n\}\\right\)\\frac\{\\mathbb\{E\}\[\\log^\{2\}\(1\+nX\_\{n\}\);\\,X\_\{n\}\>0\]\}\{\\mathbb\{E\}\[\\frac\{\(c\+n\)X\_\{n\}\}\{1\+nX\_\{n\}\};\\,X\_\{n\}\>0\]\}\.\(103\)
Let

ℓ:=log⁡n,c=1ℓ\+a\.\\ell:=\\log n,\\qquad c=\\frac\{1\}\{\\ell\+a\}\.
###### Lemma 6\(Denominator asymptotics\)\.

For fixedaa,

𝔼\[\(c\+n\)​Xn1\+n​Xn;Xn\>0\]=Pn\+O\(n−1/2\)\.\\mathbb\{E\}\\\!\\left\[\\frac\{\(c\+n\)X\_\{n\}\}\{1\+nX\_\{n\}\};\\,X\_\{n\}\>0\\right\]=P\_\{n\}\+O\\bigl\(n^\{\-1/2\}\\bigr\)\.\(104\)

###### Proof\.

For0<Xn≤1/c0<X\_\{n\}\\leq 1/c,

0≤1−\(c\+n\)​Xn1\+n​Xn=1−c​Xn1\+n​Xn≤11\+n​Xn\.0\\leq 1\-\\frac\{\(c\+n\)X\_\{n\}\}\{1\+nX\_\{n\}\}=\\frac\{1\-cX\_\{n\}\}\{1\+nX\_\{n\}\}\\leq\\frac\{1\}\{1\+nX\_\{n\}\}\.Choose

δn=n−1/2\.\\delta\_\{n\}=n^\{\-1/2\}\.On0<Xn≤δn0<X\_\{n\}\\leq\\delta\_\{n\}, Lemma[4](https://arxiv.org/html/2609.30296#Thmtheorem4)gives probabilityO⁡\(δn\)O\(\\delta\_\{n\}\)\. OnXn\>δnX\_\{n\}\>\\delta\_\{n\},

11\+n​Xn≤1n​δn=n−1/2\.\\frac\{1\}\{1\+nX\_\{n\}\}\\leq\\frac\{1\}\{n\\delta\_\{n\}\}=n^\{\-1/2\}\.Hence the difference between the left\-hand side of \([104](https://arxiv.org/html/2609.30296#S6.E104)\) andPnP\_\{n\}isO\(n−1/2\)O\(n^\{\-1/2\}\)\. ∎

###### Lemma 8\(Numerator asymptotics\)\.

For fixedaa,

𝔼\[log2\(1\+nXn\);Xn\>0\]=ℓ2Pn\+2ℓLn\+Qn\+O\(ℓ2n−1/2\)\.\\mathbb\{E\}\[\\log^\{2\}\(1\+nX\_\{n\}\);\\,X\_\{n\}\>0\]=\\ell^\{2\}P\_\{n\}\+2\\ell L\_\{n\}\+Q\_\{n\}\+O\\bigl\(\\ell^\{2\}n^\{\-1/2\}\\bigr\)\.\(105\)

###### Proof\.

Letδn=n−1/2\\delta\_\{n\}=n^\{\-1/2\}\.

*RegionXn\>δnX\_\{n\}\>\\delta\_\{n\}\.*Here

log\(1\+nXn\)=ℓ\+logXn\+rn,0≤rn=log\(1\+1n​Xn\)≤1n​Xn≤n−1/2,\\log\(1\+nX\_\{n\}\)=\\ell\+\\log X\_\{n\}\+r\_\{n\},\\qquad 0\\leq r\_\{n\}=\\log\\left\(1\+\\frac\{1\}\{nX\_\{n\}\}\\right\)\\leq\\frac\{1\}\{nX\_\{n\}\}\\leq n^\{\-1/2\},so that

log2⁡\(1\+n​Xn\)=\(ℓ\+log⁡Xn\)2\+2​\(ℓ\+log⁡Xn\)​rn\+rn2\.\\log^\{2\}\(1\+nX\_\{n\}\)=\(\\ell\+\\log X\_\{n\}\)^\{2\}\+2\(\\ell\+\\log X\_\{n\}\)r\_\{n\}\+r\_\{n\}^\{2\}\.By Lemma[5](https://arxiv.org/html/2609.30296#Thmtheorem5)the quantities𝔼⁡\[\|log⁡Xn\|;Xn\>0\]\\mathbb\{E\}\[\|\\log X\_\{n\}\|;X\_\{n\}\>0\]are bounded uniformly innn, so

𝔼\[\|2\(ℓ\+logXn\)rn\|\+rn2;Xn\>δn\]=O\(ℓn−1/2\)\.\\mathbb\{E\}\\bigl\[\\,\\bigl\|2\(\\ell\+\\log X\_\{n\}\)r\_\{n\}\\bigr\|\+r\_\{n\}^\{2\};\\,X\_\{n\}\>\\delta\_\{n\}\\bigr\]=O\\bigl\(\\ell n^\{\-1/2\}\\bigr\)\.
*Region0<Xn≤δn0<X\_\{n\}\\leq\\delta\_\{n\}\.*The uniform density bound of Lemma[4](https://arxiv.org/html/2609.30296#Thmtheorem4)gives

𝔼\[\(ℓ\+logXn\)2;0<Xn≤δn\]\\displaystyle\\mathbb\{E\}\\\!\\left\[\(\\ell\+\\log X\_\{n\}\)^\{2\};\\,0<X\_\{n\}\\leq\\delta\_\{n\}\\right\]≤C​∫0δn\(ℓ\+log⁡x\)2​𝑑x\\displaystyle\\leq C\\int\_\{0\}^\{\\delta\_\{n\}\}\(\\ell\+\\log x\)^\{2\}\\,dx=Cδn\[\(ℓ\+logδn\)2−2\(ℓ\+logδn\)\+2\]=O\(ℓ2n−1/2\),\\displaystyle=C\\delta\_\{n\}\\Bigl\[\(\\ell\+\\log\\delta\_\{n\}\)^\{2\}\-2\(\\ell\+\\log\\delta\_\{n\}\)\+2\\Bigr\]=O\\bigl\(\\ell^\{2\}n^\{\-1/2\}\\bigr\),\(106\)usingℓ\+log⁡δn=ℓ/2\\ell\+\\log\\delta\_\{n\}=\\ell/2\. Likewise, sincen​Xn≤n​δn=n1/2nX\_\{n\}\\leq n\\delta\_\{n\}=n^\{1/2\}on this region,

𝔼\[log2\(1\+nXn\);0<Xn≤δn\]≤C∫0δnlog2\(1\+nx\)dx=O\(ℓ2n−1/2\)\.\\mathbb\{E\}\[\\log^\{2\}\(1\+nX\_\{n\}\);\\,0<X\_\{n\}\\leq\\delta\_\{n\}\]\\leq C\\int\_\{0\}^\{\\delta\_\{n\}\}\\log^\{2\}\(1\+nx\)\\,dx=O\\bigl\(\\ell^\{2\}n^\{\-1/2\}\\bigr\)\.
Adding the two regions and recognising𝔼⁡\[\(ℓ\+log⁡Xn\)2;Xn\>0\]=ℓ2​Pn\+2​ℓ​Ln\+Qn\\mathbb\{E\}\[\(\\ell\+\\log X\_\{n\}\)^\{2\};X\_\{n\}\>0\]=\\ell^\{2\}P\_\{n\}\+2\\ell L\_\{n\}\+Q\_\{n\}proves \([105](https://arxiv.org/html/2609.30296#S6.E105)\)\. ∎

### 6\.4Main asymptotic theorem

###### Theorem 9\(Asymptotic defect of the rational trial family\)\.

Fixa∈ℝa\\in\\mathbb\{R\}\. For all sufficiently largekk, put

n=k−1,ck=1log⁡n\+a\.n=k\-1,\\qquad c\_\{k\}=\\frac\{1\}\{\\log n\+a\}\.For all sufficiently largekk,ck\>0c\_\{k\}\>0\. LetFk,ckF\_\{k,c\_\{k\}\}be the single\-channel rational trial function \([60](https://arxiv.org/html/2609.30296#S6.E60)\)\. Then

R⁡\(Fk,ck\)=log⁡k−𝒞⁡\(a\)\+o⁡\(1\),k→∞,R\(F\_\{k,c\_\{k\}\}\)=\\log k\-\\mathcal\{C\}\(a\)\+o\(1\),\\qquad k\\to\\infty,\(107\)where

𝒞⁡\(a\)=a−2​𝔼​\[log⁡Xa∣Xa\>0\]\\boxed\{\\mathcal\{C\}\(a\)=a\-2\\,\\mathbb\{E\}\\\!\\left\[\\log X\_\{a\}\\mid X\_\{a\}\>0\\right\]\}\(108\)and

Xa=\(a\+1\)−ΛX\_\{a\}=\(a\+1\)\-\\LambdawithΛ\\Lambdadefined by \([70](https://arxiv.org/html/2609.30296#S6.E70)\)\.

Equivalently,

limk→∞\(log⁡k−R⁡\(Fk,ck\)\)=𝒞⁡\(a\)\.\\lim\_\{k\\to\\infty\}\\left\(\\log k\-R\(F\_\{k,c\_\{k\}\}\)\\right\)=\\mathcal\{C\}\(a\)\.\(109\)

###### Proof\.

WriteNum\\mathrm\{Num\}andDen\\mathrm\{Den\}for the numerator and denominator in \([103](https://arxiv.org/html/2609.30296#S6.E103)\)\. By Lemmas[6](https://arxiv.org/html/2609.30296#Thmtheorem6)and[8](https://arxiv.org/html/2609.30296#Thmtheorem8),

Num=ℓ2Pn\+2ℓLn\+Qn\+O\(ℓ2n−1/2\)=O\(ℓ2\),Den=Pn\+O\(n−1/2\),\\mathrm\{Num\}=\\ell^\{2\}P\_\{n\}\+2\\ell L\_\{n\}\+Q\_\{n\}\+O\\bigl\(\\ell^\{2\}n^\{\-1/2\}\\bigr\)=O\(\\ell^\{2\}\),\\qquad\\mathrm\{Den\}=P\_\{n\}\+O\\bigl\(n^\{\-1/2\}\\bigr\),andPn→Pa\>0P\_\{n\}\\to P\_\{a\}\>0by Lemma[5](https://arxiv.org/html/2609.30296#Thmtheorem5), soDen\\mathrm\{Den\}is bounded away from00for largenn\. Hence

NumDen=ℓ2​Pn\+2​ℓ​Ln\+QnPn\+O\(ℓ2n−1/2\)\.\\frac\{\\mathrm\{Num\}\}\{\\mathrm\{Den\}\}=\\frac\{\\ell^\{2\}P\_\{n\}\+2\\ell L\_\{n\}\+Q\_\{n\}\}\{P\_\{n\}\}\+O\\bigl\(\\ell^\{2\}n^\{\-1/2\}\\bigr\)\.Multiplying byc⁡\(1\+1n\)​\(1\+cn\)=c⁡\(1\+O⁡\(n−1\)\)c\(1\+\\tfrac\{1\}\{n\}\)\(1\+\\tfrac\{c\}\{n\}\)=c\\bigl\(1\+O\(n^\{\-1\}\)\\bigr\)and usingc=\(ℓ\+a\)−1c=\(\\ell\+a\)^\{\-1\},

R\(Fk,c\)=ℓ2\+2​ℓ​Ln/Pn\+Qn/Pnℓ\+a\+O\(ℓn−1/2\)\.R\(F\_\{k,c\}\)=\\frac\{\\ell^\{2\}\+2\\ell\\,L\_\{n\}/P\_\{n\}\+Q\_\{n\}/P\_\{n\}\}\{\\ell\+a\}\+O\\bigl\(\\ell\\,n^\{\-1/2\}\\bigr\)\.\(110\)A direct rearrangement yields the finite\-nnexpansion

ℓ−R⁡\(Fk,c\)\\displaystyle\\ell\-R\(F\_\{k,c\}\)=a−2​LnPn\\displaystyle=a\-2\\frac\{L\_\{n\}\}\{P\_\{n\}\}−Qn/Pn−2​a​Ln/Pn\+a2ℓ\+a\+O\(ℓn−1/2\)\.\\displaystyle\\quad\-\\frac\{Q\_\{n\}/P\_\{n\}\-2aL\_\{n\}/P\_\{n\}\+a^\{2\}\}\{\\ell\+a\}\+O\\bigl\(\\ell\\,n^\{\-1/2\}\\bigr\)\.\(111\)By Lemma[5](https://arxiv.org/html/2609.30296#Thmtheorem5),

LnPn⟶LaPa=𝔼⁡\[log⁡Xa∣Xa\>0\],\\frac\{L\_\{n\}\}\{P\_\{n\}\}\\longrightarrow\\frac\{L\_\{a\}\}\{P\_\{a\}\}=\\mathbb\{E\}\[\\log X\_\{a\}\\mid X\_\{a\}\>0\],while the fraction on the second line of \([111](https://arxiv.org/html/2609.30296#S6.E111)\) tends to zero\. Hence

ℓ−R⁡\(Fk,c\)⟶a−2​𝔼​\[log⁡Xa∣Xa\>0\]\.\\ell\-R\(F\_\{k,c\}\)\\longrightarrow a\-2\\mathbb\{E\}\[\\log X\_\{a\}\\mid X\_\{a\}\>0\]\.Finally,

log⁡k−log⁡\(k−1\)=O⁡\(n−1\),\\log k\-\\log\(k\-1\)=O\(n^\{\-1\}\),soℓ\\ellmay be replaced bylog⁡k\\log k\. This proves \([107](https://arxiv.org/html/2609.30296#S6.E107)\) and \([109](https://arxiv.org/html/2609.30296#S6.E109)\)\. ∎

### 6\.5Consequence forMkM\_\{k\}

Define

𝒞∗:=infa∈ℝ𝒞⁡\(a\)\.\\mathcal\{C\}\_\{\*\}:=\\inf\_\{a\\in\\mathbb\{R\}\}\\mathcal\{C\}\(a\)\.\(112\)
###### Corollary 10\(Asymptotic lower bound forMkM\_\{k\}\)\.

The Maynard variational constant satisfies

lim infk→∞\(Mk−log⁡k\)≥−𝒞∗\.\\liminf\_\{k\\to\\infty\}\\bigl\(M\_\{k\}\-\\log k\\bigr\)\\geq\-\\mathcal\{C\}\_\{\*\}\.\(113\)Equivalently,

Mk≥log⁡k−𝒞∗−o⁡\(1\)\.M\_\{k\}\\geq\\log k\-\\mathcal\{C\}\_\{\*\}\-o\(1\)\.\(114\)

###### Proof\.

For everyε\>0\\varepsilon\>0, choose a fixedaε∈ℝa\_\{\\varepsilon\}\\in\\mathbb\{R\}such that

𝒞⁡\(aε\)≤𝒞∗\+ε\.\\mathcal\{C\}\(a\_\{\\varepsilon\}\)\\leq\\mathcal\{C\}\_\{\*\}\+\\varepsilon\.SinceMkM\_\{k\}is the supremum over admissible trial functions, Theorem[9](https://arxiv.org/html/2609.30296#Thmtheorem9)gives

Mk≥R⁡\(Fk,ck\)=log⁡k−𝒞⁡\(aε\)\+o⁡\(1\)M\_\{k\}\\geq R\(F\_\{k,c\_\{k\}\}\)=\\log k\-\\mathcal\{C\}\(a\_\{\\varepsilon\}\)\+o\(1\)for the choice

ck−1=log⁡\(k−1\)\+aε\.c\_\{k\}^\{\-1\}=\\log\(k\-1\)\+a\_\{\\varepsilon\}\.Thus

lim infk→∞\(Mk−log⁡k\)≥−𝒞∗−ε\.\\liminf\_\{k\\to\\infty\}\(M\_\{k\}\-\\log k\)\\geq\-\\mathcal\{C\}\_\{\*\}\-\\varepsilon\.Lettingε↓0\\varepsilon\\downarrow 0proves \([113](https://arxiv.org/html/2609.30296#S6.E113)\)\. ∎

The infimum in \([112](https://arxiv.org/html/2609.30296#S6.E112)\) may be restricted to a half\-line bounded above, since𝒞\\mathcal\{C\}grows without bound in the positive direction\.

###### Lemma 11\(Growth of𝒞\\mathcal\{C\}for large positive shift\)\.

Withκ=𝔼⁡\[\(−Λ\)\+\]<∞\\kappa=\\mathbb\{E\}\[\(\-\\Lambda\)^\{\+\}\]<\\inftyas in \([73](https://arxiv.org/html/2609.30296#S6.E73)\), there isa0a\_\{0\}such that for alla≥a0a\\geq a\_\{0\},

𝒞⁡\(a\)≥a−2​log⁡\(2​\(a\+1\+κ\)\)→a→∞∞\.\\mathcal\{C\}\(a\)\\ \\geq\\ a\-2\\log\\bigl\(2\(a\+1\+\\kappa\)\\bigr\)\\ \\xrightarrow\[a\\to\\infty\]\{\}\\ \\infty\.\(115\)In particularinfa∈ℝ𝒞⁡\(a\)=infa≤a1𝒞⁡\(a\)\\inf\_\{a\\in\\mathbb\{R\}\}\\mathcal\{C\}\(a\)=\\inf\_\{a\\leq a\_\{1\}\}\\mathcal\{C\}\(a\)for some finitea1a\_\{1\}\.

###### Proof\.

SinceXa\+=\(\(a\+1\)−Λ\)\+≤\(a\+1\)\+\(−Λ\)\+X\_\{a\}^\{\+\}=\(\(a\+1\)\-\\Lambda\)^\{\+\}\\leq\(a\+1\)\+\(\-\\Lambda\)^\{\+\}fora≥−1a\\geq\-1,

𝔼⁡\[Xa;Xa\>0\]≤\(a\+1\)\+κ\.\\mathbb\{E\}\[X\_\{a\};X\_\{a\}\>0\]\\leq\(a\+1\)\+\\kappa\.AlsoPa=ℙ⁡\(Λ<a\+1\)→1P\_\{a\}=\\mathbb\{P\}\(\\Lambda<a\+1\)\\to 1, soPa≥12P\_\{a\}\\geq\\tfrac\{1\}\{2\}fora≥a0a\\geq a\_\{0\}\. By Jensen’s inequality applied to the concave functionlog\\logunder the conditional law,

𝔼⁡\[log⁡Xa∣Xa\>0\]≤log⁡𝔼⁡\[Xa∣Xa\>0\]=log⁡𝔼⁡\[Xa;Xa\>0\]Pa≤log⁡\(2​\(a\+1\+κ\)\),\\mathbb\{E\}\[\\log X\_\{a\}\\mid X\_\{a\}\>0\]\\leq\\log\\mathbb\{E\}\[X\_\{a\}\\mid X\_\{a\}\>0\]=\\log\\frac\{\\mathbb\{E\}\[X\_\{a\};X\_\{a\}\>0\]\}\{P\_\{a\}\}\\leq\\log\\bigl\(2\(a\+1\+\\kappa\)\\bigr\),which gives \([115](https://arxiv.org/html/2609.30296#S6.E115)\)\. ∎

### 6\.6Numerical interpretation

Theorem[9](https://arxiv.org/html/2609.30296#Thmtheorem9)is an analytic statement and does not depend on the numerical discovery or certification pipeline\. Numerical quadrature of the limiting stable law can, however, be used to evaluate𝒞⁡\(a\)\\mathcal\{C\}\(a\)and to locate a favorable fixed shift\. Two independent routes are available: direct quadrature against the density ofΛ\\Lambda, and inversion of the explicit Laplace transform \([71](https://arxiv.org/html/2609.30296#S6.E71)\)\. Our numerical evaluation indicates a minimum near a small negative value ofaaand an asymptotic defect of approximately

𝒞∗≈0\.3343\.\\mathcal\{C\}\_\{\*\}\\approx 0\.3343\.\(116\)The value in \([116](https://arxiv.org/html/2609.30296#S6.E116)\) is a numerical estimate of the infimum, not a certified enclosure established by this asymptotic analysis\. Certifying a finite\-kkRayleigh quotient and certifying the limiting constant are separate tasks\.

For an explicit asymptotic lower bound it suffices to choose a single fixed shifta0a\_\{0\}and establish, by rigorous interval quadrature with controlled truncation errors, an upper bound𝒞⁡\(a0\)≤U\\mathcal\{C\}\(a\_\{0\}\)\\leq U\. Theorem[9](https://arxiv.org/html/2609.30296#Thmtheorem9)then gives

lim infk→∞\(Mk−log⁡k\)≥−U,\\liminf\_\{k\\to\\infty\}\(M\_\{k\}\-\\log k\)\\geq\-U,without any proof of global minimality\. By contrast, a rigorous enclosure of𝒞∗\\mathcal\{C\}\_\{\*\}itself also requires a lower bound valid for everya∈ℝa\\in\\mathbb\{R\}\. Any numerical search restricted to a finite interval must therefore be supplemented by rigorous control of both excluded tails, includinga→−∞a\\to\-\\infty\. The upper bound on the left tail ofΛ\\Lambdaproved above does not by itself supply that control\.

The finite\-kkoptimization exhibits the same qualitative scaling:

ck∼1log⁡k,R⁡\(Fk,ck\)=log⁡k−O⁡\(1\)\.c\_\{k\}\\sim\\frac\{1\}\{\\log k\},\\qquad R\(F\_\{k,c\_\{k\}\}\)=\\log k\-O\(1\)\.The fact that the discrepancylog⁡k−R⁡\(Fk,ck\)\\log k\-R\(F\_\{k,c\_\{k\}\}\)varies only slowly over more than five orders of magnitude inkkis therefore explained by the stable\-law limit rather than being an empirical coincidence\.

### 6\.7Finite\-size corrections

Expansion \([111](https://arxiv.org/html/2609.30296#S6.E111)\) also isolates the finite\-size correction\. All error terms accumulated in its derivation areO\(ℓn−1/2\)O\(\\ell\\,n^\{\-1/2\}\), henceo⁡\(ℓ−m\)o\(\\ell^\{\-m\}\)for every fixedmm; the replacement oflog⁡n\\log nbylog⁡k\\log kcontributesO⁡\(n−1\)O\(n^\{\-1\}\)\. Consequently

log⁡k−R⁡\(Fk,c\)\\displaystyle\\log k\-R\(F\_\{k,c\}\)=a−2LnPn−Qn/Pn−2​a​Ln/Pn\+a2ℓ\+a\+O\(ℓn−1/2\),\\displaystyle=a\-2\\frac\{L\_\{n\}\}\{P\_\{n\}\}\-\\frac\{Q\_\{n\}/P\_\{n\}\-2aL\_\{n\}/P\_\{n\}\+a^\{2\}\}\{\\ell\+a\}\+O\\bigl\(\\ell\\,n^\{\-1/2\}\\bigr\),\(117\)Define

dn​\(a\):=LnPn−LaPa,B⁡\(a\):=QaPa−2​a​LaPa\+a2\.d\_\{n\}\(a\):=\\frac\{L\_\{n\}\}\{P\_\{n\}\}\-\\frac\{L\_\{a\}\}\{P\_\{a\}\},\\qquad B\(a\):=\\frac\{Q\_\{a\}\}\{P\_\{a\}\}\-2a\\frac\{L\_\{a\}\}\{P\_\{a\}\}\+a^\{2\}\.SincePn→Pa\>0P\_\{n\}\\to P\_\{a\}\>0,Ln→LaL\_\{n\}\\to L\_\{a\}, andQn→QaQ\_\{n\}\\to Q\_\{a\}, the expansion above implies

log⁡k−R⁡\(Fk,c\)=𝒞⁡\(a\)−2​dn​\(a\)−B⁡\(a\)ℓ\+o⁡\(ℓ−1\)\.\\log k\-R\(F\_\{k,c\}\)=\\mathcal\{C\}\(a\)\-2d\_\{n\}\(a\)\-\\frac\{B\(a\)\}\{\\ell\}\+o\(\\ell^\{\-1\}\)\.\(118\)Heredn​\(a\)→0d\_\{n\}\(a\)\\to 0, but no quantitative rate for this convergence has been proved\. Thusdn​\(a\)d\_\{n\}\(a\)may dominate the displayed1/ℓ1/\\ellterm, and \([118](https://arxiv.org/html/2609.30296#S6.E118)\) is not yet a two\-term asymptotic expansion with a determined coefficient\.

A sufficient additional estimate is

dn​\(a\)=o⁡\(ℓ−1\),d\_\{n\}\(a\)=o\(\\ell^\{\-1\}\),in which case

log⁡k−R⁡\(Fk,c\)=𝒞⁡\(a\)−B⁡\(a\)log⁡k\+o⁡\(\(log⁡k\)−1\)\.\\log k\-R\(F\_\{k,c\}\)=\\mathcal\{C\}\(a\)\-\\frac\{B\(a\)\}\{\\log k\}\+o\\bigl\(\(\\log k\)^\{\-1\}\\bigr\)\.More generally, ifℓ​dn​\(a\)→b⁡\(a\)\\ell d\_\{n\}\(a\)\\to b\(a\), the coefficient of1/log⁡k1/\\log kis−\[B⁡\(a\)\+2​b​\(a\)\]\-\[B\(a\)\+2b\(a\)\]\. A bound of onlydn​\(a\)=O⁡\(ℓ−1\)d\_\{n\}\(a\)=O\(\\ell^\{\-1\}\)does not determine that coefficient\. Accordingly, a quantitative stable\-limit argument must establish an appropriate rate for these conditional logarithmic moments, or their first\-order asymptotics; a distributional convergence rate alone requires an additional argument to handle the logarithmic singularity\. Any fitted coefficient of1/log⁡k1/\\log kremains a numerical finite\-size observation until such an estimate is supplied\.

### 6\.8Interpretation

The asymptotic mechanism can be summarized as follows\. Squaring the single rational channel induces the one\-dimensional density

qn​\(v\)=c\+n\(1\+n​v\)2,q\_\{n\}\(v\)=\\frac\{c\+n\}\{\(1\+nv\)^\{2\}\},whose centerednn\-fold sum converges to a spectrally positive11\-stable law\. Spectral positivity refers to the jumps; the limiting law has support on all ofℝ\\mathbb\{R\}\. Choosing

c−1=log⁡n\+ac^\{\-1\}=\\log n\+acenters the boundary variable

Xn=1c−∑j=1nVjX\_\{n\}=\\frac\{1\}\{c\}\-\\sum\_\{j=1\}^\{n\}V\_\{j\}at a non\-degenerate limit\. Conditional on the eventXn\>0X\_\{n\}\>0, which is exactly the simplex eventSn<1S\_\{n\}<1, the numerator of the Rayleigh quotient contains

log2⁡\(1\+n​Xn\)=\(log⁡n\+log⁡Xn\+o⁡\(1\)\)2,\\log^\{2\}\(1\+nX\_\{n\}\)=\\bigl\(\\log n\+\\log X\_\{n\}\+o\(1\)\\bigr\)^\{2\},whereas the denominator converges to the same survival probability\. The leadinglog2⁡n\\log^\{2\}nterm is therefore divided byc−1∼log⁡nc^\{\-1\}\\sim\\log n, producing a Rayleigh quotient of sizelog⁡n\\log n\. The conditional logarithmic moment of the limiting stable variable supplies the non\-trivialO⁡\(1\)O\(1\)defect\.

Thus the rational family is not merely a finite\-dimensional numerical ansatz that happens to perform well at largekk\. Its observed scaling is the finite\-kkmanifestation of a stable\-law limit, and the family gives the explicit asymptotic lower bound

Mk≥log⁡k−𝒞∗−o⁡\(1\)\.M\_\{k\}\\geq\\log k\-\\mathcal\{C\}\_\{\*\}\-o\(1\)\.

## 7Supplementary Methods: higher\-order Delsarte hierarchy

Notation follows the Methods\. Throughout,𝖧⁡\(⋅\)\\mathsf\{H\}\(\\cdot\)denotes Shannon entropy in bits andH2​\(p\)=−p​log2​p−\(1−p\)​log2⁡\(1−p\)H\_\{2\}\(p\)=\-p\\log\_\{2\}p\-\(1\-p\)\\log\_\{2\}\(1\-p\)the binary entropy function;GGis a normalized configuration on𝔽2ℓ\\mathbb\{F\}\_\{2\}^\{\\ell\}, i\.e\. a probability distribution, andΦ⁡\(G,v\)=∑uG⁡\(u\)​G​\(u\+v\)\\Phi\(G,v\)=\\sum\_\{u\}\\sqrt\{G\(u\)G\(u\+v\)\}is the translation affinity at a shiftv≠0v\\neq 0\.

### 7\.1Scale\-up of the hierarchy tor≥3r\\geq 3

The finite\-length hierarchy was first implemented atr=2r=2, whereGL⁡\(2,2\)≅S3\\mathrm\{GL\}\(2,2\)\\cong S\_\{3\}acts as the full symmetric group on the three non\-zero vectors of𝔽22\\mathbb\{F\}\_\{2\}^\{2\}and orbit representatives can be generated by sorting\. This shortcut cannot be generalized tor=3r=3:\|GL⁡\(3,2\)\|=168\|\\mathrm\{GL\}\(3,2\)\|=168whereas\|S7\|=5040\|S\_\{7\}\|=5040\. Sorting therefore over\-merges genuinely distinct Fano\-plane orbits; atn=6n=6,4545true orbits collapse to3030sorted classes\. We instead enumerateGL⁡\(r,2\)\\mathrm\{GL\}\(r,2\)explicitly and construct genuine type orbits under the group action\.

ForS⊂\[r\]S\\subset\[r\], partial Fourier transforms factor after splitting𝔽2r=𝔽2S⊕𝔽2Sc\\mathbb\{F\}\_\{2\}^\{r\}=\\mathbb\{F\}\_\{2\}^\{S\}\\oplus\\mathbb\{F\}\_\{2\}^\{S^\{c\}\}:

KαS​\(β\)=∏w∈𝔽2ScKα⁡\(⋅,w\)\(\|S\|\)​\(β⁡\(⋅,w\)\),K^\{S\}\_\{\\alpha\}\(\\beta\)=\\prod\_\{w\\in\\mathbb\{F\}\_\{2\}^\{S^\{c\}\}\}K^\{\(\|S\|\)\}\_\{\\alpha\(\\cdot,w\)\}\\\!\\left\(\\beta\(\\cdot,w\)\\right\),with zero value unless theScS^\{c\}\-marginals ofα\\alphaandβ\\betaagree\. Each factor is a lower\-dimensional full Krawtchouk transform, so the implementation recurses through the same exact row generator\. Counting pairs of matrices of types\(α,β\)\(\\alpha,\\beta\)in the two possible orders gives the reciprocity

\(nβ\)​Kα​\(β\)=\(nα\)​Kβ​\(α\)\.\\binom\{n\}\{\\beta\}K\_\{\\alpha\}\(\\beta\)=\\binom\{n\}\{\\alpha\}K\_\{\\beta\}\(\\alpha\)\.Combined withGL⁡\(r,2\)\\mathrm\{GL\}\(r,2\)equivariance, this allows a row to be generated from one representative expansion per valid orbit\. Atr=3r=3,n=14n=14,d=6d=6, this reduces the number of expansions from12351235to112112\.

The resultingr=3r=3implementation reproduced ther=2r=2program when run atr=2r=2and produced certified strict level\-33separations, includingV2​\(9,5\)=44/9\>V3​\(9,5\)=4V\_\{2\}\(9,5\)=44/9\>V\_\{3\}\(9,5\)=4andV2​\(12,5\)=24\.260255​…\>V3​\(12,5\)=16V\_\{2\}\(12,5\)=24\.260255\\ldots\>V\_\{3\}\(12,5\)=16\.

### 7\.2Exact dual certification and implementation checks

The floating\-point solver is used as a proposal mechanism for an active dual set\. The corresponding basic system is reconstructed at arbitrary precision, rounded to dyadic rationals, and verified against every dual inequality\. The repair identity quoted in the Methods supplies an exactly feasible direction and an explicit rational repair cost, so the certified value is an exact rational upper bound on the level\-rroptimum obtained by weak duality\.

Three independent anchors were used\. First, full and partial Krawtchouk rows were compared with brute\-force character sums over𝔽2r×n\\mathbb\{F\}\_\{2\}^\{r\\times n\}for smallnn\. Second,r=1r=1was checked against an independent classical Delsarte implementation\. Third, the publishedr=2r=2values of Loyfer and Linial were reproduced, including24\.260255​…24\.260255\\ldotsat\(13,6\)\(13,6\),131\.720587​…131\.720587\\ldotsat\(16,6\)\(16,6\),3232at\(16,8\)\(16,8\)and\(17,8\)\(17,8\), and256256at\(17,6\)\(17,6\)and\(20,8\)\(20,8\)\.

Structural invariant tests detected defects that value reproduction did not\. An early active\-set implementation factorized a matrix shifted by one column relative to the solved system, yielding32\.0000887​…32\.0000887\\ldotsinstead of3232at\(17,8\)\(17,8\)\. Two precision\-state errors in dyadic rounding and polishing partially masked this defect\. These issues were corrected before the experiments reported here; all final bounds use the repaired exact\-verification path\.

### 7\.3Lift experiments and finite\-length structural map

Coregliano et al\. construct an explicit lift of a level\-one dual to levelℓ\\ellwith objectiveV1ℓV\_\{1\}^\{\\ell\}\[[5](https://arxiv.org/html/2609.30296#bib.bib13)\]\. Because the level\-ℓ\\ellhierarchy bounds\|C\|ℓ\|C\|^\{\\ell\}, the lift must be compared in this normalization\. An initial level\-one\-normalized comparison reproducedV1V\_\{1\}but violated higher\-level feasibility by\+9949\+9949, which served as a normalization control before the systematic sweep\.

The lift has no mass on partial families below the full transform\. We therefore compared the full\-Fourier\-only hierarchy with the complete partial\-Fourier program\. Across twenty tested instances the full\-only optimum never exceededV1rV\_\{1\}^\{r\}\. Atr=2r=2it was frequently exactly equal to the lift, for example

36=62\(9,5\),576=242\(11,5\),144=122\(10,5\),36=6^\{2\}\\quad\(9,5\),\\qquad 576=24^\{2\}\\quad\(11,5\),\\qquad 144=12^\{2\}\\quad\(10,5\),while the partial\-Fourier program improved strictly\. Atr=3r=3, by contrast, the full\-only hierarchy itself generally improved on the lift\. The correction to a lifted level\-one solution is therefore not a common object across levels\.

A separate scan at approximately fixed relative distance showed why these finite instances were not used for asymptotic learning\. Atδ=0\.40\\delta=0\.40, the accessible finite\-nnrates were0\.330\.33–0\.500\.50, compared with the first\-MRRW value0\.08150\.0815, and the level\-one\-to\-level\-two gain decreased withnnat fixeddd\. Extrapolation from the accessiblennrange would therefore have mixed strong finite\-size effects with the asymptotic target\.

### 7\.4Normalization of the configuration objective

For a level\-ℓ\\elldual,

\|C\|ℓ≤f⁡\(0\),\|C\|^\{\\ell\}\\leq f\(0\),so an asymptotic bound obtained from a normalized configurationGGis

R⁡\(C\)≤𝖧⁡\(G\)ℓ\+O⁡\(log⁡nn\),R\(C\)\\leq\\frac\{\\mathsf\{H\}\(G\)\}\{\\ell\}\+O\\\!\\left\(\\frac\{\\log n\}\{n\}\\right\),and𝖧⁡\(G\)\\mathsf\{H\}\(G\)must not be compared directly with an ordinary code\-rate bound\. Re\-deriving this normalization exposed two discrepancies in the cited preprint\. First, the leading small\-ε\\varepsiloncoefficient of the first MRRW expression is12​ε2​log2⁡\(1/ε\)\\tfrac\{1\}\{2\}\\varepsilon^\{2\}\\log\_\{2\}\(1/\\varepsilon\)rather than14​ε2​log2⁡\(1/ε\)\\tfrac\{1\}\{4\}\\varepsilon^\{2\}\\log\_\{2\}\(1/\\varepsilon\): withpε=ε24\+O⁡\(ε4\)p\_\{\\varepsilon\}=\\tfrac\{\\varepsilon^\{2\}\}\{4\}\+O\(\\varepsilon^\{4\}\)andlog2⁡\(1/pε\)=2​log2⁡\(1/ε\)\+2\+O⁡\(ε2\)\\log\_\{2\}\(1/p\_\{\\varepsilon\}\)=2\\log\_\{2\}\(1/\\varepsilon\)\+2\+O\(\\varepsilon^\{2\}\), one getsh⁡\(ε\)=12​ε2​log2⁡\(1/ε\)\+O⁡\(ε2\)h\(\\varepsilon\)=\\tfrac\{1\}\{2\}\\varepsilon^\{2\}\\log\_\{2\}\(1/\\varepsilon\)\+O\(\\varepsilon^\{2\}\)\. Second, the closed form stated in their Lemma 6\.10\(i\) is a lower bound on the walk count rather than an equality, since its proof passes through\(m/2F\)2≥\(m/2F\)\\binom\{m/2\}\{F\}^\{2\}\\geq\\binom\{m/2\}\{F\}\. Neither affects the qualitative conclusion of that work: the stated rate bound is off by a factor2​ℓ2\\ellrelative to𝖧⁡\(G\)/ℓ\\mathsf\{H\}\(G\)/\\ell— theℓ\\ellbeing the rate normalization above and the22the same halving — but the factor is applied on both sides of the comparison made there\.

### 7\.5Reduced configuration search

WithG=p2G=p^\{2\}, the configuration problem becomes

min‖p‖2=1,p≥0⁡1ℓ​𝖧​\(p2\)\\min\_\{\\\|p\\\|\_\{2\}=1,\\;p\\geq 0\}\\frac\{1\}\{\\ell\}\\mathsf\{H\}\(p^\{2\}\)subject to quadratic translation\-affinity constraintsp𝖳​Sv​p≥εp^\{\\mathsf\{T\}\}S\_\{v\}p\\geq\\varepsilon\. The largest search space considered \(ℓ=4\\ell=4\) has1515free probability coordinates\. We therefore used direct multistart sequential quadratic programming with analytic gradients rather than a neural network\. The campaign comprised2828instance solves overℓ≤4\\ell\\leq 4andε∈\{0\.30,0\.20,0\.10,0\.05\}\\varepsilon\\in\\\{0\.30,0\.20,0\.10,0\.05\\\}, using400400–600600randomized starts per instance \(13,60013\{,\}600local optimizations in total\)\. One start per instance used the CJJ vertex\-uniform configuration and the remaining starts were randomized positive\-sphere points with varying concentration\.

No run produced a configuration below the first MRRW benchmark\. Recovered objectives differed from MRRW by between−3\.2×10−14\-3\.2\\times 10^\{\-14\}and−2\.1×10−16\-2\.1\\times 10^\{\-16\}, consistent with floating\-point noise, and the recovered configurations agreed coordinatewise with the tensor\-product quasirandom family to within10−1110^\{\-11\}\.

### 7\.6The tensorization obstruction

#### 7\.6\.1Setting and hypotheses

LetGGbe a normalized configuration on𝔽2ℓ\\mathbb\{F\}\_\{2\}^\{\\ell\}, that is,G⁡\(u\)≥0G\(u\)\\geq 0and∑uG⁡\(u\)=1\\sum\_\{u\}G\(u\)=1\. Write𝖧⁡\(G\)\\mathsf\{H\}\(G\)for its Shannon entropy in bits,H2H\_\{2\}for binary entropy, and

Φ⁡\(G,v\)=∑uG⁡\(u\)​G​\(u\+v\)\\Phi\(G,v\)=\\sum\_\{u\}\\sqrt\{G\(u\)G\(u\+v\)\}for its translation affinity\. By Cauchy–Schwarz,0≤Φ⁡\(G,v\)≤10\\leq\\Phi\(G,v\)\\leq 1\.

For an even integerm≥2m\\geq 2andv≠0v\\neq 0, the*paired walk weight*is

Wm​\(G,v\)=∑Fm\!∏uF⁡\(u\)\!​∏uG​\(u\)F⁡\(u\),W\_\{m\}\(G,v\)=\\sum\_\{F\}\\frac\{m\!\}\{\\prod\_\{u\}F\(u\)\!\}\\prod\_\{u\}G\(u\)^\{F\(u\)\},\(119\)where the sum runs overF:𝔽2ℓ→ℤ≥0F:\\mathbb\{F\}\_\{2\}^\{\\ell\}\\to\\mathbb\{Z\}\_\{\\geq 0\}with∑uF⁡\(u\)=m\\sum\_\{u\}F\(u\)=mandF⁡\(u\)=F⁡\(u\+v\)F\(u\)=F\(u\+v\)for alluu, and00=10^\{0\}=1\. For fixedG,v,mG,v,m, this is the leading coefficient in the configuration asymptotics of\[[5](https://arxiv.org/html/2609.30296#bib.bib13)\]:

Avm​Λn,G​\(X\)=nm​Wm​\(G,v\)\+o⁡\(nm\),X∈confign,ℓ−1​\(n​G\)\.A\_\{v\}^\{m\}\\Lambda\_\{n,G\}\(X\)=n^\{m\}W\_\{m\}\(G,v\)\+o\(n^\{m\}\),\\qquad X\\in\\mathrm\{config\}\_\{n,\\ell\}^\{\-1\}\(nG\)\.\(120\)HereAvA\_\{v\}is the adjacency operator of the column\-translation graph andΛn,G\\Lambda\_\{n,G\}is the indicator of the configuration fibre\. The combinatorial definition \([119](https://arxiv.org/html/2609.30296#S7.E119)\) applies to every probability distributionGG; no limit innnis needed for Lemmas[14](https://arxiv.org/html/2609.30296#Thmtheorem14)–[17](https://arxiv.org/html/2609.30296#Thmtheorem17)or Theorem[18](https://arxiv.org/html/2609.30296#Thmtheorem18)\.

To interpret \([120](https://arxiv.org/html/2609.30296#S7.E120)\) for a fixedGG, let

NG=\{n∈ℕ:n​G∈Confign,ℓ\}N\_\{G\}=\\\{n\\in\\mathbb\{N\}:nG\\in\\mathrm\{Config\}\_\{n,\\ell\}\\\}and assume thatNGN\_\{G\}is infinite\. This holds wheneverGGhas rational entries\. Alongn→∞n\\to\\inftyinNGN\_\{G\}, every positive entry satisfiesn​G​\(u\)=Ω⁡\(n\)nG\(u\)=\\Omega\(n\), as required by the configuration asymptotics\. This lattice condition is needed only for the connection to finite configuration fibres, not for the entropy inequality below\.

Fixε∈\(0,1\)\\varepsilon\\in\(0,1\)\. We isolate the following two hypotheses:

- \(A1\)*Spanning\.*The setV⊆𝔽2ℓ∖\{0\}V\\subseteq\\mathbb\{F\}\_\{2\}^\{\\ell\}\\setminus\\\{0\\\}of tested shifts spans𝔽2ℓ\\mathbb\{F\}\_\{2\}^\{\\ell\}\.
- \(A2\)*Finite\-walk lower bound\.*For everyv∈Vv\\in Vthere exist an even integermv≥2m\_\{v\}\\geq 2and a constantcv≥1c\_\{v\}\\geq 1such that Wmv​\(G,v\)≥cv​εmv\.W\_\{m\_\{v\}\}\(G,v\)\\geq c\_\{v\}\\varepsilon^\{m\_\{v\}\}\.The ordersmvm\_\{v\}and constantscvc\_\{v\}are fixed independently ofnn\.

The spanning condition is equivalent to requiring that, for everyi≠0i\\neq 0, somev∈Vv\\in Vsatisfies⟨i,v⟩=1\\langle i,v\\rangle=1: failure to span is equivalent to containment ini⟂i^\{\\perp\}for some non\-zeroii\. Condition \(25\) of\[[5](https://arxiv.org/html/2609.30296#bib.bib13)\]is a sufficient finite\-nncondition for the cited spectral construction\. When it holds for a fixedGGand a fixed even walk order along infinitely many admissible blocklengths, \([120](https://arxiv.org/html/2609.30296#S7.E120)\) implies \(A1\)–\(A2\) withcv=22​ℓ−1c\_\{v\}=2^\{2\\ell\-1\}\. Indeed, there are only finitely many possible choices of witnessing shifts, so they can be held fixed on an infinite subsequence; division bynmn^\{m\}and passage to the limit then gives \(A2\)\. We use onlycv≥1c\_\{v\}\\geq 1\. Conversely, \(A2\) alone does not assert finite\-nnfeasibility of the spectral construction\.

#### 7\.6\.2The obstruction

###### Lemma 14\(Walk–affinity comparison\)\.

For every normalized configurationGG, everyv≠0v\\neq 0and every even integerm≥2m\\geq 2,

Wm​\(G,v\)≤Φ​\(G,v\)m\.W\_\{m\}\(G,v\)\\leq\\Phi\(G,v\)^\{m\}\.The inequality is strict wheneverΦ⁡\(G,v\)\>0\\Phi\(G,v\)\>0\.

###### Proof\.

Translation byv≠0v\\neq 0is a fixed\-point\-free involution of𝔽2ℓ\\mathbb\{F\}\_\{2\}^\{\\ell\}, so its orbits are the2ℓ−12^\{\\ell\-1\}unordered pairsP=\{u,u\+v\}P=\\\{u,u\+v\\\}\. A multiplicity functionFFadmissible in \([119](https://arxiv.org/html/2609.30296#S7.E119)\) is determined byaP:=F⁡\(u\)=F⁡\(u\+v\)a\_\{P\}:=F\(u\)=F\(u\+v\), with∑PaP=m/2\\sum\_\{P\}a\_\{P\}=m/2\. WritingxP=G⁡\(u\)​G​\(u\+v\)≥0x\_\{P\}=G\(u\)G\(u\+v\)\\geq 0gives

Wm​\(G,v\)=∑∑PaP=m/2m\!∏P\(aP\!\)2​∏PxPaP\.W\_\{m\}\(G,v\)=\\sum\_\{\\sum\_\{P\}a\_\{P\}=m/2\}\\frac\{m\!\}\{\\prod\_\{P\}\(a\_\{P\}\!\)^\{2\}\}\\prod\_\{P\}x\_\{P\}^\{a\_\{P\}\}\.\(121\)Both elements ofPPcontributexP\\sqrt\{x\_\{P\}\}toΦ⁡\(G,v\)\\Phi\(G,v\), soΦ⁡\(G,v\)=2​∑PxP\\Phi\(G,v\)=2\\sum\_\{P\}\\sqrt\{x\_\{P\}\}and

Φ​\(G,v\)m=2m​∑∑PcP=mm\!∏PcP\!​∏PxPcP/2\.\\Phi\(G,v\)^\{m\}=2^\{m\}\\\!\\\!\\sum\_\{\\sum\_\{P\}c\_\{P\}=m\}\\frac\{m\!\}\{\\prod\_\{P\}c\_\{P\}\!\}\\prod\_\{P\}x\_\{P\}^\{c\_\{P\}/2\}\.\(122\)Every term of \([122](https://arxiv.org/html/2609.30296#S7.E122)\) is non\-negative\. Discard those with somecPc\_\{P\}odd and match the remaining terms to \([121](https://arxiv.org/html/2609.30296#S7.E121)\) bycP=2​aPc\_\{P\}=2a\_\{P\}\. The ratio of the corresponding coefficients is

m\!/∏P\(aP\!\)22m​m\!/∏P\(2​aP\)\!=∏P\(2​aPaP\)4aP≤1\.\\frac\{m\!\\big/\\prod\_\{P\}\(a\_\{P\}\!\)^\{2\}\}\{2^\{m\}m\!\\big/\\prod\_\{P\}\(2a\_\{P\}\)\!\}=\\prod\_\{P\}\\frac\{\\binom\{2a\_\{P\}\}\{a\_\{P\}\}\}\{4^\{a\_\{P\}\}\}\\leq 1\.In fact this ratio is strictly less than one, sincem≥2m\\geq 2implies that someaP≥1a\_\{P\}\\geq 1, and\(2​aa\)<4a\\binom\{2a\}\{a\}<4^\{a\}fora≥1a\\geq 1\. IfΦ⁡\(G,v\)\>0\\Phi\(G,v\)\>0, somexP\>0x\_\{P\}\>0; assigningaP=m/2a\_\{P\}=m/2to this pair and zero to the others gives a positive matched term\. Summing therefore proves strict inequality in this case\. IfΦ⁡\(G,v\)=0\\Phi\(G,v\)=0, both sides vanish\. ∎

###### Lemma 15\(Affinity floor\)\.

Under\(A2\),Φ⁡\(G,v\)\>ε\\Phi\(G,v\)\>\\varepsilonfor everyv∈Vv\\in V\. In particular,Φ⁡\(G,v\)≥ε\\Phi\(G,v\)\\geq\\varepsilon\.

###### Proof\.

Hypothesis \(A2\) givesWmv​\(G,v\)≥cv​εmv\>0W\_\{m\_\{v\}\}\(G,v\)\\geq c\_\{v\}\\varepsilon^\{m\_\{v\}\}\>0, soΦ⁡\(G,v\)\>0\\Phi\(G,v\)\>0\. The strict part of Lemma[14](https://arxiv.org/html/2609.30296#Thmtheorem14)yields

Φ​\(G,v\)mv\>Wmv​\(G,v\)≥cv​εmv≥εmv\.\\Phi\(G,v\)^\{m\_\{v\}\}\>W\_\{m\_\{v\}\}\(G,v\)\\geq c\_\{v\}\\varepsilon^\{m\_\{v\}\}\\geq\\varepsilon^\{m\_\{v\}\}\.Takingmvm\_\{v\}\-th roots proves the claim\. No limit inmmornnis used\. ∎

###### Lemma 16\(Entropy profile\)\.

Letpε=12​\(1−1−ε2\)p\_\{\\varepsilon\}=\\tfrac\{1\}\{2\}\(1\-\\sqrt\{1\-\\varepsilon^\{2\}\}\)andh⁡\(ε\)=H2​\(pε\)h\(\\varepsilon\)=H\_\{2\}\(p\_\{\\varepsilon\}\)forε∈\[0,1\]\\varepsilon\\in\[0,1\]\. Thenhhis strictly increasing and strictly convex on\(0,1\)\(0,1\),h⁡\(0\)=0h\(0\)=0,h⁡\(1\)=1h\(1\)=1, and

h⁡\(ε\)=H2​\(12−δ⁡\(1−δ\)\),δ=1−ε2\.h\(\\varepsilon\)=H\_\{2\}\\\!\\left\(\\tfrac\{1\}\{2\}\-\\sqrt\{\\delta\(1\-\\delta\)\}\\right\),\\qquad\\delta=\\tfrac\{1\-\\varepsilon\}\{2\}\.\(123\)

###### Proof\.

Puts=1−ε2s=\\sqrt\{1\-\\varepsilon^\{2\}\}, sopε=\(1−s\)/2p\_\{\\varepsilon\}=\(1\-s\)/2andδ⁡\(1−δ\)=s2/4\\delta\(1\-\\delta\)=s^\{2\}/4\. Thus12−δ⁡\(1−δ\)=pε\\tfrac\{1\}\{2\}\-\\sqrt\{\\delta\(1\-\\delta\)\}=p\_\{\\varepsilon\}\. For0<ε<10<\\varepsilon<1, differentiation gives

H2′​\(pε\)=2​artanh⁡\(s\)ln⁡2,d​pεd​ε=ε2​s,H\_\{2\}^\{\\prime\}\(p\_\{\\varepsilon\}\)=\\frac\{2\\operatorname\{artanh\}\(s\)\}\{\\ln 2\},\\qquad\\frac\{dp\_\{\\varepsilon\}\}\{d\\varepsilon\}=\\frac\{\\varepsilon\}\{2s\},and hence

h′​\(ε\)=ε​artanh⁡\(s\)s​ln⁡2\>0,h′′​\(ε\)=artanh⁡\(s\)−ss3​ln⁡2\>0\.h^\{\\prime\}\(\\varepsilon\)=\\frac\{\\varepsilon\\operatorname\{artanh\}\(s\)\}\{s\\ln 2\}\>0,\\qquad h^\{\\prime\\prime\}\(\\varepsilon\)=\\frac\{\\operatorname\{artanh\}\(s\)\-s\}\{s^\{3\}\\ln 2\}\>0\.The second inequality follows fromartanh⁡\(s\)=s\+s3/3\+⋯\>s\\operatorname\{artanh\}\(s\)=s\+s^\{3\}/3\+\\cdots\>sfor0<s<10<s<1;h′′h^\{\\prime\\prime\}extends continuously toh′′​\(1\)=1/\(3​ln⁡2\)h^\{\\prime\\prime\}\(1\)=1/\(3\\ln 2\)\. The endpoint values follow fromp0=0p\_\{0\}=0andp1=1/2p\_\{1\}=1/2\. ∎

###### Lemma 17\(Conditional affinity and Jensen\)\.

LetX=\(X1,…,Xℓ\)∼GX=\(X\_\{1\},\\dots,X\_\{\\ell\}\)\\sim Gand, for fixedii, writeY=X−iY=X\_\{\-i\}andpY=Pr⁡\(Xi=1∣Y\)p\_\{Y\}=\\Pr\(X\_\{i\}=1\\mid Y\)\. Then

Φ⁡\(G,ei\)\\displaystyle\\Phi\(G,e\_\{i\}\)=𝔼Y​\[2​pY​\(1−pY\)\],\\displaystyle=\\mathbb\{E\}\_\{Y\}\\\!\\left\[2\\sqrt\{p\_\{Y\}\(1\-p\_\{Y\}\)\}\\right\],\(124\)𝖧⁡\(Xi∣X−i\)\\displaystyle\\mathsf\{H\}\(X\_\{i\}\\mid X\_\{\-i\}\)=𝔼Y​\[h⁡\(2​pY​\(1−pY\)\)\]\.\\displaystyle=\\mathbb\{E\}\_\{Y\}\\\!\\left\[h\\\!\\left\(2\\sqrt\{p\_\{Y\}\(1\-p\_\{Y\}\)\}\\right\)\\right\]\.Consequently,𝖧⁡\(Xi∣X−i\)≥h⁡\(Φ⁡\(G,ei\)\)\\mathsf\{H\}\(X\_\{i\}\\mid X\_\{\-i\}\)\\geq h\\big\(\\Phi\(G,e\_\{i\}\)\\big\)\.

###### Proof\.

Identifyuuwith\(y,xi\)\(y,x\_\{i\}\)\. For eachyyof positive probability, the two points\(y,0\)\(y,0\)and\(y,1\)\(y,1\)each contribute

G⁡\(y,0\)​G​\(y,1\)=Pr⁡\(Y=y\)​py​\(1−py\)\\sqrt\{G\(y,0\)G\(y,1\)\}=\\Pr\(Y=y\)\\sqrt\{p\_\{y\}\(1\-p\_\{y\}\)\}toΦ⁡\(G,ei\)\\Phi\(G,e\_\{i\}\), giving the first identity\. Values ofpyp\_\{y\}on zero\-probability events may be chosen arbitrarily\. Forb⁡\(p\):=2​p⁡\(1−p\)b\(p\):=2\\sqrt\{p\(1\-p\)\}, symmetry of binary entropy givesH2​\(p\)=h⁡\(b⁡\(p\)\)H\_\{2\}\(p\)=h\(b\(p\)\)for everyp∈\[0,1\]p\\in\[0,1\]\. Averaging overYYgives the second identity\. Jensen’s inequality and Lemma[16](https://arxiv.org/html/2609.30296#Thmtheorem16)give the conclusion\. ∎

###### Theorem 18\(Tensorization obstruction\)\.

Fixℓ≥1\\ell\\geq 1andε∈\(0,1\)\\varepsilon\\in\(0,1\)\. LetGGbe a normalized configuration on𝔽2ℓ\\mathbb\{F\}\_\{2\}^\{\\ell\}, and suppose that a spanning setV⊆𝔽2ℓ∖\{0\}V\\subseteq\\mathbb\{F\}\_\{2\}^\{\\ell\}\\setminus\\\{0\\\}satisfiesΦ⁡\(G,v\)≥ε\\Phi\(G,v\)\\geq\\varepsilonfor everyv∈Vv\\in V\. Then

Jℓ​\(G\)=𝖧⁡\(G\)ℓ≥h⁡\(ε\)=H2​\(12−δ⁡\(1−δ\)\),δ=1−ε2\.J\_\{\\ell\}\(G\)=\\frac\{\\mathsf\{H\}\(G\)\}\{\\ell\}\\geq h\(\\varepsilon\)=H\_\{2\}\\\!\\left\(\\tfrac\{1\}\{2\}\-\\sqrt\{\\delta\(1\-\\delta\)\}\\right\),\\qquad\\delta=\\tfrac\{1\-\\varepsilon\}\{2\}\.The right\-hand side is the first MRRW expression\. This bound is sharp for the relaxed problem with constraintsΦ⁡\(G,ei\)≥ε\\Phi\(G,e\_\{i\}\)\\geq\\varepsilon,1≤i≤ℓ1\\leq i\\leq\\ell:Gε=Bernoulli⁡\(pε\)⊗ℓG\_\{\\varepsilon\}=\\operatorname\{Bernoulli\}\(p\_\{\\varepsilon\}\)^\{\\otimes\\ell\}attains equality in that problem\. Under the stronger hypotheses\(A1\)–\(A2\), one hasJℓ​\(G\)\>h⁡\(ε\)J\_\{\\ell\}\(G\)\>h\(\\varepsilon\)\.

###### Proof\.

Choose a basisv1,…,vℓv\_\{1\},\\dots,v\_\{\\ell\}contained inVVand an invertible linear mapTTwithT​ei=viTe\_\{i\}=v\_\{i\}\. DefineG~​\(u\)=G​\(T​u\)\\widetilde\{G\}\(u\)=G\(Tu\)\. Then

𝖧⁡\(G~\)=𝖧⁡\(G\),Φ⁡\(G~,ei\)=Φ⁡\(G,vi\)≥ε\.\\mathsf\{H\}\(\\widetilde\{G\}\)=\\mathsf\{H\}\(G\),\\qquad\\Phi\(\\widetilde\{G\},e\_\{i\}\)=\\Phi\(G,v\_\{i\}\)\\geq\\varepsilon\.LetX∼G~X\\sim\\widetilde\{G\}\. The entropy chain rule and conditioning give

𝖧⁡\(G\)=∑i=1ℓ𝖧⁡\(Xi∣X1,…,Xi−1\)≥∑i=1ℓ𝖧⁡\(Xi∣X−i\)≥∑i=1ℓh⁡\(Φ⁡\(G~,ei\)\)≥ℓ​h​\(ε\),\\mathsf\{H\}\(G\)=\\sum\_\{i=1\}^\{\\ell\}\\mathsf\{H\}\(X\_\{i\}\\mid X\_\{1\},\\dots,X\_\{i\-1\}\)\\geq\\sum\_\{i=1\}^\{\\ell\}\\mathsf\{H\}\(X\_\{i\}\\mid X\_\{\-i\}\)\\geq\\sum\_\{i=1\}^\{\\ell\}h\\big\(\\Phi\(\\widetilde\{G\},e\_\{i\}\)\\big\)\\geq\\ell h\(\\varepsilon\),where the second inequality uses Lemma[17](https://arxiv.org/html/2609.30296#Thmtheorem17)and the last uses monotonicity ofhh\. Division byℓ\\ellproves the bound\.

For sharpness of the relaxed problem, affinity is multiplicative over tensor products, so

Φ⁡\(Gε,ei\)=2​pε​\(1−pε\)=ε,𝖧⁡\(Gε\)=ℓ​H2​\(pε\)=ℓ​h​\(ε\)\.\\Phi\(G\_\{\\varepsilon\},e\_\{i\}\)=2\\sqrt\{p\_\{\\varepsilon\}\(1\-p\_\{\\varepsilon\}\)\}=\\varepsilon,\\qquad\\mathsf\{H\}\(G\_\{\\varepsilon\}\)=\\ell H\_\{2\}\(p\_\{\\varepsilon\}\)=\\ell h\(\\varepsilon\)\.This proves attainment in the stated relaxation, not under \(A2\)\.

Finally, under \(A1\)–\(A2\), Lemma[15](https://arxiv.org/html/2609.30296#Thmtheorem15)givesΦ⁡\(G,vi\)\>ε\\Phi\(G,v\_\{i\}\)\>\\varepsilonfor every basis shift\. Strict monotonicity ofhhmakes the last inequality in the entropy chain strict, yieldingJℓ​\(G\)\>h⁡\(ε\)J\_\{\\ell\}\(G\)\>h\(\\varepsilon\)\. ∎

###### Corollary 19\.

Consider the entropy\-based rate estimate

R⁡\(C\)≤Jℓ​\(G\)\+O⁡\(log⁡n/n\)R\(C\)\\leq J\_\{\\ell\}\(G\)\+O\(\\log n/n\)for the fixed single\-orbit spectral construction \(Section[7\.4](https://arxiv.org/html/2609.30296#S7.SS4)\), withℓ\\ellandGGfixed and with its feasibility hypotheses satisfied\. IfGGsatisfies\(A1\)–\(A2\), then its leading term obeysJℓ​\(G\)\>h⁡\(ε\)J\_\{\\ell\}\(G\)\>h\(\\varepsilon\)\. Consequently, minimizing this entropy\-based estimate over configurations satisfying\(A1\)–\(A2\)cannot give an asymptotic rate bound below the first MRRW expression\. The relaxed affinity\-constrained minimum equalsh⁡\(ε\)h\(\\varepsilon\), but this does not assert attainment by a configuration satisfying the finite\-walk conditions\.

###### Proof\.

The rate normalization follows because a level\-ℓ\\elldual bounds\|C\|ℓ\|C\|^\{\\ell\}\. Apply Theorem[18](https://arxiv.org/html/2609.30296#Thmtheorem18)to the leading termJℓ​\(G\)J\_\{\\ell\}\(G\)of the stated estimate\. ∎

#### 7\.6\.3Scope and relation to the cited construction

### 7\.7Computational verification of the algebraic reductions

None of the computations in this section is used as part of the analytic proof\. They test that the implementation computes the objects Theorem[18](https://arxiv.org/html/2609.30296#Thmtheorem18)concerns, which is where a hidden factor ofm\!m\!, a2m2^\{m\}, or an ordered\-versus\-unordered pair convention would appear\.

The literal sum \([119](https://arxiv.org/html/2609.30296#S7.E119)\) over multiplicity functions was enumerated in exact rational arithmetic and matched the paired form \([121](https://arxiv.org/html/2609.30296#S7.E121)\) as an exact rational identity, not to a tolerance, for all testedℓ∈\{1,2,3\}\\ell\\in\\\{1,2,3\\\}, allv≠0v\\neq 0and all evenm≤8m\\leq 8\.

All downstream computation takes place on the configuration lattice, whereasAvA\_\{v\}acts on𝔽2ℓ×n\\mathbb\{F\}\_\{2\}^\{\\ell\\times n\}\. To confirm that this reduction loses nothing,AvA\_\{v\}was constructed literally on the full vertex set and applied as a matrix, rather than assuming thatAvm​Λ​\(X\)A\_\{v\}^\{m\}\\Lambda\(X\)depends onXXonly throughconfig⁡\(X\)\\operatorname\{config\}\(X\)\. In every tested instance the result was constant on the entire fibreconfign,ℓ−1⁡\(g0\)\\operatorname\{config\}^\{\-1\}\_\{n,\\ell\}\(g\_\{0\}\)and agreed exactly with an independent configuration\-lattice walk implementation \(Supplementary Table[5](https://arxiv.org/html/2609.30296#S7.T5)\)\. Blocklengths are necessarily small, the largest cases being45=10244^\{5\}=1024and84=40968^\{4\}=4096matrices; allv≠0v\\neq 0were tested and all agree, withv=e1v=e\_\{1\}shown for legibility\.

Supplementary Table 5:The operatorAvm​ΛA\_\{v\}^\{m\}\\Lambdacomputed literally on𝔽2ℓ×n\\mathbb\{F\}\_\{2\}^\{\\ell\\times n\}, against the configuration\-lattice walk count\. “Matrix values” is the set of values taken byAvm​Λ​\(X\)A\_\{v\}^\{m\}\\Lambda\(X\)over the entire fibreconfign,ℓ−1⁡\(g0\)\\operatorname\{config\}^\{\-1\}\_\{n,\\ell\}\(g\_\{0\}\); that this set is a singleton in every row is precisely the assertion that the quantity depends onXXonly throughconfig⁡\(X\)\\operatorname\{config\}\(X\)\. All entries are exact integers\.The finite\-nnwalk count was compared with the leading termnm​Wm​\(G,v\)n^\{m\}W\_\{m\}\(G,v\)of \([120](https://arxiv.org/html/2609.30296#S7.E120)\)\. The ratios converge to11, and do so*from above*, so the finite\-nncorrection is positive:

ℓ=2,m=2:n=81632641281\.66671\.33331\.16671\.08331\.0417ℓ=3,m=2:n=122448961921\.75001\.37501\.18751\.09381\.0469\\begin\{array\}\[\]\{lccccc\}\\ell=2,\\ m=2:&n=8&16&32&64&128\\\\ &1\.6667&1\.3333&1\.1667&1\.0833&1\.0417\\\\\[3\.0pt\] \\ell=3,\\ m=2:&n=12&24&48&96&192\\\\ &1\.7500&1\.3750&1\.1875&1\.0938&1\.0469\\end\{array\}with the same behaviour atm=4m=4\(forℓ=3\\ell=3,3\.1067→1\.09983\.1067\\rightarrow 1\.0998over the same sequence\)\. The sign of this correction is not used anywhere in the argument, which requires only theo⁡\(nm\)o\(n^\{m\}\)bound, but it is recorded because the direction is not obvious a priori\.

The inequality of Lemma[14](https://arxiv.org/html/2609.30296#Thmtheorem14)was checked on480480exact rational cases, with largest observed ratio0\.49570\.4957\. The extremal value12\\tfrac\{1\}\{2\}is attained already atℓ=1\\ell=1,m=2m=2, where there is a single pair and a single admissible multiplicity, so the coefficient comparison reduces exactly to\(21\)/4\\binom\{2\}\{1\}/4\.

The entropy inequalities carry complete proofs, so the remaining checks are consistency tests rather than evidence\. In6060\-digit arithmetic, the closed forms of Lemma[16](https://arxiv.org/html/2609.30296#Thmtheorem16)agreed with numerical differentiation to10−5710^\{\-57\}; the entropy chain of Theorem[18](https://arxiv.org/html/2609.30296#Thmtheorem18)held on16001600random distributions acrossℓ=1,…,4\\ell=1,\\dots,4with minimum gap−5\.6×10−61\-5\.6\\times 10^\{\-61\}, i\.e\. zero to working precision; and the conditional\-affinity identity \([124](https://arxiv.org/html/2609.30296#S7.E124)\) held on800800further distributions with maximum deviation4\.6×10−614\.6\\times 10^\{\-61\}\.

## Supplementary Methods: sign\-uncertainty discovery and diagnostics

### Rigorous tripwires

Two lower bounds are evaluated at run time and used to invalidate the numerics rather than to prove anything\. Ford=1d=1,s=\+1s=\+1,

A\+​\(1\)≥12​\(1\+λBCK\)=0\.4107675​…,A\_\{\+\}\(1\)\\geq\\frac\{1\}\{2\(1\+\\lambda\_\{\\mathrm\{BCK\}\}\)\}=0\.4107675\\ldots,withλBCK=−minx⁡sin⁡x/x\\lambda\_\{\\mathrm\{BCK\}\}=\-\\min\_\{x\}\\sin x/x\[[1](https://arxiv.org/html/2609.30296#bib.bib26)\]\. More generally, Theorem 2\.1 of\[[3](https://arxiv.org/html/2609.30296#bib.bib12)\]gives

u0≥λ⁡\(d,n\),u\_\{0\}\\geq\\lambda\(d,n\),whereλ⁡\(d,n\)\\lambda\(d,n\)is the smallest root ofL⌊n/2⌋\+1\(d/2−1\)L^\{\(d/2\-1\)\}\_\{\\lfloor n/2\\rfloor\+1\}\. The latter follows from Gauss quadrature againste−u​ud/2−1​d​ue^\{\-u\}u^\{d/2\-1\}\\,du: positivity of the quadrature weights forces a sign change at or beyond the first node\. Any run reporting feasibility below either applicable bound is therefore rejected\. During development, an early grid\-only relaxation returnedρ=0\.398942=1/2​π\\rho=0\.398942=1/\\sqrt\{2\\pi\}ind=1d=1, immediately falsified by the BCK tripwire\.

### Discovery campaign and numerical diagnosis

The initial search deliberately explored exact Fourier\-eigenfunction mixtures with trainable Gaussian scales and contact locations, using local optimization and Adam\-based learned proposals\. Low\-rank constructions reproduced the expected Bourgain–Clozel–Kahane controls and rapidly improved the objective, but thed=1d=1search plateaued nearρ=0\.5785\\rho=0\.5785, above the published Cohn–Gonçalves value\.

The failure was diagnostic\. When scales crowded or the imposed contact count was incompatible with the local root structure, the coefficient solve became extremely ill\-conditioned\. Candidate Adam trajectories that appeared to improve the objective entered systems with condition numbers of order101710^\{17\}–101810^\{18\}\. At this scale, forward accuracy could not be guaranteed in the floating\-point formats used here\. Multiprecision reevaluation showed that the apparent descent was dominated by numerical error rather than a reproducible improvement\. Learned widths and contacts were therefore retained only as proposals, with a hard conditioning guard before verification\.

The experiments also showed that the number of double contacts should not be rigidly tied to the number of basis functions\. Decoupling the two made comparisons across basis sizes meaningful and showed that additional widths at fixed contact count contributed little\. Multistart runs converged to distinct plateaus near0\.57850\.5785,0\.59290\.5929and higher values, and verified evaluations could jump when the outermost root pattern changed\. This behaviour is consistent with the discontinuity mechanism reported by Cohn and Gonçalves for Newton continuation atd≤2d\\leq 2\. We use this only as a diagnostic, not as an optimality theorem\.

### Implementation details of the convex exchange solver

Three implementation choices proved load\-bearing\. First, bare LP feasibility with zero objective returns arbitrary vertices that hug zero at many grid points and can dip between them\. Maximizing a relative margin instead produces strict positivity between most nodes and sharply reduces the number of exchange rounds\. Second, column scaling by the supremum over the full grid is dominated by the far tail and can exclude genuine feasible solutions; rows are therefore equilibrated while columns are left unscaled\. Third, multiple far tail anchors are nearly parallel after normalization and degrade the simplex\. The tail is imposed instead through a single structural sign constraint on the leading coefficient\.

At each bisection valueu0u\_\{0\}, the LP is solved on the current grid, the candidate derivativeP′P^\{\\prime\}is evaluated at high precision, and all real critical points on the active interval are located\. Any negative critical value is inserted as a new cutting row\. The process stops only after the candidate passes the independent continuous scan to the prescribed numerical tolerance\.

### Convergence status of the exchange loop

The oracle is Remez\-like in mechanics but not in guarantee\. The truncated Laguerre spaceVNV\_\{N\}is not a Haar system on\(0,∞\)\(0,\\infty\):L2​N\(α\)​\(2​u\)L^\{\(\\alpha\)\}\_\{2N\}\(2u\)belongs toVNV\_\{N\}and has2​N2Npositive roots, exceeding the Haar budget of an\(N\+1\)\(N\+1\)\-dimensional space\. Consequently, uniqueness of the contact configuration and classical equioscillation arguments are not available\.

The relevant framework is convex semi\-infinite programming\. Every finite grid gives an outer relaxation\. Under compactness and a Slater point, the exchange sequence drives the maximum constraint violation to zero and the corresponding margins converge to the semi\-infinite value\. Together with monotonicity and continuity inu0u\_\{0\}, this gives a numerical bracket for the fixed\-degree boundary\. This concerns optimality within the chosen finite space only\. Validity of every quoted upper bound is established by the exact rational certificate and does not depend on convergence of the exchange loop\.

### Contact degeneracy and leave\-one\-out reconstruction

The active contact count is discovered at the feasibility boundary\. The certifiedd=1d=1degree\-2222,3030and3838constructions usem=5m=5,77and99double contacts, respectively\. The non\-Haar structure can make the active set non\-unique\. At degree3838, the exchange oracle reports ten near\-active contacts although the square reconstruction uses nine\. Collocating all ten produces an admissible but irrelevant function whose outer sign change moves tou=274u=274\(ρ=9\.34\\rho=9\.34\)\.

We therefore perform leave\-one\-out reconstruction: each candidate contact is removed in turn, the remaining contacts are collocated in the square Laguerre system, and only reconstructions whose exact certificate agrees with the convex estimate are retained\. At degree3838, four of the ten subsets certify the same0\.57258870\.5725887value to nine digits\. The same repair is required atd=2d=2\. The multiplicity of successful subsets is consistent with a degenerate optimal facet and is not interpreted as uniqueness of the extremizer\.

### Exact Sturm implementation

After rational reconstruction,

P⁡\(u\)=∏i\(u−τi\)2​R​\(u\),P\(u\)=\\prod\_\{i\}\(u\-\\tau\_\{i\}\)^\{2\}R\(u\),and only the sign ofRRon the ray remains to be certified\. We build a primitive pseudo\-remainder sequence overℤ\\mathbb\{Z\}, using only positive rescalings so that Sturm sign\-variation counts are preserved\. This controls coefficient growth substantially better than naive rational Euclidean remainders, whose bit lengths become impractical by degree3838\.

The verifier checks the exact Laguerre\-to\-monomial expansion, zero remainder after division by every squared contact factor, the sign of the leading coefficient, the exact valueR⁡\(u¯\)\>0R\(\\bar\{u\}\)\>0, and equality of Sturm sign\-variation counts atu¯\\bar\{u\}and\+∞\+\\infty\. No grid, floating\-point root finder or interval sampling appears in the final admissibility decision\.

### Relation to the non\-Gaussian search

A numerical A/B analysis explains why the earlier Gaussian\-mixture search did not improve the final polynomial construction at smalldd\. Both spaces are evaluated through the same convex machinery, so the comparison concerns function spaces rather than optimizer quality\. Projecting individual mixture directions onto the degree\-2222polynomial space over the active region gives relative sup\-norm residuals of approximately

1\.3×10−15​\(a=1\.5\),3\.0×10−10​\(a=2\.5\),9\.2×10−7​\(a=4\),1\.3\\times 10^\{\-15\}\\ \(a=1\.5\),\\quad 3\.0\\times 10^\{\-10\}\\ \(a=2\.5\),\\quad 9\.2\\times 10^\{\-7\}\\ \(a=4\),6\.3×10−5​\(a=6\),1\.6×10−3​\(a=10\)\.6\.3\\times 10^\{\-5\}\\ \(a=6\),\\qquad 1\.6\\times 10^\{\-3\}\\ \(a=10\)\.Wider directions are genuinely distinct, but their optimized coefficients are of order10−310^\{\-3\}and did not improve the objective beyond the numerical noise floor\. This containment analysis is diagnostic only and is not used in the certificate\.

### Higher\-dimensional asymptotic diagnostic

Using the same convex machinery, the scale\-free quantityρ​2​π/d\\rho\\sqrt\{2\\pi/d\}along one degree–dimension path was

1\.435267​\(d=1\),1\.340341​\(d=2\),1\.212474​\(d=4\),1\.079183​\(d=8\),1\.026193​\(d=12\),1\.005480​\(d=16\)\.1\.435267\\ \(d\{=\}1\),\\quad 1\.340341\\ \(d\{=\}2\),\\quad 1\.212474\\ \(d\{=\}4\),\\quad 1\.079183\\ \(d\{=\}8\),\\quad 1\.026193\\ \(d\{=\}12\),\\quad 1\.005480\\ \(d\{=\}16\)\.The sequence approaches the sublinear\-degree asymptote11rather than crossing it\. This is only a diagnostic path, not an asymptotic theorem\. Atd=8d=8, increasing degree from1414to3838changes the reported value non\-monotonically over a range of approximately5\.6×10−45\.6\\times 10^\{\-4\}even though the exact fixed\-degree optima are nested\. We therefore interpret the spread as a numerical resolution limit of the discovery oracle\.

### Degree ceiling of the numerical stage

The current floating\-point exchange solver is reliable only to polynomial degree approximately3838\. Beyond this, the raw polynomial values span an extreme dynamic range: at degree5454,PPreaches about9\.7×10739\.7\\times 10^\{73\}at the far end of its own grid\. Row equilibration cannot repair the resulting column\-direction scaling, and the LP begins returning points that violate its own constraints\.

Near the ceiling, the finite grid can also acquire a recession direction that is non\-negative on all nodes while dipping between them, producing an “unbounded” rather than infeasible LP, and the cut pool can accumulate near\-duplicate rows across the bisection\. We cap the margin variable, refresh the cut pool when necessary, and discard runs whose bisection collides with a rigorous lower tripwire\. Multiplication by the positive factore−ue^\{\-u\}reduces the dynamic range but relocates the failure to underflow in the far tail\. Higher\-dimensional asymptotic work will therefore require higher precision or iterative/exact refinement inside the optimizer\. The certificates reported here are unaffected because all lie at degree at most3838and are validated by the exact rational stage\.

## 8NeuralCert structure

NeuralCert is organized around explicit interfaces and data boundaries rather than a single shared numerical backend\. A problem\-specific discovery module produces an explicit candidate representation together with the metadata required to interpret it\. A separate certification module reconstructs and evaluates that candidate using rigorous arithmetic and emits a self\-contained certificate\. An independent verifier consumes only this certificate and recomputes the claimed result without relying on the discovery or certification implementation\. Discovery methods can therefore be modified or replaced without changing the trust assumptions underlying certificates that have already been issued\. The three applications considered here share this discovery–certification–verification architecture, but not their numerical representations or proof mechanisms\. For the Maynard problem, finite\-kkpolynomial candidates are certified through exact rational evaluation of the associated Gram quadratic forms, whereas large\-kkrational candidates are certified through Fourier\-domain evaluation in Arb ball arithmetic\. Delsarte bounds are certified by exact rational verification of dual feasibility, while sign\-uncertainty candidates are reduced to rational polynomial factorizations and verified using exact Sturm\-sequence arguments\.

```
NeuralCert repository
|-- neuralcert/                  reusable framework
|   |-- core/                   problem and result contracts
|   |-- discovery/              neural models and optimization
|   |-- distill/                structural compression
|   |-- refine/                 deterministic local refinement
|   |-- exact/                  reusable exact arithmetic
|   |-- verify/                 independent verification primitives
|   |-- data/                   reproducible datasets
|   |-- pipeline.py             typed stage orchestration
|   ‘-- problems/
|       |-- maynard/            Maynard plugin and adapters
|       |-- delsarte/           coding-theory LP plugin
|       ‘-- sign_uncertainty/   Gaussian/Laguerre/Sturm plugin
‘-- maynard_tools/              specialized compatibility layer
    |-- discovery/              factored, gated, ratio
    |-- certification/          Karatsuba, CRT, FLINT, Arb
    ‘-- verifier/               independent Maynard checks
```

Supplementary Figure 4:Compact directory view of the NeuralCert package architecture\.
## 9Computational performance

The variational problem shards along the pair axis: for a pair\(j,ℓ\)\(j,\\ell\)the entire convolution chain touches only channelsjjandℓ\\elland produces two scalars, so the compute\-to\-communication ratio permits distribution over commodity interconnect\. Because the Hellmann–Feynman route detaches the Ritz vector, both quadratic forms are plain sums over pairs, and the exact gradient follows from two scalar all\-reductions and one gradient all\-reduction per iteration; no autograd\-aware collectives are required\. Memory is dominated by activations indexed by pair, growing asm⁡\(m\+1\)/2m\(m\+1\)/2, so parameter\-sharding schemes \(FSDP, ZeRO\) address the wrong resource; DDP is also inapplicable, since it averages gradients while the reduction here is a sum inside a nonlinear ratio\. Memory is not dominated by the assembled Gram matrices \(m×mm\\times m\), but by the activations of the convolution chain, which are indexed by pair and scale asN⋅nq⋅m⁡\(m\+1\)/2N\\cdot n\_\{q\}\\cdot m\(m\+1\)/2per chain step\. Because Hellmann–Feynman requires these intermediates for the backward pass, they are retained across all∼2​log2​k\\sim 2\\log\_\{2\}ksteps of the doubling chain\. Supplementary Table[6](https://arxiv.org/html/2609.30296#S9.T6)shows GPU memory use with different settings for the Maynard\-Tao polynomial fitting discovery procedure\.

Certification cost:The CRT backend at k=3000, degree 60 needs 155,000 primes; ball arithmetic replaces it with one adaptive\-precision evaluation\. The resolution requirement as a scaling law\.nr​e​pn\_\{rep\}grows roughly as≈k0\.78\\approx k^\{0\.78\}to hold theg≡xg\\equiv xcontrol at10−310^\{\-3\}\.

Supplementary Table 6:Overview of memory use and runtime for different settings and options for the Maynard\-TaoPolynomial discoveryregime since other implementations in the package have seconds runtime while using MegaBytes\.Note\.Heavy compute analyses were performed using multiple NVIDIA H100 NVL or NVIDIA H200 SXM GPU cards with respectively 96 or 141 GB VRAM\. GPU\-consumption = GB VRAM at it’s peak\.Ng​r​i​dN\_\{grid\}= Representation nodesNN; The Chebyshev–Lobatto nodes per channel which is the resolution of a one\-dimensional function, and thekk\-dimensional integral is handled by the separable structure; memory scales asN⋅3​m​\(m\+1\)/2N\\cdot 3m\(m\+1\)/2\. This is the vanilla setup \(ϵ=0\\epsilon=0\)\. Runtime ink=100,Ng​r​i​d=3000,Channels=32k=100,N\_\{grid\}=3000,\\text\{Channels\}=32is 1 hour and 14 minutes for 4000 Adam iterations and 8 minutes for 20 steps LBFGS\.

## 10Anatomy of a certificate

A certificate is a compact, exactly specified record from which a claimed inequality can be independently recomputed\. It is expensive to produce and cheap to check: verification requires neither the discovery pipeline nor the hardware used to generate the candidate, and a reader who distrusts every stage of the numerical search can nonetheless confirm the stated bound by running a short standalone program\. The asymmetry is deliberate: generating a competitive trial function requires the full optimisation and certification stack, whereas checking the resulting inequality requires only the exact trial parameters and a few seconds of interval arithmetic on commodity hardware\.

This note makes that concrete by walking through one complete example atk=5k=5— the smallest case, chosen so that every artefact fits on the page with nothing elided\. All files shown are shipped verbatim indemo/and are the unmodified output of the two commands in Panels A and C\.

### Panel A — The claim

The object being certified is one explicit function, fixed in advance:

F⁡\(t1,…,t5\)=∏i=15g⁡\(ti\)​1ℛ5​\(t\),g⁡\(t\)=1c\+4​t,c=552637315007978860977902\.F\(t\_\{1\},\\dots,t\_\{5\}\)=\\prod\_\{i=1\}^\{5\}g\(t\_\{i\}\)\\,\\mathbf\{1\}\_\{\\mathcal\{R\}\_\{5\}\}\(t\),\\qquad g\(t\)=\\frac\{1\}\{c\+4t\},\\qquad c=\\frac\{552637315007\}\{978860977902\}\.\(125\)BecauseMkM\_\{k\}is a supremum over admissible trial functions,*any*suchFFyields a lower bound, and the certified claim is

M5≥1\.9717711764\.M\_\{5\}\\ \\geq\\ 1\.9717711764\.Three properties of this statement are worth making explicit before the mechanics, because they are what the certificate does and does not assert\.

1. 1\.It is a statement aboutFF, not about the optimiser\. The value ofccwas found by numerical search; the search is not part of the claim, and a different \(or worse\)ccwould give a different \(or worse\) valid bound\.
2. 2\.It is a lower bound and is deliberately not sharp\. The discovery stage reportedR=1\.9966686079R=1\.9966686079for this sameFF; directed rounding of the aliasing and truncation budgets costs about2\.5×10−22\.5\\times 10^\{\-2\}, and the certified number is the conservative end of that interval\.
3. 3\.It is weaker than the best published constructions atk=5k=5, and this is the expected behaviour of an honest lower\-bound machine\. The rank\-one rational family is not the optimal family here: Maynard’s explicit polynomial trial givesM5≥1417255/708216=2\.00116​…M\_\{5\}\\geq 1417255/708216=2\.00116\\ldots, and Bogaert’s Krylov computation gives2\.00714514442\.0071451444\. The gap matters mathematically, since under the Elliott–Halberstam conjecture the thresholdM5\>2M\_\{5\}\>2impliesDHL⁡\[5,2\]\\mathrm\{DHL\}\[5,2\]; the certified rank\-one value does not reach it, whereas the polynomial trial does\. Atk=5k=5the certificate is therefore a demonstration of the machinery rather than a source of number\-theoretic content\.

### Panel B — The two files

The pipeline writes two records, and the separation between them is the trust boundary\.

B1\. The discovery record \(k5\-eps0\-ratio\.npz\)\.

Produced by the search stage\. It carries the trial function in exact rational form and nothing else of mathematical weight; the floating\-point fields are provenance, not evidence\.

k:5

epsilon:0/1\(vanillafunctional\)

mu:\[1\]

c\_num,c\_den:\[’552637315007’\],\[’978860977902’\]

power:\[1\]g\(t\)=\(c\+nt\)^\{\-1\}

w\_num,w\_den:\[’1’\],\[’1’\]mixingweight\(rankone\)

R\_discovery:1\.99666861float,notcertified

ceiling:2\.01179739k/\(k\-1\)logk

canonical:k=5\|552637315007/978860977902^\-1\*1

sha256:dfbd528eb3900e18771468e642b8cbcde4b2ed9074a593701419a856e2c596e6

Thecanonicalstring is the complete mathematical content of the file: dimension, exact pole, power, weight\. Everything the certifier needs is recoverable from that one line, and thesha256binds the certificate below to this exact trial function\.

B2\. The certificate \(k5\-eps0\-ratio\.json\)\.

Produced by the certification stage and reproduced here in full\.

\{

"format":"maynard\-Mk\-certificate/2",

"trial\_function":\{

"g":"g\(t\)=1/\(c\+nt\),n=k\-1",

"form":"F=prod\_\{i=1\.\.k\}g\(t\_i\)onthesimplex",

"c\_exact":"552637315007/978860977902",

"u\_exact":"1",

"note":"Risinvariantunderscalingoftheweight;anyfixedweight

yieldsavalidlowerbound\."

\},

"parameters":\{

"P":8\.0,"prec":200,

"nnode":153,"M0":153,

"rmax":12,"npanels":80,

"A\_upper":16\.170030887466964,

"Theta\_ball":32\.34006177493393,

"theta\_far":120\.0,"block\_ratio":2\.0,

"tol\_bits":45,"assembly":"directed\-arb/2"

\},

"certified\_quantities":\{

"N\_enclosure":\{"lower":0\.0877189203290238,"upper":0\.0877189203290238\},

"D\_enclosure":\{"lower":0\.2196630290925637,"upper":0\.21966302909256372\}

\},

"error\_terms":\{

"N\_band":0\.0,"N\_tail":0\.0005786544597255379,"N\_alias":3\.51115912146848e\-07,

"D\_band":0\.0,"D\_tail":0\.0013051009573133488,"D\_alias":4\.99060886139429e\-07,

"P\_S\_gt":1\.2860941636237798e\-06,"y":7\.0

\},

"far\_field\_blocks":\[\],

"final\_arithmetic":\{

"N\_lower\_used":0\.08713991475338612,

"D\_upper\_used":0\.22096862911076323,

"R\_lower":1\.9717711763896169,

"R\_midpoint":1\.9966700971800757,

"ceiling":2\.0117973905426254

\},

"claim":\{

"k":5,

"statement":"M\_5≥\\geq1\.971771176",

"M\_k\_lower\_bound":1\.971771176,

"known\_upper\_bound\_ceiling":2\.0117973905426254,

"ceiling\_source":"M\_k<k/\(k\-1\)logk\(Polymath8b\)",

"primes\_criterion":"r\_k=ceil\(thetaM\_k/2\),theta=1/2\-eps

\(Bombieri\-Vinogradov;MaynardProp\.4\.2\)",

"primes\_implied":1

\},

"assembly\_note":"AllerrorboundsassembledinArbballarithmetic;floats

exitonlythroughprovablyoutward\-roundedconversions

\(nudge\-verifiedagainsttheball\)\.",

"source":\{

"npz":"\./k5\-eps0\-ratio\.npz",

"sha256":"dfbd528eb3900e18771468e642b8cbcde4b2ed9074a593701419a856e2c596e6",

"canonical":"k=5\|552637315007/978860977902^\-1\*1",

"R\_discovery":1\.9966686078849576,

"u\_exact":"1"

\},

"self\_hash\_sha256":"7f12647dfe9d7c9e02fc1f636eb8aec4ae05ac359675a0412cdbb551a328ee20"

\}

That the file fits on one page is itself the point: there is nowhere for an unstated assumption to hide\. The fields divide into three kinds, and Panel C explains why the distinction matters\.

### Panel C — Verification

C1\. The transcript\.

The certifier reads the discovery record and writes the certificate:

$neuralcertcertify\-\-methodratio\-\-npzk5\-eps0\-ratio\.npz\\

\-\-cert\-jsonk5\-eps0\-ratio\.json

trialfunction:k=5\|552637315007/978860977902^\-1\*1

sha256=dfbd528eb3900e18\.\.\.\(verified\)

k=5R\_discovery=1\.9966686079

A=16\.17Theta\_ball=32\.3401theta\_far=120nodes=307farblocks=0

aliasing:P\(S\>7\)<=1\.286e\-06\(Markov,r<=12\)

Nin\[0\.08771892032902380\+/\-6\.19e\-18\]err\(band,tail,alias\)=\(0\.00e\+00,5\.79e\-04,3\.51e\-07\)

Din\[0\.2196630290925637\+/\-2\.42e\-17\]err=\(0\.00e\+00,1\.31e\-03,4\.99e\-07\)

Rmidpoint=1\.9966700972\|mid\-discovery\|=1\.49e\-06

CERTIFIED\(directed\):M\_5≥\\geq1\.9717711764\[0s\]

=\>atleast1primesinfinitelyoften\(theta=1/2\-eps\)

C2\. What is recomputed, and what is ignored\.

Every quantity below is regenerated by the verifier fromc\_exactand the computational settings alone\. None of the stored values is read as input; they are recomputed and compared\.

QuantityReconstructed fromValueA=2​\(c\+n\)/cA=2\(c\+n\)/ccc16\.17003088746696416\.170030887466964Θball=2​A\\Theta\_\{\\mathrm\{ball\}\}=2AAA32\.3400617749339332\.34006177493393Δ​θ=2​π/P\\Delta\\theta=2\\pi/PP=8P=80\.78539816340\.7853981634⌊Θball/Δ​θ⌋\\lfloor\\Theta\_\{\\mathrm\{ball\}\}/\\Delta\\theta\\rfloorA,PA,P4141M0=max⁡\(41,nnear\)M\_\{0\}=\\max\(41,\\,n\_\{\\mathrm\{near\}\}\)above,nnear=153n\_\{\\mathrm\{near\}\}=153153153Band errorM0=nnearM\_\{0\}=n\_\{\\mathrm\{near\}\}*empty*NtailN\_\{\\mathrm\{tail\}\},DtailD\_\{\\mathrm\{tail\}\}Lemma \(E2\) withA,Δ​θ,M0A,\\Delta\\theta,M\_\{0\}5\.7865×10−45\.7865\\times 10^\{\-4\},1\.3051×10−31\.3051\\times 10^\{\-3\}ℙ⁡\(S\>7\)\\mathbb\{P\}\(S\>7\)certified moments,r≤12r\\leq 121\.286×10−61\.286\\times 10^\{\-6\}Rlow=k​Nlow/DupR\_\{\\mathrm\{low\}\}=kN\_\{\\mathrm\{low\}\}/D\_\{\\mathrm\{up\}\}above1\.97177117638961\.9717711763896
Two features of this table are worth reading off\. First, the far\-field band is empty becauseM0=nnear=153M\_\{0\}=n\_\{\\mathrm\{near\}\}=153exceeds⌊Θball/Δ​θ⌋=41\\lfloor\\Theta\_\{\\mathrm\{ball\}\}/\\Delta\\theta\\rfloor=41: the analytic tail begins exactly where the computed sum ends, which is the choice discussed in the truncation bound and which atk=5k=5is worth more than an order of magnitude in the final bound\. Second, the entire chain descends from the single rationalcc; the verifier never needs the neural stage, a GPU, or any stored enclosure\.

Test 1 shows that an inflated claim cannot survive\. Test 2 shows the converse and less obvious half: a stored intermediate cannot be used to strengthen a claim either, because the verifier does not read it\. Onlyc\_exact\(guarded by thesha256against the discovery record\) and the computational settings enter, and altering the latter changes only how sharp the recomputed bound is, never whether the recomputation is valid\.

C4\. Independent reimplementation\.

The strongest form of the check does not use our code at all\. Every entry in the Panel C2 table is an elementary formula incc,PP,nnearn\_\{\\mathrm\{near\}\}andrmaxr\_\{\\max\}, and can be reproduced in a few dozen lines with any interval\-arithmetic library\. We have verified the present example this way as well as through the shipped verifier\.

## 11Maynard–Tao Discovery and Certification Tables

Supplementary Table 7:Vanilla discovery and certification results \(ϵ=0\\epsilon=0\)\.kkDiscoveryC​hdiscCh\_\{\\mathrm\{disc\}\}C​hcertCh\_\{\\mathrm\{cert\}\}DegreeCertificationF​rceilingFr\_\{\\mathrm\{ceiling\}\}203\.1283459912293633603\.1275579504967310891992910,992213\.1701068372983630603\.1696877441350859087106440,992223\.2104485790103629603\.2097914252497743010296050,991233\.2498124502483630603\.2490181028734139228343340,991243\.2874775738483632603\.2862399940663731101680970,991253\.3225087270723628603\.3221511359607705603797340,991263\.3580331090703630603\.3565934089774853980513690,991273\.3907974124993630603\.3898823271870247626780300\.990283\.4235505013693629603\.4218911753676461213063160\.990293\.4539128995353633603\.4528797749549881048892540\.990303\.4845199057404235603\.4829208049083564991738980\.990313\.5141547415704235603\.5124054121982641664197450\.990323\.5429480972894234603\.5404780957101035355699370\.990333\.5694083036474235603\.5681680821408169220574090\.990343\.5958560148814236603\.5948622152059478311192580,989353\.6226371404604236603\.6207297846091151500231890,989363\.6490493769294236603\.6459394491332373147811110,989373\.6726490218374236603\.6707493581918008979536160,989383\.6974505141324238603\.6942126596176723354303140,989393\.7199688334324237603\.7180575159272732834884550,989403\.7434121447284845603\.7408501291898449293844970150460,989413\.7654781322494839603\.7633262482113631881952269786700,989423\.7864436408154840603\.7848363987108520901663630452630,989433\.8083933915264844603\.8061128993324498435137253489990,988443\.8310892495764841603\.8267342861145008425505873647870,988453\.8505489107934844603\.8475823968146918420889471019990,988463\.8707082530984845603\.8670622744607915700550535872040,988473\.8909815489564842603\.8870500397121554654582577398110,988483\.9116069764914840603\.9061509109048129348474255590320,988493\.9276171398014842603\.9252181779466287140630385562460,988503\.9437072178584033503\.9424459717861973445587348622540,988513\.9622011563424034503\.9617429099431768823844405776150,988523\.9800859208964033503\.9793617433528050487798878272290,988533\.9973920904654035503\.9965695466035207921891538650450,988544\.0144173494484033504\.0134213927283374686858808633640,987554\.0311580930474033504\.0299303158833736616480751715870,987564\.0487401961544032504\.0472111834986403147709856561140,987574\.0646759432914033504\.0636589112659303660158970900050,987584\.0808972416504033504\.0796625398037689496587348687620,987594\.0961076757234033504\.0949290275206735228212348720440,987604\.1112242212714032504\.1102446799000463072365172320520,987614\.1231804138894037504\.1225933968925705784875927048360,986624\.1414158068454035504\.1404212714403973882990911513430,987634\.1514528526154037504\.1505837322661636233224240421140,986644\.1708436654004033504\.1697330352464908567219162977300,987654\.1851179583924036504\.1842004562910097109389691008240,987664\.1992581110854034504\.1982123526310114232706515437780,987674\.2128013006504032504\.2117343752599384674624916562280,987684\.2269137073474034504\.2256715372322035729142982510050,987694\.2406410549704027504\.2395308434032994718068598416860,987704\.2520430445244034504\.2507095954584235092742719037500,986714\.2620772832754034504\.2607814711800625276128108206000,985724\.2801706714934031504\.2784866594830717405245074666770,987734\.2916523310794036504\.2904399063779092962842954836810,986744\.3020395674184032504\.3007666579628258795835712736850,986754\.3150372906534036504\.3138198041639895944089826842600,986764\.3308390057454033504\.3288845634609768174567081079300,986774\.3425885811014032504\.3409031395632538399776950487580,986784\.3534906586604030504\.3514500265253147033937293017220,986794\.3663320322204025504\.3635470694673143529357616328100,986804\.3758706553635036504\.3753055758751952807968099506410,986814\.3883910842615034504\.3878228366326321712723030864050,986824\.3947748707015040504\.3943269018443651762524670347600,985834\.4018792019375043504\.4014726715183735089125452346140,984844\.4173737050405040504\.4169530949497574114309145069200,985854\.3464418140155041504\.3464944883796581855005634605630,967864\.4311051488495042504\.4306269694812805841462147240730,983874\.4529089611955040504\.4519133969906252931276985594510,985884\.4643029796125041504\.4634261572577084173553753159390,986894\.4723550815935042504\.4719885309808107548262277703130,985904\.4871193717755037504\.4859107342109005510452805456260,986914\.4966985275925037504\.4956450935326295105099394719850,986924\.5058296930795039504\.5045477059766281504654315740440,985934\.5183473399545036504\.5163392083779479490804433220710,986944\.5280590404705029504\.5266359371404596144069720860930,986954\.5338067721775041504\.5324362259273725848396141179230,985964\.5454040660275026504\.5440842211524060684838767432170,985974\.5543206036745031504\.5535888074591291652535742033200,985984\.5652789478505028504\.5640045279397188019841564604970,985994\.5093544188615035504\.5085615728629630882248360654940,9711004\.5853893711155039504\.5835591691203798611859473894710,9851004\.5826365291145039504\.5819939856703818525095531804910,9851304\.8223415201155013504\.8087138241257188362877343687170\.9801504\.961245663975506504\.9553667181829810770612007980910\.9822005\.090813348651504605\.0872806884052788041646395368170\.9552205\.2583894424175012605\.2479977864157325913672142563220\.9692505\.384166994782505605\.3499290877054556379959519389080\.9653505\.610666510192503604\.9217794424830776058849435452590\.8385006\.080822857000503605\.0326552408588428487574972073580\.808Note\.Valuesk=25k=25andk\>30k\>30have certified values above Bogaert’s Krylov subspace method and the published numbers on the Polymath8b Wiki\.C​hdiscCh\_\{\\mathrm\{disc\}\}; amount of channels in the discovery,C​hcertCh\_\{\\mathrm\{cert\}\}; amount of channels in the certification \(after pruning\)\.F​rceilingFr\_\{\\mathrm\{ceiling\}\}; the gap to the Cauchy Schwartz ceiling\.Supplementary Table 8:Discovery and certification results forε=1/25\\varepsilon=1/25\.kkDiscoveryC​hdiscCh\_\{\\mathrm\{disc\}\}C​hcertCh\_\{\\mathrm\{cert\}\}DegreeCertification203\.2109360697863633603\.210366921521749737183691731395213\.2526803356443630603\.252298834127242084124289323548223\.2927792667223629603\.292126369692432458798975243945233\.3311734421933630603\.330322094752801530355410455276243\.3680367706353630603\.366908489372213299919519038109253\.4023713657543629603\.402010619443212328887383163464263\.4370280986363631603\.436157517278351176844746714320273\.4688270242443629603\.468320542253183831988945583237283\.5004158764523630603\.499444954997692030882286887925293\.5311865165023632603\.530010877716944327578014764437303\.5617001793404239603\.559458746545551314730876537135313\.5906897449724240603\.588557568639265632860385653635323\.6187329638794238603\.616127200711337138001580976885333\.6442532624454236603\.642965802069061441160922851345343\.6714505994074239603\.668624924763460294356636286090353\.6962500847824240603\.694175287948157320412197179779363\.7209085180934239603\.718279485810969785557788969075373\.7442734266464238603\.742226163630752703816176588194383\.7681349693684241603\.766021610226650587509688536386393\.7902137306784239603\.788245621497978381906985116415403\.8135217642344845603\.810310054297441303862285913738413\.8336166279614841603\.831637106843201917337701123609423\.8559066284024845603\.852600769154481058291704449249433\.8767535500314842603\.873483797869255723761005318927443\.8955779138754846603\.893301604331862091582323108956453\.9156497330224845603\.913474652772154201185956494902463\.9369371391394843603\.932552485373460445967272733645473\.9567030332424846603\.952079215054289455935133148337483\.9758235961684844603\.970095831528526602285865965250493\.9914507309294843603\.988676716542276674807876834435504\.0060263863274032504\.005637471340251334822223632678514\.0213379132834021504\.020552375066701497376356751144524\.0406351192784031504\.040258681163316069160504782304534\.0582155515464024504\.057203993845967521227902017906544\.0729692250604028504\.072322886063011063853846972776554\.0848022028824018504\.084085711139175414397956372460564\.1000287836794026504\.099206855766737568002165290524574\.1155858447644023504\.115039486353743806900256901164584\.1351521913124025504\.134242300759367593439951066209594\.1491878951984030504\.148472609734504851305763152187604\.1589669326004019504\.157998757293132279032736615835614\.1759755078794031504\.175077906373857205061979626025624\.1898627349734032504\.188904070211809938022792508286634\.2105163367264033504\.208857616045980796408818491228644\.2194151614374030504\.218124437987688545452024305223654\.2385018006154036504\.237384124837738623814213247461664\.2521205452984029504\.250793037094066225539732443814674\.2614027486084022504\.259777456025320664617972564576684\.2703796189424026504\.269124177371744769988813084968694\.283984925481405504\.277274261946282382619508640328704\.3015601173064034504\.298733418769167808786843497725714\.3168850832424031504\.314413443106201423239218270654724\.3295033346054035504\.327034144580560001136968528718734\.3349089747724030504\.332768096731747934119777486576744\.3420399958484028504\.339799356192207942523849650386754\.3589172467934032504\.356375192119249086787668845356764\.3759022551184031504\.372974259197422394873126180262774\.3880926894564035504\.385396831879956294549952588437784\.3991845357124022504\.395990292345199428975680930752794\.4072144162414030504\.404183984805812816254054716642804\.1959419887675040504\.197608268366179620141468586531814\.0751173662645035504\.093066546579362186151057435374824\.1929786893285043504\.195259401348774146934234241013834\.4256442011765035504\.424852131550489892382995112584844\.0968254734755031504\.099802658570555784696980423382854\.4074752056275039504\.406760295616284828171584604987864\.4570725091215031504\.458412761836136236841801895336874\.4532652659895045504\.453060114331590189551387134208884\.3444345719085040504\.344422869750479566119971669552894\.4105388547015039504\.417151906430070350388796841891904\.4707454637235045504\.471149582958362588083170625461914\.1751218412425041504\.176550564513205748419816639538924\.3818721851975045504\.389037025811565058867953246127934\.3129175702695044504\.314963190359444288509182754059944\.2726521511095038504\.273951035321736008358360794466954\.4202172559295038504\.419415306457317943090466947241964\.3116148040445034504\.311428453792439589002649755954974\.2036130027675035504\.206802724014883123574465837334984\.2171145510525039504\.225342000024770175903490961201994\.5482988523845039504\.5469231906827621839444643948531004\.4965360058025042504\.4955121338105546633247784952901004\.4669836327655040504\.4668025228129856391582551857701304\.8479865621515010504\.8441564802877458119118149596501504\.977475423203508504\.9619473343636139668668017987772005\.094346768152507605\.0873936900448346745778158853682204\.978911765299508604\.9767644895570454819524993191422204\.978911765299508604\.9767644895570454819524993191423004\.086884354936501603\.6899639558542960356160074271844005\.3819485549005013605\.3429887493843686407474052600755004\.5121152070755014604\.192789933536702947199563856796Note\.C​hdiscCh\_\{\\mathrm\{disc\}\}; amount of channels in the discovery,C​hcertCh\_\{\\mathrm\{cert\}\}; amount of channels in the certification \(after pruning\)\.Supplementary Table 9:Certified enlarged\-support sweep underlying theε\\varepsilon\-crossover analysis\.For eachkk, the table reports the certified lower bound obtained for the corresponding enlarged\-support parameterε\\varepsilon, together with its within\-kkranking among the tested values\. Rank11denotes the strongest certified value at thatkk\.kkε\\varepsilonCertified valueRank50003\.9421277424904536046637619855585501/51/53\.7125092395459212043175290713688501/61/63\.7645049913245597029994669327107501/81/83\.8743841801534644147104683985246501/121/123\.9590889834063599216454334830594501/161/163\.9887184381225635633925228852523501/251/253\.9979977403376278440755838845621501/501/503\.994094475900581244335373138334260004\.1078676901678660553014641995754601/51/53\.8231383589912181458796331899228601/61/63\.9011089731246428231544208241997601/81/84\.0086384291881180376062146855176601/121/124\.0789193122205895741361957302015601/161/164\.1374294062600095166322800911563601/251/254\.1540302102680539928862637605132601/501/504\.159671531351647672765832327114170004\.2507689960134570356176703378654701/51/53\.9328746217758662657761043920998701/61/63\.9924676057628459199875633827017701/81/84\.0971810322388706816232436654876701/121/124\.2038733499670480266463588486025701/161/164\.2647515582303815636551237548343701/251/254\.2841819547257592297688828249612701/501/504\.298346167280743718826612134536180004\.3650762866895428640358280779544801/51/54\.0107395107914323896104990561948801/61/64\.0467135734564691142357276374377801/81/84\.1984339110294657566119695010536801/121/124\.3007152184255789089131336645435801/161/164\.3688468404743456303808422043953801/251/254\.3976416332296763537599315577602801/501/504\.409932699888494327027154402066190004\.4753941789306063003029300283153901/51/54\.0511934592934467040085040245507901/61/64\.1166207103671561188538734139866901/81/84\.0261338610588922363588806678828901/121/124\.2776407421114849134504944852565901/161/164\.4263009477437443446073854005484901/251/254\.5021083221224457179670067687422901/501/504\.5148126684992810343775377005731100004\.57186516176605592097795734251621001/51/54\.13117496620231909457325208692981001/61/64\.22866193022130054799028298021671001/81/84\.37616419838276327228227987243961001/121/124\.45574710958496614332161561742751001/161/164\.48307260431631638152306182410541001/251/254\.52059276732582497695574371960731001/501/504\.6003982521597685120248061287651

Similar Articles