AI智能体能否重新发现Blaschke曲线的不变量?

arXiv cs.AI 论文

摘要

一篇研讨会论文提出了一个单实例案例研究:让AI智能体从数值数据中重新发现四次Blaschke曲线的一个隐藏不变量。智能体成功导出了一个齐次三次式,能够以机器精度的残差预测未见过的构型;论文也诚实地指出其局限,例如一个确定性的多项式拟合基线也能得到与智能体相同的结果。

arXiv:2609.38369v1 Announce Type: new Abstract: We study generalized Blaschke curves as a controlled environment for AI-assisted mathematical rediscovery. For one fixed degree-four Blaschke product, an agent receives numerical coordinates of the six pair-lines determined by each of 80 boundary configurations. The target theorem is withheld from the task instructions. The saved research log reports rejected geometric hypotheses and a homogeneous cubic fitted to polygon sides. Its frozen coefficients predict 480 lines from 80 unseen parameter values, with a recorded RMS scale-free residual of $8.88\times10^{-17}$. Discovery-set diagonals provide an out-of-fit consistency check, not a fully held-out test. A separate one-configuration run reports insufficient evidence for invariance. A post-review deterministic degree-search baseline also recovers the cubic, so the experiment does not establish an advantage over polynomial fitting. We present this single-instance case study as a protocol for separating conjecture, numerical validation, and proof, with explicit limitations concerning agent metadata, prior knowledge, and reproducibility.
查看原文
查看缓存全文

缓存时间: 2026/10/01 09:40

# Can an AI Agent Rediscover a Blaschke-Curve Invariant?
Source: [https://arxiv.org/html/2609.38369](https://arxiv.org/html/2609.38369)
\\workshoptitle

The 6th Workshop on Mathematical Reasoning and AI

###### Abstract

We study generalized Blaschke curves as a controlled environment for AI\-assisted mathematical rediscovery\. For one fixed degree\-four Blaschke product, an agent receives numerical coordinates of the six pair\-lines determined by each of 80 boundary configurations\. The target theorem is withheld from the task instructions\. The saved research log reports rejected geometric hypotheses and a homogeneous cubic fitted to polygon sides\. Its frozen coefficients predict 480 lines from 80 unseen parameter values, with a recorded RMS scale\-free residual of8\.88×10−178\.88\\times 10^\{\-17\}\. Discovery\-set diagonals provide an out\-of\-fit consistency check, not a fully held\-out test\. A separate one\-configuration run reports insufficient evidence for invariance\. A post\-review deterministic degree\-search baseline also recovers the cubic, so the experiment does not establish an advantage over polynomial fitting\. We present this single\-instance case study as a protocol for separating conjecture, numerical validation, and proof, with explicit limitations concerning agent metadata, prior knowledge, and reproducibility\.

## 1Introduction

An experimentally suggested equation is not yet a theorem, and a successful numerical fit is not by itself evidence of an autonomous discovery process\. These distinctions motivate a small, inspectable test case in Blaschke geometry\. We ask whether a mathematical agent can formulate a parameter\-independent relation from numerical observations and commit to that relation before testing new configurations\.

Scientific rediscovery from data has an established history\. For example, AI Feynman evaluates symbolic regression on known physical equations\([Udrescu and Tegmark, 2020](https://arxiv.org/html/2609.38369#bib.bib8)\)\. FunSearch and AlphaGeometry illustrate different combinations of learned search, mathematical structure, and verification\([Romera\-Paredes et al\., 2024](https://arxiv.org/html/2609.38369#bib.bib6);[Trinh et al\., 2024](https://arxiv.org/html/2609.38369#bib.bib7)\)\. Our contribution is narrower: a geometric case study with an explicit information boundary and a frozen numerical prediction\. We do not introduce symbolic regression, prove a new Blaschke theorem, or estimate an agent success rate\.

Finite Blaschke products connect complex analysis, Poncelet geometry, and numerical ranges\([Daepp et al\., 2019](https://arxiv.org/html/2609.38369#bib.bib1);[Mirman and Shukla, 2005](https://arxiv.org/html/2609.38369#bib.bib4)\)\. In the normalized degree\-three case, the roots of a boundary level equation form triangles tangent to a fixed ellipse\. Higher degrees lead to generalized algebraic envelopes\([Siebeck, 1864](https://arxiv.org/html/2609.38369#bib.bib5);[Linfield, 1920](https://arxiv.org/html/2609.38369#bib.bib3);[Hunziker et al\., 2022](https://arxiv.org/html/2609.38369#bib.bib2)\)\. This gives a rediscovery target that can be hidden from task instructions while remaining mathematically grounded\. The resulting experiment separates three questions: whether a candidate equation fits, whether it predicts new parameter values, and what the available record establishes about the agent’s process\.

## 2A geometric discovery environment

Write𝕋=\{z∈ℂ:\|z\|=1\}\\mathbb\{T\}=\\\{z\\in\\mathbb\{C\}:\|z\|=1\\\}and fix

B⁡\(z\)=z​∏j=13z−aj1−aj¯​z,\(a1,a2,a3\)=\(0\.22\+0\.18​i,−0\.31\+0\.09​i,0\.08−0\.38​i\)\.B\(z\)=z\\prod\_\{j=1\}^\{3\}\\frac\{z\-a\_\{j\}\}\{1\-\\overline\{a\_\{j\}\}z\},\\qquad\(a\_\{1\},a\_\{2\},a\_\{3\}\)=\(0\.22\+0\.18i,\-0\.31\+0\.09i,0\.08\-0\.38i\)\.\(1\)For eachλ∈𝕋\\lambda\\in\\mathbb\{T\}, the four roots ofB⁡\(z\)=λB\(z\)=\\lambdalie on𝕋\\mathbb\{T\}\. In cyclic order, they determine four sides and two diagonals\. We encode each pair\-line by\[u:v:w\]\[u:v:w\], whereu​x\+v​y\+w=0ux\+vy\+w=0, usingu2\+v2=1u^\{2\}\+v^\{2\}=1and a deterministic sign\. A point in this dual projective plane represents a line in the original plane\.

The known geometric structure is a fixed cubic relation among these line coordinates\. Equivalently, the pair\-lines are tangent to a parameter\-independent algebraic envelope of class three; the class refers to the degree of its dual curve, not necessarily its degree in the original plane\. The experiment seeks numerical evidence for this relation without supplying the theorem or its expected degree\. It does not seek the zeros from raw observations: both the representation and the data\-generation problem were chosen by the experimenter\.

![Refer to caption](https://arxiv.org/html/2609.38369v1/discovery_80_lambda.png)Figure 1:The 480 stored discovery lines from exactly 80 values ofλ\\lambda\. Blue lines are sides and red lines are diagonals; black points are the zeros ofBB\. This post\-review visualization uses the experimental discovery CSV, not an additional sampling grid\. No envelope figure was supplied in the original discovery task\.#### Data construction\.

Forj=0,…,159j=0,\\ldots,159, letθj=2​π​\(j\+0\.173\)/160\\theta\_\{j\}=2\\pi\(j\+0\.173\)/160\. Even indices provide discovery data and odd indices provide evaluation data\. Each split contains 80 configurations and 480 lines\. We solve

z​∏j=13\(z−aj\)−ei​θ​∏j=13\(1−aj¯​z\)=0z\\prod\_\{j=1\}^\{3\}\(z\-a\_\{j\}\)\-e^\{i\\theta\}\\prod\_\{j=1\}^\{3\}\(1\-\\overline\{a\_\{j\}\}z\)=0numerically, project each root radially onto𝕋\\mathbb\{T\}, sort by argument, and compute the six pair\-lines\. The stored modulus diagnostic is measured*after*projection and is not an independent root\-accuracy estimate\. The split tests new values of the same parameter for the same product, not new products or extrapolation beyond the sampled circle\.

## 3Agent protocol and recorded trajectory

#### Information boundary and audit\.

The archived protocol describes a separate Codex sub\-agent receiving the discovery CSV and a task prompt\. Recovered local session records identify both agents asgpt\-5\.6\-solwithlowreasoning effort\. Both launches specifyfork\_turns: none, excluding a fork of the parent conversation\. The CSV includes parameter values, side/diagonal labels, and line coordinates; the prompt names degree\-four Blaschke products and asks for stable algebraic or geometric structure\. Numerical or symbolic calculations, fitting, and plotting are allowed\. The theorem, expected degree, held\-out file, other workspace sources, external lookup, and parent\-agent hints are prohibited during discovery\. File restrictions are instruction\-level: the recorded execution policy permits unrestricted filesystem access\.

The protocol states that the parent agent audited the written conjecture before releasing the held\-out file, with coefficients and thresholds then fixed\. The human experimenter selected the mathematical setting; the computational scaffold generated data and managed release\. The recovered parent record contains two experiment launches and an evaluation follow\-up for each\. Child records preserve tool calls and usage metadata, but a complete launch\-message and intervention audit remains unavailable\. Temperature, a sampling seed, a dated model snapshot, and a prespecified execution budget are not established\. We report two documented trajectories without claiming an exhaustive account\-wide search for attempts\. Appendix[A](https://arxiv.org/html/2609.38369#A1)gives recovered metadata and remaining limitations\.

#### Reported hypothesis development\.

The full\-data log first rejects a fixed intersection point for the two diagonals\. It then fits a homogeneous quadratic to the 320 side lines, obtaining an algebraic RMS residual of9\.25×10−39\.25\\times 10^\{\-3\}in its chosen coefficient scaling\. Splitting sides by the sign ofwwreduces but does not eliminate this error\. These checks reject those particular descriptions, not every conceivable conic decomposition\.

The log next reports a search over homogeneous polynomial degrees\. At degree three, the side\-only design matrix has smallest and next\-smallest singular values3\.21×10−153\.21\\times 10^\{\-15\}and6\.64×10−26\.64\\times 10^\{\-2\}\. Four interlaced discovery\-parameter folds yield small omitted\-fold residuals\. The 160 discovery diagonals also satisfy the side\-fitted cubic, with raw RMS residual8\.29×10−178\.29\\times 10^\{\-17\}\. Since those diagonals were visible from the outset, this is an*out\-of\-fit consistency check*, not a hidden\-label generalization test\.

#### Frozen numerical conjecture\.

The monomials are ordered as

\(u3,u2​v,u2​w,u​v2,u​v​w,u​w2,v3,v2​w,v​w2,w3\)\.\(u^\{3\},u^\{2\}v,u^\{2\}w,uv^\{2\},uvw,uw^\{2\},v^\{3\},v^\{2\}w,vw^\{2\},w^\{3\}\)\.The archived equation is

P⁡\(u,v,w\)=\\displaystyle P\(u,v,w\)=\{\}0\.0032683124​u3−0\.0494655836​u2​v\+0\.491767409872​u2​w\\displaystyle 0\.0032683124u^\{3\}\-0\.0494655836u^\{2\}v\+0\.491767409872u^\{2\}w−0\.0171636876​u​v2−0\.0198​u​v​w\+0\.01​u​w2−0\.0202735836​v3\\displaystyle\-0\.0171636876uv^\{2\}\-0\.0198uvw\+0\.01uw^\{2\}\-0\.0202735836v^\{3\}\+0\.502767409872​v2​w\+0\.11​v​w2−w3=0\.\\displaystyle\+0\.502767409872v^\{2\}w\+0\.11vw^\{2\}\-w^\{3\}=0\.\(2\)The agent conjectures this relation for every pair\-line of the fixed product\. Its prespecified test uses

r⁡\(u,v,w\)=\|P⁡\(u,v,w\)\|\(u2\+v2\+w2\)3/2,r\(u,v,w\)=\\frac\{\|P\(u,v,w\)\|\}\{\(u^\{2\}\+v^\{2\}\+w^\{2\}\)^\{3/2\}\},\(3\)requiring every evaluation row below10−1010^\{\-10\}and RMS below10−1110^\{\-11\}, both overall and separately for sides and diagonals\. This residual is invariant under rescaling the*line*coordinates\. Its value still depends on polynomial scaling, fixed here by the coefficient−1\-1ofw3w^\{3\}\.

## 4Results and deterministic controls

#### Original frozen evaluation\.

The archived report records 480/480 passing rows, overall RMS8\.88×10−178\.88\\times 10^\{\-17\}, and maximum3\.18×10−163\.18\\times 10^\{\-16\}\. Side and diagonal RMS values are8\.86×10−178\.86\\times 10^\{\-17\}and8\.94×10−178\.94\\times 10^\{\-17\}, respectively\. These are numerical predictions at previously unseen parameter values, not proof of an identity for allλ\\lambda\. A post\-review reevaluation of the stored coefficients and CSV gives RMS8\.85×10−178\.85\\times 10^\{\-17\}; the tiny difference is at floating\-point roundoff scale\. We retain the original recorded result rather than replace its history\.

#### One\-configuration ablation\.

The separate ablation retained*only one*discovery parameter value and its six lines; it did not remove just one configuration from the full set\. The saved log diagnoses inadequate evidence for parameter invariance, records unit\-circle and incidence relations, and produces no invariant cubic\. Its construction\-level check passes the 80 evaluation configurations but is not the target discovery\. With ten homogeneous cubic coefficients and at most six independent linear constraints, one configuration cannot uniquely determine a cubic up to scale from incidence data alone\. This dimension count does not rule out identification under additional prior assumptions\.

The ablation changes both the number of parameter values and the number of observations\. Its contrast with the successful trajectory therefore does not isolate a causal effect of parameter diversity, establish that every one\-configuration prompt fails, or quantify reliability across agents\.

#### Post\-review SVD baseline\.

We added a deterministic comparison after review; it was not part of the original agent protocol\. For each degreed=1,…,5d=1,\\ldots,5, we form the matrix of all degree\-ddmonomials evaluated at the same 320 discovery side rows\. We count singular values above10−12​σmax10^\{\-12\}\\sigma\_\{\\max\}and select the first degree with a one\-dimensional numerical nullspace\. Degrees one through five have nullities0,0,1,3,60,0,1,3,6\. The higher\-degree relations are consistent with multiples of the cubic\. No diagonal or evaluation row enters this degree selection or coefficient fit\.

Table 1:Original agent outcome and post\-review deterministic comparison\. Both polynomials use coefficient−1\-1forw3w^\{3\}and residual \([3](https://arxiv.org/html/2609.38369#S3.E3)\)\. Roundoff\-level differences are not evidence of superior discovery performance\.The baseline selects degree three and passes the same numerical thresholds \(Table[1](https://arxiv.org/html/2609.38369#S4.T1)\)\. After unit\-norm scaling and sign alignment, its coefficient vector differs from the frozen vector by6\.20×10−156\.20\\times 10^\{\-15\}in Euclidean norm\. Thus the experiment does not demonstrate that an agent is needed to recover the cubic from this representation\. Its process\-level interest lies in formulating and recording hypotheses and validation criteria, not in outperforming SVD; these aspects require stronger evaluation in a larger study\.

#### Parameter\-count diagnostic\.

As a further numerical control, we select one, two, five, ten, twenty, forty, or eighty evenly spread discovery configurations and repeat their side rows to give 320 rows in every condition\. The cubic nullities are6,2,1,1,1,1,16,2,1,1,1,1,1, respectively\. Repetition changes row count but not information or rank\. For each testedk≥5k\\geq 5, the fitted cubic predicts evaluation lines with unit\-coefficient RMS below2\.3×10−152\.3\\times 10^\{\-15\}\. These deterministic diagnostics show why replicated observations cannot repair insufficient independent constraints\. They are not repeated agent runs, a matched\-information experiment, or a claim that five configurations are minimal\. Appendix[B](https://arxiv.org/html/2609.38369#A2)specifies the computation\.

## 5Interpretation and limitations

The study supports a modest conclusion: a documented agent trajectory produced a numerical equation that predicts new configurations of one fixed Blaschke product\. Parameter\-separated evaluation avoids training and testing on different lines of the same configuration, and freezing the equation makes the numerical claim directly checkable\. Nevertheless, held\-out parameters alone do not establish novelty, proof, or a uniquely agentic mechanism\.

Withholding a theorem from the prompt is not equivalent to withholding it from pretraining\. The task explicitly names Blaschke products, and the agent could have relevant prior knowledge\. Research logs record reported hypotheses; they cannot certify faithful internal reasoning or rule out recalled results\. This distinction is supported by studies of explanation faithfulness\([Turpin et al\., 2023](https://arxiv.org/html/2609.38369#bib.bib9);[Chen et al\., 2025](https://arxiv.org/html/2609.38369#bib.bib10)\)\. We therefore use rediscovery in the operational sense of producing and testing a withheld target relation, not as a claim about the origin of the model’s knowledge\.

The principal limitations are a single product, one documented trajectory per condition, unrecorded sampling settings and execution budgets, and incomplete process auditing\. Numerical rank deficiency is not a symbolic proof\. The mathematically informed dual\-coordinate representation also makes deterministic recovery straightforward\. Future experiments should prerecord complete configurations and budgets, compare repeated runs against numerical baselines, hide diagonals entirely during discovery, vary products and representations, and distinguish symbolic verification from numerical accuracy\. The present case study is a starting point for such a workflow, not its completed validation\.

## Acknowledgments and Disclosure of Funding

The author has no grant funding to acknowledge and declares no competing interests\.

AI tools assisted the experiment, analysis, and preparation of this manuscript\. The author is responsible for checking the mathematics, reported evidence, citations, and final text\. The post\-review controls are identified separately from the archived agent runs\.

## References

- Chenet al\.\(2025\)Y\. Chen, J\. Benton, A\. Radhakrishnan, J\. Uesato, C\. Denison, J\. Schulman, A\. Somani, P\. Hase, M\. Wagner, F\. Roger, V\. Mikulik, S\. R\. Bowman, J\. Leike, J\. Kaplan, and E\. PerezReasoning models don’t always say what they think\.Note:arXiv:2505\.05410External Links:[Link](https://arxiv.org/abs/2505.05410)Cited by:[§5](https://arxiv.org/html/2609.38369#S5.p2.1)\.
- Daeppet al\.\(2019\)U\. Daepp, P\. Gorkin, A\. Shaffer, and K\. VossFinding ellipses: what blaschke products, poncelet’s theorem, and the numerical range know about each other\.Carus Mathematical Monographs, Vol\.34,American Mathematical Society\.Cited by:[§1](https://arxiv.org/html/2609.38369#S1.p3.1)\.
- Hunzikeret al\.\(2022\)M\. Hunziker, A\. Martinez\-Finkelshtein, T\. Poe, and B\. SimanekPoncelet–darboux, kippenhahn, and szego: interactions between projective geometry, matrices, and orthogonal polynomials\.Journal of Mathematical Analysis and Applications511,pp\. 126049\.Cited by:[§1](https://arxiv.org/html/2609.38369#S1.p3.1)\.
- Linfield \(1920\)B\. Z\. LinfieldOn the relation of the roots and poles of a rational function to the roots of its derivative\.Bulletin of the American Mathematical Society27,pp\. 17–21\.Cited by:[§1](https://arxiv.org/html/2609.38369#S1.p3.1)\.
- Mirman and Shukla \(2005\)B\. Mirman and P\. ShuklaA characterization of complex plane poncelet curves\.Linear Algebra and its Applications408,pp\. 86–119\.Cited by:[§1](https://arxiv.org/html/2609.38369#S1.p3.1)\.
- Romera\-Paredeset al\.\(2024\)B\. Romera\-Paredes, M\. Barekatain, A\. Novikov, M\. Balog, M\. P\. Kumar, E\. Dupont, F\. J\. R\. Ruiz, J\. S\. Ellenberg, P\. Wang, O\. Fawzi, P\. Kohli, and A\. FawziMathematical discoveries from program search with large language models\.Nature625,pp\. 468–475\.Cited by:[§1](https://arxiv.org/html/2609.38369#S1.p2.1)\.
- Siebeck \(1864\)J\. SiebeckUeber eine neue analytische behandlungweise der brennpunkte\.Journal für die reine und angewandte Mathematik64,pp\. 175–182\.Cited by:[§1](https://arxiv.org/html/2609.38369#S1.p3.1)\.
- Trinhet al\.\(2024\)T\. H\. Trinh, Y\. Wu, Q\. V\. Le, H\. He, and T\. LuongSolving olympiad geometry without human demonstrations\.Nature625,pp\. 476–482\.Cited by:[§1](https://arxiv.org/html/2609.38369#S1.p2.1)\.
- Turpinet al\.\(2023\)M\. Turpin, J\. Michael, E\. Perez, and S\. R\. BowmanLanguage models don’t always say what they think: unfaithful explanations in chain\-of\-thought prompting\.InAdvances in Neural Information Processing Systems,External Links:[Link](https://arxiv.org/abs/2305.04388)Cited by:[§5](https://arxiv.org/html/2609.38369#S5.p2.1)\.
- Udrescu and Tegmark \(2020\)S\. Udrescu and M\. TegmarkAI Feynman: a physics\-inspired method for symbolic regression\.Science Advances6\(16\),pp\. eaay2631\.External Links:[Document](https://dx.doi.org/10.1126/sciadv.aay2631)Cited by:[§1](https://arxiv.org/html/2609.38369#S1.p2.1)\.

## Appendix AAvailable records and reproducibility boundaries

#### Code and data availability\.

The code, numerical data, saved task prompts and research logs, and recovered run metadata are publicly available in the following repository:

[https://github\.com/yezeytuncu/blaschke\-ai\-rediscovery](https://github.com/yezeytuncu/blaschke-ai-rediscovery)\.

The experimental artifacts used here are preserved in repository commit[578b784](https://github.com/yezeytuncu/blaschke-ai-rediscovery/tree/578b7840744874426c8809759509c6969598a662)\. Raw application conversations are not included\.

The project records contain the generator, discovery and ablation CSV files, task prompts, and separate protocol, research\-log, and evaluation files\. On September 29, 2026, we additionally recovered metadata and tool\-call records from the original local application sessions dated August 26\. The prompts below are the saved task prompts, not a reconstruction of the full system prompt or complete conversation\. The original summaries remain unchanged; a separate metadata audit records the recovered information without publishing raw app logs\.

Both child sessions record provideropenai, model identifiergpt\-5\.6\-sol, reasoning effortlow, and harness version0\.150\.0\-alpha\.8\. The model identifier is not a dated weights snapshot\. Parent launch calls specifyfork\_turns: noneand no explicit model override\. The full\-data run used shell commands and R scripts after a failed Python/NumPy import; the sparse\-data run used Python standard\-library calculations\. Both wrote their reports through file\-editing tools\. The visible child tool calls introduce the evaluation CSV only in the second task turn\. This supports the recorded discovery/evaluation ordering but is not a complete audit of launch\-message content or all possible context exposure\.

Table 2:Recovered original\-run accounting\. Elapsed times span recorded task start to completion\. Output\-token counts are the harness\-reported field, not a prespecified budget or a cost estimate\.The audit exports only selected metadata, timestamps, usage counters, launch options, source line numbers, and source\-file hashes\. It excludes raw messages and reasoning content\. Temperature, random seed, original R/Python versions, and a prespecified execution budget remain unestablished\. The inspected parent session contains two experiment launches; this is not an exhaustive account\-wide attempt count or evidence of a prospective run\-selection rule\. These limitations prevent exact agent\-level reproduction even though the numerical computations are independently checkable\.

#### Data and numerical computation\.

The original generator is namedblaschke\_invariant\.pyin the experiments directory\. The CSV files are stored inoutputs/degree4\_baseline/\. Each row containstheta, the real and imaginary parts ofλ\\lambda,chord\_type, and\(u,v,w\)\(u,v,w\)\. The sign convention makes the first line coefficient with magnitude above10−1210^\{\-12\}positive after normalization byu2\+v2\\sqrt\{u^\{2\}\+v^\{2\}\}\. This convention fixes a CSV representation, not an additional geometric condition\.

The generator itself includes a cubic\-fitting calculation\. This experimenter\-side computation is not evidence of agent discovery and was not an allowed discovery input under the archived protocol\. The relevant agent outcome is the separately recorded frozen polynomial and its evaluation\. The discovery and evaluation CSV files remain unchanged in the camera\-ready preparation\.

### Saved task prompt for the full\-data trajectory

### Saved task prompt for the sparse\-data trajectory

#### Protocol scaffold\.

The protocol permits only the respective prompt and discovery CSV before freezing; it prohibits evaluation data, other workspace files, browsing, external theorem lookup, and parent\-agent hints\. The full\-data agent must record discarded hypotheses, a precise conjecture, coefficients, and thresholds\. The parent agent then audits the written log and releases only the evaluation CSV, with no refitting allowed\. These are the recorded instructions and phase boundaries\. We do not infer a hardware\-enforced isolation mechanism or a faithful hidden reasoning trace from them\.

## Appendix BPost\-review baseline and parameter\-count controls

The new scriptexperiments/camera\_ready\_baseline\.pyreads the stored CSV files without regenerating them\. Its monomial order enumerates exponents\(i,j,d−i−j\)\(i,j,d\-i\-j\)withiidecreasing fromddto zero and, for eachii,jjdecreasing fromd−id\-ito zero\. Atd=3d=3, this is precisely the order used in equation \([2](https://arxiv.org/html/2609.38369#S3.E2)\)\. Singular\-value decomposition is performed on the raw monomial matrix, without column standardization\. The relative numerical rank threshold is10−1210^\{\-12\}\. If there are fewer rows than columns, nullity includes the additional right\-nullspace dimensions\. Coefficients are initially normalized to Euclidean norm one\. For Table[1](https://arxiv.org/html/2609.38369#S4.T1), the cubic is rescaled so that itsw3w^\{3\}coefficient is−1\-1\.

Table 3:Numerical degree search on the 320 discovery side rows\.For the parameter\-count diagnostic, sort the 80 discovery angles and select indices⌊80​j/k⌋\\lfloor 80j/k\\rfloor,j=0,…,k−1j=0,\\ldots,k\-1\. Each selected configuration supplies its four side lines, and each selected row is repeated80/k80/ktimes\. All testedkkdivide 80\. This creates 320 rows in every condition\. The numerical nullity is checked against that of the unreplicated subset; the two agree in every reported case\. Evaluation is reported only for a one\-dimensional cubic nullspace, since an arbitrary vector from a larger nullspace is not a uniquely identified equation\.

Table 4:Post\-review deterministic controls, not additional agent ablations\. Residuals here use unit Euclidean coefficient norm, unlike Table[1](https://arxiv.org/html/2609.38369#S4.T1)\.The controls were run with Python 3\.12\.4, NumPy 2\.5\.3, and Matplotlib 3\.11\.2 on macOS arm64\. They use no random sampling\. The output JSON records software versions, input SHA\-256 hashes, singular\-value diagnostics, coefficient comparisons, residuals, and repeated\-row controls\. These are versions for the new computations, not recovered versions for the original agent runs\. Floating\-point summation order and numerical libraries can change the final few digits at roundoff scale\.

The baseline was designed with knowledge of the original result\. Although its code fits and selects using discovery sides alone, this is a retrospective comparison, not a newly blinded experiment\. Its role is to determine whether a simple non\-agent procedure suffices on the published instance\. Repeated rows are deliberately redundant and cannot substitute for a future study of distinct observations, independent product instances, or repeated agent trials\.

相似文章

探索物理问题中的结构:AI智能体能发现统计力学映射吗?

arXiv cs.AI

本文介绍了StatMechBench-v0,一个用于评估基于LLM的AI智能体能否从原始配分函数中发现统计力学映射到可处理表示的基准。结果表明,智能体常常能通过数值检查,但会错误识别底层结构,这凸显了当前LLM推理的局限性以及对更丰富验证方式的需求。

基于对偶智能体的凸松弛的AI辅助发现

arXiv cs.AI

本文介绍了一种使用LLM智能体(编码智能体和理论智能体)发现尖锐常数不等式的凸松弛的自动研究范式,从而改进了两个优化常数的认证下界。

平坦性代理并非函数时:鲁棒性证书与训练干预手段

arXiv cs.LG

本文指出,仅仅作为有效的曲率上界,并不足以证明使用末层相对平坦性代理作为对抗鲁棒性证书或可微训练正则化手段的合理性。作者推导出一种规范不变的修复方法,展示了该代理在保持对称性的 softmax 平移下是无界的,并通过 CIFAR-10 实验表明:经商空间正则化后的预测器保持了对齐性,而直接进行原始正则化的预测器的泛化能力可能被抑制。