Physics-R1: An Audited Olympiad Corpus and Recipe for Visual Physics Reasoning
Summary
This paper audits multimodal physics evaluation pipelines, revealing issues like train-eval contamination, translation drift, and MCQ saturation. It releases new datasets (PhysCorp-A, PhysR1Corp, PhysOlym-A) and a training recipe (Physics-R1) that significantly improves performance on held-out olympiad problems.
View Cached Full Text
Cached at: 05/15/26, 06:18 AM
# An Audited Olympiad Corpus and Recipe for Visual Physics Reasoning
Source: [https://arxiv.org/html/2605.14040](https://arxiv.org/html/2605.14040)
###### Abstract
We audit the multimodal\-physics evaluation pipeline end\-to\-end and document three undetected construction practices that distort how the field measures vision\-language reasoning: train–eval contamination, translation drift, and MCQ saturation\. \(1\) Public training pools \(UGPhysics\-Train, SciInstruct, MMK12\) pass single\-stage 5\-gram\-Jaccard audits with zero hits across all six public physics evals; a three\-stage audit \(Jaccard→\\tomxbai\-embed\-largecosine→\\toHaiku\-4\.5 LLM\-judge\) surfaces𝟏𝟑𝟒\\mathbf\{134\}near\-duplicates and𝟒,𝟖𝟒𝟔\\mathbf\{4\{,\}846\}paraphrase candidates in SciInstruct alone\. \(2\) A 17\-pp Sonnet\-4\.5\(Anthropic,[2025](https://arxiv.org/html/2605.14040#bib.bib53)\)delta on 59 paired Estonian\-English olympiad problems \(30\.5%30\.5\\%vs\.13\.6%13\.6\\%; sign testp=0\.011p\{=\}0\.011, McNemarp=0\.021p\{=\}0\.021, paired bootstrap95%95\\%CI\[\+5\.1,\+28\.9\]\[\+5\.1,\+28\.9\]pp\)\. \(3\) A 46\-pp format\-and\-novelty gradient on identical Sonnet weights between MCQ \(79\.7%79\.7\\%on PhyX\) and open\-ended olympiad evaluation \(33\.4%33\.4\\%onPhysOlym\-A\)\. We release four artifacts addressing these gaps:PhysCorp\-A\(6,4326\{,\}432\-record three\-stage\-audited multimodal corpus\),PhysR1Corp\(2,2682\{,\}268\-record closed\-form RL pool\),PhysOlym\-A\(500500\-problem,99\.8%99\.8\\%novel\-source held\-out olympiad eval with native difficulty labels and an EN/ET bilingual subset\), and Physics\-R1, a reference GSPO\+DAPO recipe cold\-started from Qwen3\-VL\-8B\-Thinking\. Across33seeds \(§[5](https://arxiv.org/html/2605.14040#S5)\), Physics\-R1 lifts the audited corpus over the 8B base by\+18\.3\+18\.3pp onPhysOlym\-Aliberal \(8\.0→26\.3±1\.78\.0\{\\to\}\\mathbf\{26\.3\}\{\\pm\}1\.7;7\.17\.1pp behind Sonnet 4\.5\),\+15\.7\+15\.7pp on PhysReason \(23\.9→39\.6±6\.423\.9\{\\to\}\\mathbf\{39\.6\}\{\\pm\}6\.4; ahead of Qwen3\-VL\-32B and Gemini 2\.5 Pro\),\+6\.9\+6\.9pp on OlympiadBench\-Physics \(46\.2±1\.5\\mathbf\{46\.2\}\{\\pm\}1\.5\), and\+4\.1\+4\.1pp on PhyX MCQ \(77\.8±0\.3\\mathbf\{77\.8\}\{\\pm\}0\.3\)\.
## 1Introduction
Multimodal physics reasoning is increasingly tracked via vision\-language benchmarks, but how those benchmarks are constructed is rarely audited\. Researcher\-curated training pools aggregate physics problems from publicly available sources whose paraphrase relationships evade conventional n\-gram dedup; multilingual benchmarks distribute English translations of problems first composed in another language; MCQ\-format splits saturate against the closed\-frontier ceiling\. Each represents a methodological gap in how the field constructs benchmarks, and together they distort cross\-model comparisons, inflate frontier\-model rankings on public leaderboards, and obscure the format\-and\-novelty axis along which capability actually diverges\.
We argue that defensible measurement of multimodal physics reasoning requires an end\-to\-end audit of the evaluation pipeline\. This paper performs that audit, surfaces three measurement findings, and constructs released artifacts directly against the gap each finding identifies\. Physics\-R1, a reference GSPO\+DAPO recipe\(Zhenget al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib12); Yuet al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib11)\)cold\-started from Qwen3\-VL\-8B\-Thinking\(Qwen Team,[2025](https://arxiv.org/html/2605.14040#bib.bib17)\)and building on MM\-Eureka\(Menget al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib33)\)and DeepSeek\-R1’s binary correctness signal\(DeepSeek\-AI,[2025](https://arxiv.org/html/2605.14040#bib.bib34); Shaoet al\.,[2024](https://arxiv.org/html/2605.14040#bib.bib10)\), accompanies the corpus as evidence\-of\-trainability rather than as the primary contribution: it lifts the audited held\-out eval over the 8B base while still trailing the closed frontier \(§[5\.2](https://arxiv.org/html/2605.14040#S5.SS2.SSS0.Px5)\)\.
#### Finding 1: single\-stage 5\-gram\-Jaccard audit reports public physics\-VL training pools as clean, but a three\-stage audit \(Jaccard→\\tomxbai cosine→\\toLLM\-judge\) surfaces𝟏𝟑𝟒\\mathbf\{134\}near\-duplicates among4,8464\{,\}846Stage\-2 candidates in SciInstruct alone\.
Across the three published physics\-VL training pools we re\-audit against six public evals \(UGPhysics\-Train, SciInstruct’s 42K\-record en\_phy\_chem split, MMK12’s 15K\-record train pool\), conventional 5\-gram\-Jaccard atJ≥0\.4J\\geq 0\.4\(Stage\-1\) reports*zero*hits for every pool against all six evals—a single\-stage audit calls them all clean\. Stage\-2mxbai\-embed\-largecosine at≥0\.85\\geq 0\.85then surfaces𝟒,𝟖𝟒𝟔\\mathbf\{4\{,\}846\}paraphrase\-class candidate pairs from SciInstruct alone \(PhysReason\-full2,6872\{,\}687, PhysUniBench\-en1,0271\{,\}027dominant\),99from UGPhysics\-Train, and6666from MMK12 \(Table[2](https://arxiv.org/html/2605.14040#S3.T2)\)\. Stage\-3, a Haiku\-4\.5 LLM\-judge, classifies each Stage\-2 candidate as a*close duplicate*or a*same\-topic neighbor*: of the4,8464\{,\}846SciInstruct candidates,𝟏𝟑𝟒\\mathbf\{134\}\(2\.8%2\.8\\%\) are close duplicates and the duplicate fraction is sharply cosine\-driven \(100%100\\%atcos≥0\.95\\cos\\geq 0\.95,1\.5%1\.5\\%atcos∈\[0\.85,0\.87\)\\cos\\in\[0\.85,0\.87\)\)\. On a1,6791\{,\}679\-record researcher\-curated sample ofPhysCorp\-pre\-audit\(14,29414\{,\}294records\) under the field\-default within\-pool dedup workflow,345345records \(20\.5%\\mathbf\{20\.5\\%\}\) leak at Stage\-1 alone against the six public evals \(concentrated in PhysUniBench\-en,339339, and MMMU\-Pro Physics,2020\); the joint Stage\-1∨\\veeStage\-2 sweep on this same sample against an internal analysis eval reaches8\.8%\\mathbf\{8\.8\\%\}at the published operating point and27\.1%27\.1\\%atcos≥0\.80\\cos\\geq 0\.80\(Table[4](https://arxiv.org/html/2605.14040#A1.T4)\)\.
#### Finding 2: translation introduces a measurable score delta on identical physics problems\.
On 59 paired Estonian/English Physics Olympiad problems, Sonnet 4\.5\(Anthropic,[2025](https://arxiv.org/html/2605.14040#bib.bib53)\)attains30\.5%\\mathbf\{30\.5\\%\}strict on Estonian originals against only13\.6%\\mathbf\{13\.6\\%\}on English translations of the same problems \(sign test on 16 discordant pairsp=0\.011p\{=\}0\.011; McNemar exactp=0\.021p\{=\}0\.021; bootstrap95%95\\%CI\[\+5\.1,\+28\.9\]\[\+5\.1,\+28\.9\]pp\)\. Estonian PhO problems were composed in Estonian first; English versions are translations whose physics vocabulary, grammatical case mapping, and subtlety of scope degrade information content\. For Sonnet 4\.5, whose cross\-lingual transfer covers Estonian, published numbers on the English\-translation benchmark systematically*underestimate*model ability relative to original\-language gold; for models with weaker training in the original language, the relationship is expected to reverse \(App\.[H\.4](https://arxiv.org/html/2605.14040#A8.SS4)\(viii\), pre\-registered\) \(§[3\.2](https://arxiv.org/html/2605.14040#S3.SS2), §[5\.1](https://arxiv.org/html/2605.14040#S5.SS1.SSS0.Px3)\)\.
#### Finding 3: same\-model evaluation across three physics benchmarks reveals a 46\-point format\-and\-novelty gradient\.
Evaluated in the same week on identical Sonnet 4\.5 weights, the score sweeps from79\.7%\\mathbf\{79\.7\\%\}on PhyX\(Shenet al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib1)\)\(4\-way MCQ\) down to50\.4%\\mathbf\{50\.4\\%\}liberal on OlympiadBench\-Physics\(Heet al\.,[2024](https://arxiv.org/html/2605.14040#bib.bib2)\)and33\.4%\\mathbf\{33\.4\\%\}liberal on our held\-out audited eval—format\-and\-novelty alone move the score by4646points on fixed weights \(§[3\.2](https://arxiv.org/html/2605.14040#S3.SS2); scoring in §[5](https://arxiv.org/html/2605.14040#S5)\)\.
Together the three findings imply that defensible physics\-VL measurement requires three properties at construction time: a three\-stage audit \(n\-gram Jaccard→\\toembedding cosine→\\toLLM\-judge precision filter\), original\-language gold, and open\-ended novel\-source evaluation\. Four released artifacts instantiate this protocol: \(a\)PhysCorp\-A, the audited multimodal physics corpus produced by the three\-stage pipeline \(Algorithm[1](https://arxiv.org/html/2605.14040#alg1)\), and the closed\-form RL training poolPhysR1Corpon which Physics\-R1 is trained \(§[3](https://arxiv.org/html/2605.14040#S3)\); \(b\)PhysOlym\-A, the open\-ended held\-out olympiad benchmark with native difficulty calibration, an EN/ET bilingual subset, and a Sonnet\-as\-judge protocol whose unjudgeable rate \(13\.9%13\.9\\%\) we disclose \(§[3\.2](https://arxiv.org/html/2605.14040#S3.SS2), §[5\.1](https://arxiv.org/html/2605.14040#S5.SS1)\); \(c\) Physics\-R1, a reference RL recipe whose audited held\-out lift onPhysOlym\-Avalidates the corpus as trainable rather than memorized \(Table[3](https://arxiv.org/html/2605.14040#S5.T3)\); we recommend a binary correctness reward as the default—variance\-optimal under GSPO with group\-normalized advantages, Goodhart\-robust against unit/conservation/format proxies, and harness\-portable \(§[4](https://arxiv.org/html/2605.14040#S4), properties P1–P4\)—and report the dense five\-component physics\-native reward as a shape ablation; and \(d\) the audit protocol itself, released asaudit\_three\_stage\.pywith saved best\-overlap scores and Stage\-3 judge labels \(Appendix[A](https://arxiv.org/html/2605.14040#A1)\)\. The 3\-seed sensitivity sweep \(seeds\{42,17,23\}\\\{42,17,23\\\}on the auditedPhysR1Corp\) is reported in Table[3](https://arxiv.org/html/2605.14040#S5.T3)withσ≤3\.3\\sigma\\leq 3\.3pp on PUB\-OE, OlymBench\-Phys, andPhysOlym\-A, andσ=6\.4\\sigma\{=\}6\.4pp on PhysReason \(seed\-42 outlier\); the reward\-component drop\-out ablation \(Table[11](https://arxiv.org/html/2605.14040#A8.T11)\) is left to follow\-up work\.
## 2Related Work
#### Rule\-based RL for reasoning\.
DeepSeek\-R1\(DeepSeek\-AI,[2025](https://arxiv.org/html/2605.14040#bib.bib34)\)established that simple rule\-based rewards \(binary correctness \+ format\) suffice to train competitive math reasoners directly from a base model without SFT, using GRPO\(Shaoet al\.,[2024](https://arxiv.org/html/2605.14040#bib.bib10)\)\. MM\-Eureka\(Menget al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib33)\)extended the recipe to VLMs with a difficulty curriculum; DAPO\(Yuet al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib11)\)added decoupled clipping and dynamic sampling; GSPO\(Zhenget al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib12)\)replaced token\-level with sequence\-level importance weighting\. Physics\-R1 inherits MM\-Eureka’s structural choices and the binary correctness reward unchanged: although physics intermediate steps carry units, conservation laws, and symbolic equations that*a priori*admit per\-step verification, we find that under GSPO with group\-normalized advantages a binary reward is variance\-optimal and robust to the within\-wrong\-group Goodhart channel that physics\-native shaping opens \(§[4](https://arxiv.org/html/2605.14040#S4)\); the dense physics\-native reward is reported as an ablation\.
#### Physics QA benchmarks\.
PhyX\(Shenet al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib1)\), OlympiadBench\-Physics\(Heet al\.,[2024](https://arxiv.org/html/2605.14040#bib.bib2)\), UGPhysics\(Xuet al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib3)\), PhysReason\(Zhanget al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib4)\), MMMU/MMMU\-Pro\(Yueet al\.,[2024a](https://arxiv.org/html/2605.14040#bib.bib5),[b](https://arxiv.org/html/2605.14040#bib.bib41)\), MMK12\(Menget al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib33)\), PHYBench\(Qiuet al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib36)\), and PhysUniBench\(Wanget al\.,[2025b](https://arxiv.org/html/2605.14040#bib.bib37)\)are the canonical references\. Top entries cluster within ten points of the closed\-frontier ceiling on MCQ formats; only PHYBench, OIBench, and PutnamBench publish a contamination protocol, and none publish the three\-stage \(n\-gram, embedding, LLM\-judge\) pairwise audit we introduce in §[3\.3](https://arxiv.org/html/2605.14040#S3.SS3)\. Table[1](https://arxiv.org/html/2605.14040#S2.T1)maps our released audited corpus andPhysOlym\-Aagainst related benchmarks on seven axes\.
Table 1:Released artifacts vs related benchmarks across eight axes\.*Audit:*2\-stage \(n\-gram\+embedding\) / 1\-stage / orig\. \(constructed\-novel\) / none\.*T/T leak:*train→\\totest joint\-stage \(J≥0\.4∨cos≥0\.85J\\geq 0\.4\\vee\\cos\\geq 0\.85\) audit against six public physics evals;✓\\checkmarkall 6 = clean\.*Diff:*organizer difficulty\.*X\-L:*paired cross\-lingual\.*Use:*E/T = eval/train\.*RL\-ready:*closed\-form gold \+ audit\-clean \+ RL recipe\. “⋅\\cdot” = eval\-only; “n/r” = train pool, no cross\-corpus audit\. Only this work reports train/test contamination: after re\-audit cleanup,PhysCorp\-A\(6,432\) andPhysR1Corp\(2,268\) are clean againstall sixevals \(Table[2](https://arxiv.org/html/2605.14040#S3.T2)\)\.BenchmarkSizeFormatMMAuditT/T leakDiffX\-LUseRL\-ready*Physics\-domain benchmarks*PHYBench\(Qiuet al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib36)\)500open\+EED–orig\.⋅\\cdot––E–PhysUniBench\(Wanget al\.,[2025b](https://arxiv.org/html/2605.14040#bib.bib37)\)3,304open MM✓\\checkmark1\-stage⋅\\cdot✓\\checkmark–E–UGPhysics\(Xuet al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib3)\)5,520open text–1\-stagen/r–EN/ZHT–PhysReason\(Zhanget al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib4)\)1,200step open MM✓\\checkmarknone⋅\\cdot––E–OlympiadBench\(Heet al\.,[2024](https://arxiv.org/html/2605.14040#bib.bib2)\)8,952open MM✓\\checkmarknone⋅\\cdot–EN/ZHE–*Olympiad / formal / contamination\-by\-design*PutnamBench\(Tsoukalaset al\.,[2024](https://arxiv.org/html/2605.14040#bib.bib6)\)1,692Lean/Isab\.–orig\.⋅\\cdot✓\\checkmark–E–OIBench\(Zhuet al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib38)\)250open code–2\-stage⋅\\cdot✓\\checkmarkEN/ZHE–FrontierMath\(Glazeret al\.,[2024](https://arxiv.org/html/2605.14040#bib.bib7)\)290open math–orig\.⋅\\cdot✓\\checkmark–E–HLE\(Phanet al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib39)\)2,500expert exam✓\\checkmarkorig\.⋅\\cdot––E–*Multimodal / multi\-domain*MMLU\-Pro\(Wanget al\.,[2024](https://arxiv.org/html/2605.14040#bib.bib40)\)12,03210\-MCQ–none⋅\\cdot––E–MMMU\-Pro Phys\(Yueet al\.,[2024b](https://arxiv.org/html/2605.14040#bib.bib41)\)6010\-MCQ MM✓\\checkmarknone⋅\\cdot––E–SciInstruct\(Zhanget al\.,[2024](https://arxiv.org/html/2605.14040#bib.bib42)\)254,KSFT instr\.–1\-stagen/r––T–*This work \(train→\\totest cross\-corpus audit reported; Table[2](https://arxiv.org/html/2605.14040#S3.T2)\)*PhysCorp\-A\(ours\)6,432open\+MCQ MM✓\\checkmark2\-stage✓\\checkmarkall 6✓\\checkmark✓\\checkmarkT✓\\checkmarkPhysR1Corp\(ours\)2,268MCQ \+ num MM✓\\checkmark2\-stage✓\\checkmarkall 6✓\\checkmark✓\\checkmarkT✓\\checkmarkPhysOlym\-A\(ours\)500open MM novel✓\\checkmark2\-stage⋅\\cdot\(eval; clean\)✓\\checkmarkEN/ETE–
✓\\checkmark= present; – = absent or not reported\. “2\-stage” audit = pairwise 5\-gram\-Jaccard*and*embedding\-cosine against external corpora and held\-out splits\.
#### Contamination audits and other prior work\.
PutnamBench\(Tsoukalaset al\.,[2024](https://arxiv.org/html/2605.14040#bib.bib6)\), FrontierMath\(Glazeret al\.,[2024](https://arxiv.org/html/2605.14040#bib.bib7)\), HLE\(Phanet al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib39)\), and EnigmaEval\(Wanget al\.,[2025a](https://arxiv.org/html/2605.14040#bib.bib51)\)provide release\-policy templates and dismissal grounds; methodological work spans n\-gram audits\(Sainzet al\.,[2023](https://arxiv.org/html/2605.14040#bib.bib8)\), the rephrased\-samples failure mode\(Yanget al\.,[2023](https://arxiv.org/html/2605.14040#bib.bib43)\)\(which our Stage 2 catches\), embedding\-based detection\(Singhet al\.,[2024](https://arxiv.org/html/2605.14040#bib.bib9)\), and performance\-based detection\(Dekonincket al\.,[2024](https://arxiv.org/html/2605.14040#bib.bib49)\); the survey ofRavautet al\.\([2024](https://arxiv.org/html/2605.14040#bib.bib48)\)consolidates these\. We import the math template, adding the embedding\-cosine pass because physics statements \(units, vectors, figure references\) are more paraphrase\-sensitive than typical math problems—a sensitivity Table[4](https://arxiv.org/html/2605.14040#A1.T4)quantifies\. PhysBench\(Chowet al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib47)\)evaluates intuitive\-physics dynamics from video, orthogonal scope\. Multilingual benchmarks have proliferated\(Xuanet al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib46); Ahujaet al\.,[2024](https://arxiv.org/html/2605.14040#bib.bib44); Wuet al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib50)\); our cross\-lingual finding \(§[5\.1](https://arxiv.org/html/2605.14040#S5.SS1.SSS0.Px3)\) differs methodologically by evaluating identical 59 problems in original Estonian and English translation on the same closed model with paired tests, isolating a within\-problem effect aggregate benchmarks cannot\.
## 3Data: The Audited Corpus and Held\-Out Olympiad Eval
Released artifacts:PhysCorp\-A\(6,432\-record audited corpus, including 1,609 first\-ML\-format olympiad problems—Estonian PhO with native 1–10 difficulty \+ 201 EN/ET bilingual, Kevin Zhou’s handouts, 7 international olympiads\);PhysR1Corp\(2,268\-record closed\-form RL pool, MCQ and numerical only\); the held\-outPhysOlym\-Aeval \(§[3\.2](https://arxiv.org/html/2605.14040#S3.SS2)\); the Physics\-R1 recipe \(Algorithms[2](https://arxiv.org/html/2605.14040#alg2),[3](https://arxiv.org/html/2605.14040#alg3)\); and the audit pipeline \(Algorithm[1](https://arxiv.org/html/2605.14040#alg1), Table[4](https://arxiv.org/html/2605.14040#A1.T4)\)\. All ship under per\-source licenses \(Table[5](https://arxiv.org/html/2605.14040#A5.T5)\) on HuggingFace\+GitHub\+Zenodo with Croissant 1\.0 metadata\.
### 3\.1Training Corpus Composition
The corpus is drawn from nine source families \(Table[9](https://arxiv.org/html/2605.14040#A8.T9)\)\. Five are repackaged from existing benchmarks under documented licenses \(UGPhysics\(Xuet al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib3)\), OpenStax College and University Physics\(OpenStax,[2024](https://arxiv.org/html/2605.14040#bib.bib30)\), Physics Stack Exchange\(Stack Exchange Inc\.,[2024](https://arxiv.org/html/2605.14040#bib.bib31)\), an MMMU\+o1\-CoT seed\(Yueet al\.,[2024a](https://arxiv.org/html/2605.14040#bib.bib5)\), PhysReason\(Zhanget al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib4)\)\); four contribute first\-ML\-format material: the Estonian Physics Olympiad collection\(Estonian Physics Olympiad,[2018](https://arxiv.org/html/2605.14040#bib.bib22)\)\(418 problems, 2004–2018, with organizer\-issued 1–10 difficulty labels and a 201\-problem bilingual EN\+ET subset\), Kevin Zhou’s olympiad handouts\(Zhou,[2018](https://arxiv.org/html/2605.14040#bib.bib23)\)\(692 problems, with native point values 1–5 and a 3\.2% advanced flag; some problems are drawn from books or other olympiad archives with inline attribution preserved per record, see Appendix[E](https://arxiv.org/html/2605.14040#A5)\), and refreshed scrapes of seven international olympiads \(IPhO\(International Physics Olympiad,[2025](https://arxiv.org/html/2605.14040#bib.bib24)\), NBPhO\(NBPhO Committee,[2025](https://arxiv.org/html/2605.14040#bib.bib28)\), EuPhO\(EuPhO Committee,[2025](https://arxiv.org/html/2605.14040#bib.bib29)\), APhO\(Asian Physics Olympiad Committee,[2025](https://arxiv.org/html/2605.14040#bib.bib26)\), USAPhO\(American Association of Physics Teachers,[2025](https://arxiv.org/html/2605.14040#bib.bib25)\), INPhO\(Homi Bhabha Centre for Science Education,[2025](https://arxiv.org/html/2605.14040#bib.bib27)\), IYPT\)\. Source families ship under a mix of CC BY 4\.0, CC BY\-SA 4\.0, public\-domain by competition policy \(Estonian PhO, IPhO, NBPhO, EuPhO, APhO, USAPhO, INPhO\), CC BY\-NC 4\.0 \(Kevin Zhou’s handouts; written grant 2026\-05\-03\), and CC BY\-NC\-SA 4\.0 \(UGPhysics\); per\-source licenses are listed in Appendix[E](https://arxiv.org/html/2605.14040#A5)\(Table[5](https://arxiv.org/html/2605.14040#A5.T5)\) and carried through to each released record\. The full14,29414\{,\}294\-record pre\-audit pool is released asPhysCorp\-pre\-auditso that downstream users can reproduce the audit;PhysCorp\-Ais the6,4326\{,\}432\-record subset that survives all three stages plus a re\-audit against PhysReason\-full and PhysUniBench\-en \(804 records dropped, dominated by PhysReason\-full540540and PhysUniBench\-en186186\)\. The released pool is disjoint from PhyX, MMMU\-Pro Physics, OlympiadBench\-Physics, UGPhysics\-Train, PhysReason\-full, PhysUniBench\-en, andPhysOlym\-Aat the joint operating thresholds\. The candidate\-to\-release cleanup forPhysR1Corpis detailed in §[3\.3](https://arxiv.org/html/2605.14040#S3.SS3)\.
#### LLM\-touched\-statement subset disclosure\.
Of the2,2682\{,\}268records inPhysR1Corp, approximately7373\(3\.2%3\.2\\%\) have LLM\-touched problem statements:∼11\\sim 11are derived from a8585\-record Claude\-generated synthetic\-MCQ augmentation pool \(3 verbatim, 8 numeric paraphrases\), and∼62\\sim 62are numeric\-variation paraphrases of realPhysCorp\-Arecords \(e\.g\., variant problem constants\)\. The remaining∼2,195\\sim 2\{,\}195records have unmodified problem statements from the nine source families\. LLM augmentation is documented per\-distribution in the Croissant metadata’ssyntheticDataDescriptionfield; the held\-outPhysOlym\-Aeval contains no synthetic problem content\.
### 3\.2PhysOlym\-A: Held\-Out Olympiad Eval
Standard physics\-VL benchmarks no longer resolve frontier\-class differences: PhyX clusters top entries within ten points of the80%80\\%ceiling; OlympiadBench\-Physics predates the contamination\-audit discipline; UGPhysics is itself a candidate for audited training data, not held\-out evaluation\. Physics\-R1’s stopping rule and reward\-component ablation depend on a held\-out signal that is non\-saturating and contamination\-clean against the training pool\.
PhysOlym\-A\(Physics Olympiad, Audited\) is composed of 200 problems from Kevin Zhou’s olympiad handouts, 136 from the Estonian PhO collection, 85 from an IPhO/NBPhO/EuPhO scrape, and 79 from an APhO/USAPhO/INPhO scrape \(500 total,𝟒𝟗𝟗\\mathbf\{499\}novel\-source under our four\-corpus audit\)\. Native difficulty signals:27%27\\%of records carry Estonian organizer\-issued 1–10 difficulty;38%38\\%carry Zhou’s pedagogical 1–5 point values;2%2\\%carry Zhou’s advanced \[A\] flag\. The three\-stage audit \(§[3\.3](https://arxiv.org/html/2605.14040#S3.SS3)\) certifies𝟎\\mathbf\{0\}Stage\-3 near\-duplicate overlaps between the audited training pool andPhysOlym\-A, and𝟎\\mathbf\{0\}overlaps between the novel pool and PhyX 1000q\. The single non\-novel record is an EuPhO 2020 problem also present in OlympiadBench\-Physics atJ=0\.91J\{=\}0\.91; we disclose this in Appendix[A](https://arxiv.org/html/2605.14040#A1)rather than silently drop it\. The scoring protocol \(LLM\-judge with strict/liberal accuracy,κ\\kappainter\-judge agreement, and the auxiliary held\-out splits used during training\) is described in §[5](https://arxiv.org/html/2605.14040#S5)\.
### 3\.3The Three\-Stage Audit Pipeline
Table 2:Train/test contamination across released physics\-VL training pools, three\-stage audit\.Rows: public physics eval splits; columns: training pools \(three competitor, two cleaned ours\)\. Cells:Stage\-1 / Stage\-2 raw / Stage\-3 near\-duppair counts \(Algorithm[1](https://arxiv.org/html/2605.14040#alg1)\)\. Stage\-1 = 5\-gram Jaccard≥0\.4\\geq 0\.4, Stage\-2 =mxbai\-embed\-large\-v1cosine≥0\.85\\geq 0\.85\(high recall over close\-content pairs\), Stage\-3 = Haiku\-4\.5 LLM judge separating each Stage\-2 candidate into*close duplicate*\(paraphrase / numeric variation of the same problem\) vs\.*same\-topic neighbor*\(related physics, distinct setup\)\. Competitor pools \(UGPhysics\-Train: 200\-record annotated subset; SciInstruct: en\_phy\_chem42,35242\{,\}352\-record subset of 254 K; MMK12:15,60815\{,\}608\-record MM\-Eureka train pool\) report𝟎/𝟔\\mathbf\{0/6\}Stage\-1 hits; Stage\-2 surfaces𝟒,𝟖𝟒𝟔\\mathbf\{4\{,\}846\}close\-content pairs in SciInstruct,99in UGPhysics\-Train, and6666in MMK12\.Stage\-3 LLM\-judge separates close duplicates from same\-topic neighbors:SciInstruct4,846→𝟏𝟑𝟒4\{,\}846\{\\to\}\\mathbf\{134\}near\-duplicates \(PhysReason\-full2,687→362\{,\}687\{\\to\}36, PhysUniBench\-en1,027→221\{,\}027\{\\to\}22, PhyX\-mini703→46703\{\\to\}46dominant\); UGPhysics\-Train9→𝟎9\{\\to\}\\mathbf\{0\}; MMK1266→𝟎66\{\\to\}\\mathbf\{0\}\. The close\-duplicate share is sharply cosine\-driven:100%100\\%atcos≥0\.95\\cos\\geq 0\.95vs\.1\.5%1\.5\\%atcos∈\[0\.85,0\.87\)\\cos\\in\[0\.85,0\.87\)\(Appendix[A](https://arxiv.org/html/2605.14040#A1)\)\.Both released pools are fully Stage\-3 clean against all six evals:PhysCorp\-A\(6,4326\{,\}432\) after dropping804804from a7,2367\{,\}236candidate, with𝟎/𝟎\\mathbf\{0/0\}Stage\-3 close\-duplicates against all six evals \(no S2 candidates surviving the joint S1∨\\veeS2 cleanup\);PhysR1Corp\(2,2682\{,\}268\) after dropping8787MMMU\-Pro \+7878PhyX\-mini/PhysUniBench\-en hits from a2,4332\{,\}433candidate, with all1919remaining S2 candidates classified as same\-topic neighbors by Stage\-3 \(𝟎/𝟏𝟗\\mathbf\{0/19\}near\-duplicates\), agreeing100%100\\%with manual inspection\.Eval↓\\downarrow/ Train pool→\\to*Other published pools \(we re\-audit\)**This work \(cleaned\)*UGPhysics\-Train\(200 sub\)SciInstruct\(en\_phy\_chem; 42 K\)MMK12\(MM\-Eureka; 15 K\)PhysCorp\-A\(6,432\)PhysR1Corp\(2,268\)PhysOlym\-A\(500\)0/0/00\\,/\\,0\\,/\\,\\mathbf\{0\}0/163/80\\,/\\,163\\,/\\,\\mathbf\{8\}0/0/00\\,/\\,0\\,/\\,\\mathbf\{0\}0/0/00\\,/\\,0\\,/\\,\\mathbf\{0\}0/3/00\\,/\\,3\\,/\\,\\mathbf\{0\}PhyX\-mini \(1,000\)0/1/00\\,/\\,1\\,/\\,\\mathbf\{0\}0/703/460\\,/\\,703\\,/\\,\\mathbf\{46\}0/0/00\\,/\\,0\\,/\\,\\mathbf\{0\}0/0/00\\,/\\,0\\,/\\,\\mathbf\{0\}0/0/00\\,/\\,0\\,/\\,\\mathbf\{0\}MMMU\-Pro Phys \(60\)0/0/00\\,/\\,0\\,/\\,\\mathbf\{0\}0/141/70\\,/\\,141\\,/\\,\\mathbf\{7\}0/0/00\\,/\\,0\\,/\\,\\mathbf\{0\}0/0/00\\,/\\,0\\,/\\,\\mathbf\{0\}0/1/00\\,/\\,1\\,/\\,\\mathbf\{0\}OlymBench\-Phys \(692\)0/2/00\\,/\\,2\\,/\\,\\mathbf\{0\}0/130/150\\,/\\,130\\,/\\,\\mathbf\{15\}0/0/00\\,/\\,0\\,/\\,\\mathbf\{0\}0/0/00\\,/\\,0\\,/\\,\\mathbf\{0\}0/4/00\\,/\\,4\\,/\\,\\mathbf\{0\}PhysReason\-full \(1,200\)0/4/00\\,/\\,4\\,/\\,\\mathbf\{0\}0/2,687/360\\,/\\,2\{,\}687\\,/\\,\\mathbf\{36\}0/62/00\\,/\\,62\\,/\\,\\mathbf\{0\}0/0/00\\,/\\,0\\,/\\,\\mathbf\{0\}0/11/00\\,/\\,11\\,/\\,\\mathbf\{0\}PhysUniBench\-en \(1,022\)0/2/00\\,/\\,2\\,/\\,\\mathbf\{0\}0/1,027/220\\,/\\,1\{,\}027\\,/\\,\\mathbf\{22\}0/4/00\\,/\\,4\\,/\\,\\mathbf\{0\}0/0/00\\,/\\,0\\,/\\,\\mathbf\{0\}0/0/00\\,/\\,0\\,/\\,\\mathbf\{0\}*Total S2 / S3 real*99/𝟎\\mathbf\{0\}𝟒,𝟖𝟒𝟔\\mathbf\{4\{,\}846\}/𝟏𝟑𝟒\\mathbf\{134\}6666/𝟎\\mathbf\{0\}0/𝟎\\mathbf\{0\}𝟏𝟗\\mathbf\{19\}/𝟎\\mathbf\{0\}
Format:Stage\-1 / Stage\-2 / Stage\-3pair counts \(Stage\-3 = Haiku\-4\.5 LLM\-judge classifying each Stage\-2 candidate as close duplicate vs\. same\-topic neighbor\)\. SciInstruct’s S3\-near\-dup cells reveal the close\-duplicate share is threshold\-driven:17/1717/17atcos≥0\.95\\cos\{\\geq\}0\.95,54/1,15954/1\{,\}159at\[0\.87,0\.90\)\[0\.87,0\.90\),53/3,54353/3\{,\}543at\[0\.85,0\.87\)\[0\.85,0\.87\)\(Appendix[A](https://arxiv.org/html/2605.14040#A1), Table[A](https://arxiv.org/html/2605.14040#A1.SS0.SSS0.Px1)\)\.
The pipeline constructs both the audited training pool and the held\-outPhysOlym\-Aeval under the same definition of contamination, applied pairwise across the training pool, four external corpora \(PhyX, MMMU\-Pro Physics, OlympiadBench\-Physics, UGPhysics\-Train\), and the held\-out splits\.*Stage 1 \(n\-gram\)\.*Tokenize each problem statement with a unicode word tokenizer, build the 5\-gram shingle set, and flag pairs with Jaccard≥0\.4\\geq 0\.4\.*Stage 2 \(embedding\)\.*Encode each statement withmxbai\-embed\-large\(1024\-dim,L2L\_\{2\}\-normalized\) and flag pairs with cosine≥0\.85\\geq 0\.85\. Stage\-2 has high recall on close\-content pairs, including the rephrasing\-class duplicates Stage\-1 misses, but its single\-threshold operating point also flags same\-topic\-but\-distinct\-problem pairs\.*Stage 3 \(LLM\-judge precision filter\)\.*For each Stage\-2 candidate, a Haiku\-4\.5 judge receives both problem statements and classifies the pair as a*close duplicate*\(paraphrase or numeric variation of the same problem\) or a*same\-topic neighbor*\(related physics, distinct setup\)\. Only Stage\-3 close\-duplicate records are removed from the training pool\. Pseudocode is in Algorithm[1](https://arxiv.org/html/2605.14040#alg1); worked examples in Appendix[H\.5](https://arxiv.org/html/2605.14040#A8.SS5); calibration of the embedder \+ thresholds in Appendix[A](https://arxiv.org/html/2605.14040#A1)\.
On the train/test contamination matrix of Table[2](https://arxiv.org/html/2605.14040#S3.T2), the cosine\-bucketed precision pattern \(100%100\\%close\-duplicates atcos≥0\.95\\cos\\geq 0\.95vs\.1\.5%1\.5\\%atcos∈\[0\.85,0\.87\)\\cos\\in\[0\.85,0\.87\); Appendix[A](https://arxiv.org/html/2605.14040#A1), Table[A](https://arxiv.org/html/2605.14040#A1.SS0.SSS0.Px1)\) confirms the protocol’s design hypothesis: embedding cosine alone is recall\-dominant and an LLM judge is the appropriate precision filter\.Both released training pools are fully Stage\-3 clean against all six public evals \(Table[2](https://arxiv.org/html/2605.14040#S3.T2)\):PhysCorp\-A\(6,4326\{,\}432\), built via Stage\-1∨\\veeStage\-2 audit dropping804804of a7,2367\{,\}236candidate, surfaces0Stage\-2 candidates and hence𝟎/𝟎\\mathbf\{0/0\}Stage\-3 close\-duplicates by construction;PhysR1Corp\(2,2682\{,\}268\), additionally dropping8787MMMU\-Pro and7878PhyX\-mini/PhysUniBench near\-duplicates from a2,4332\{,\}433\-record candidate \(Appendix[A\.1](https://arxiv.org/html/2605.14040#A1.SS1)\), retains1919Stage\-2 candidates classified as same\-topic neighbors by Stage\-3 with100%100\\%manual\-inspection agreement \(𝟎/𝟏𝟗\\mathbf\{0/19\}close\-duplicates\)\.
Algorithm 1Three\-stage contamination audit\.1:Train pool
TT, external corpora
\{Ek\}k=1K\\\{E\_\{k\}\\\}\_\{k=1\}^\{K\}, held\-out splits
\{Hj\}j=1J\\\{H\_\{j\}\\\}\_\{j=1\}^\{J\}, normalize fn
norm\(⋅\)\\mathrm\{norm\}\(\\cdot\), embedder
enc\(⋅\)\\mathrm\{enc\}\(\\cdot\), LLM judge
Judge\(⋅,⋅\)∈\{close\-dup,topic\-neighbor\}\\textsc\{Judge\}\(\\cdot,\\cdot\)\\in\\\{\\text\{close\-dup\},\\text\{topic\-neighbor\}\\\}, thresholds
τJ=0\.4\\tau\_\{J\}\{=\}0\.4,
τC=0\.85\\tau\_\{C\}\{=\}0\.85
2:Audited pool
T′T^\{\\prime\}disjoint from
⋃kEk∪⋃jHj\\bigcup\_\{k\}E\_\{k\}\\cup\\bigcup\_\{j\}H\_\{j\}at the joint thresholds\.
3:*Stage 1: 5\-gram Jaccard \(n\-gram audit\)\.*
4:
St←\{5\-gram shingle set ofnorm\(t\)\}S\_\{t\}\\leftarrow\\\{\\text\{5\-gram shingle set of \}\\mathrm\{norm\}\(t\)\\\}for each
t∈T∪⋃Ek∪⋃Hjt\\in T\\cup\\bigcup E\_\{k\}\\cup\\bigcup H\_\{j\}
5:for
t∈Tt\\in Tdo
6:
Jmax\(t\)←maxx∈⋃Ek∪⋃Hj\|St∩Sx\|/\|St∪Sx\|J\_\{\\max\}\(t\)\\leftarrow\\max\_\{x\\in\\bigcup E\_\{k\}\\cup\\bigcup H\_\{j\}\}\\,\|S\_\{t\}\\cap S\_\{x\}\|/\|S\_\{t\}\\cup S\_\{x\}\|
7:endfor
8:*Stage 2:mxbai\-embed\-largecosine \(paraphrase recall\)\.*
9:
𝐞t←enc\(norm\(t\)\)/‖enc\(norm\(t\)\)‖\\mathbf\{e\}\_\{t\}\\leftarrow\\mathrm\{enc\}\(\\mathrm\{norm\}\(t\)\)/\\\|\\mathrm\{enc\}\(\\mathrm\{norm\}\(t\)\)\\\|for each
tt
10:for
t∈Tt\\in Tdo
11:
Cmax\(t\)←maxx∈⋃Ek∪⋃Hj𝐞t⊤𝐞xC\_\{\\max\}\(t\)\\leftarrow\\max\_\{x\\in\\bigcup E\_\{k\}\\cup\\bigcup H\_\{j\}\}\\mathbf\{e\}\_\{t\}^\{\\top\}\\mathbf\{e\}\_\{x\}
12:endfor
13:
C\(T\)←\{t:Jmax\(t\)≥τJorCmax\(t\)≥τC\}C\(T\)\\leftarrow\\\{t:J\_\{\\max\}\(t\)\\geq\\tau\_\{J\}\\,\\textsc\{or\}\\,C\_\{\\max\}\(t\)\\geq\\tau\_\{C\}\\\}⊳\\trianglerightcandidate set, high\-recall union
14:*Stage 3: Haiku\-4\.5 LLM\-judge \(precision filter\)\.*For each
t∈C\(T\)t\\in C\(T\)with top\-matching
x∗\(t\)←argmaxx𝐞t⊤𝐞xx^\{\*\}\(t\)\\leftarrow\\arg\\max\_\{x\}\\mathbf\{e\}\_\{t\}^\{\\top\}\\mathbf\{e\}\_\{x\}, queryJudgeto classify the pair as a close duplicate \(paraphrase / numeric variation of the same problem\) or a same\-topic neighbor \(related physics, distinct setup\)\.
15:
R\(T\)←\{t∈C\(T\):Judge\(t,x∗\(t\)\)=close\-dup\}R\(T\)\\leftarrow\\\{t\\in C\(T\):\\textsc\{Judge\}\(t,x^\{\*\}\(t\)\)=\\text\{close\-dup\}\\\}
16:
T′←T∖R\(T\)T^\{\\prime\}\\leftarrow T\\setminus R\(T\)
17:return
T′T^\{\\prime\}and per\-stage counts
\|Jmax≥τJ\|\|J\_\{\\max\}\\geq\\tau\_\{J\}\|,
\|Cmax≥τC\|\|C\_\{\\max\}\\geq\\tau\_\{C\}\|,
\|R\(T\)\|\|R\(T\)\|\(Tables[2](https://arxiv.org/html/2605.14040#S3.T2),[4](https://arxiv.org/html/2605.14040#A1.T4)\)\.
#### Threshold\-sensitive leakage on a researcher\-curated baseline \(Finding 1\)\.
On a 1,679\-record sample drawn fromPhysCorp\-pre\-auditunder conventional 5\-gram\-Jaccard \+ within\-pool embedding dedup, audited against a 500\-record internal analysis eval \(distinct fromPhysOlym\-A, constructed post\-audit\), the joint Stage\-1∨\\veeStage\-2 audit raises the detected leak rate from3\.3%3\.3\\%\(Stage\-1 alone, all exact matches atJ=1\.0J\{=\}1\.0\) to8\.8%\\mathbf\{8\.8\\%\}, sweeping4\.74\.7–27\.1%27\.1\\%as the cosine threshold moves between0\.900\.90and0\.800\.80\(Appendix[A\.1](https://arxiv.org/html/2605.14040#A1.SS1), Table[4](https://arxiv.org/html/2605.14040#A1.T4)\)\. The5\.55\.5\-pp gap is the rephrasing dark\-matter that justifies the audited release as a measurement intervention\.
## 4Physics\-R1: A Multi\-Model RL Recipe
Physics\-R1 is reported as evidence the audited corpus has training utility under standard rule\-based RL, not as an algorithmic contribution\. The optimizer is GSPO\(Zhenget al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib12)\)\+DAPO\(Yuet al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib11)\), unmodified\. For each promptxx, sampleK=16K\{=\}16rollouts\{yk\}∼πθold\(⋅∣x\)\\\{y\_\{k\}\\\}\\sim\\pi\_\{\\theta\_\{\\mathrm\{old\}\}\}\(\\cdot\\mid x\), score with rewardr\(yk,x\)r\(y\_\{k\},x\), form group\-normalized advantages and the clipped sequence\-level GSPO objective
Ak\\displaystyle A\_\{k\}=r\(yk,x\)−r¯σr\+ε,wk\(θ\)=\(πθ\(yk∣x\)πθold\(yk∣x\)\)1/\|yk\|,\\displaystyle=\\frac\{r\(y\_\{k\},x\)\-\\bar\{r\}\}\{\\sigma\_\{r\}\+\\varepsilon\},\\qquad w\_\{k\}\(\\theta\)=\\left\(\\frac\{\\pi\_\{\\theta\}\(y\_\{k\}\\mid x\)\}\{\\pi\_\{\\theta\_\{\\mathrm\{old\}\}\}\(y\_\{k\}\\mid x\)\}\\right\)^\{\\\!1/\|y\_\{k\}\|\},\(1\)ℒGSPO\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{GSPO\}\}=−𝔼\[1K∑kmin\(wkAk,clip\(wk,1±ϵ\)Ak\)\]\+βKLDKL\(πθ∥πbase\),\\displaystyle=\-\\mathbb\{E\}\\\!\\left\[\\tfrac\{1\}\{K\}\\\!\\sum\_\{k\}\\\!\\min\\\!\\big\(w\_\{k\}A\_\{k\},\\ \\mathrm\{clip\}\(w\_\{k\},\\,1\{\\pm\}\\epsilon\)A\_\{k\}\\big\)\\right\]\+\\beta\_\{\\mathrm\{KL\}\}D\_\{\\mathrm\{KL\}\}\(\\pi\_\{\\theta\}\\\|\\pi\_\{\\mathrm\{base\}\}\),with\(r¯,σr\)\(\\bar\{r\},\\sigma\_\{r\}\)the group mean/std,\(ϵlo,ϵhi\)=\(0\.20,0\.28\)\(\\epsilon\_\{\\mathrm\{lo\}\},\\epsilon\_\{\\mathrm\{hi\}\}\)\{=\}\(0\.20,0\.28\),βKL=10−3\\beta\_\{\\mathrm\{KL\}\}\{=\}10^\{\-3\},πbase=\\pi\_\{\\mathrm\{base\}\}\{=\}Qwen3\-VL\-8B\-Thinking BASE\. Cold\-start from base, KL anchor, MM\-Eureka\(Menget al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib33)\)difficulty curriculum \(drop0/N0/NandN/NN/Nprompts,∼22%\{\\sim\}22\\%filtered\),12,28812\{,\}288\-token CoT budget, and held\-out PhyX\-mini\-MC early stopping fix the joint setting \(Algorithm[3](https://arxiv.org/html/2605.14040#alg3), Table[10](https://arxiv.org/html/2605.14040#A8.T10)\); implementation usesverl0\.6\.1\(Shenget al\.,[2024](https://arxiv.org/html/2605.14040#bib.bib16)\)on Qwen3\-VL\-8B\-Thinking\(Qwen Team,[2025](https://arxiv.org/html/2605.14040#bib.bib17)\)with FSDP1 sharding \(§[6](https://arxiv.org/html/2605.14040#S6)\)\.
#### Two reward shapes: binary \(recommended\) vs\. dense \(ablation\)\.
Physics rollouts admit physics\-native per\-step signals—units, conservation, symbolic form—so a denser reward looks free\. We compare:
\(binary, recommended\)rbin\(y,x\)=𝟙\[Match\(ExtractBoxed\(y\),g\(x\)\)\]∈\{0,1\},\\displaystyle r\_\{\\mathrm\{bin\}\}\(y,x\)=\\mathbb\{1\}\\\!\\big\[\\textsc\{Match\}\(\\textsc\{ExtractBoxed\}\(y\),\\,g\(x\)\)\\big\]\\in\\\{0,1\\\},\(2\)\(dense, ablation\)rdense=clip\(rans\+rfmt\+rdim\+rsym\+rcons,−1,1\)\.\\displaystyle r\_\{\\mathrm\{dense\}\}=\\mathrm\{clip\}\\\!\\big\(r\_\{\\mathrm\{ans\}\}\+r\_\{\\mathrm\{fmt\}\}\+r\_\{\\mathrm\{dim\}\}\+r\_\{\\mathrm\{sym\}\}\+r\_\{\\mathrm\{cons\}\},\\,\-1,1\\big\)\.whereMatchaccepts MCQ\-letter equality,±1%\\pm 1\\%numeric tolerance, or symbolic equivalence \(Appendix[C\.1](https://arxiv.org/html/2605.14040#A3.SS1)\); the dense components arerans≡rbinr\_\{\\mathrm\{ans\}\}\{\\equiv\}r\_\{\\mathrm\{bin\}\},rfmt∈\{0,\+0\.1\}r\_\{\\mathrm\{fmt\}\}\{\\in\}\\\{0,\{\+\}0\.1\\\}\(\\boxed\{\}present\),rdim∈\{0,\+0\.15\}r\_\{\\mathrm\{dim\}\}\{\\in\}\\\{0,\{\+\}0\.15\\\}\(sympy\.physics\.units\),rsym∈\{0,\+0\.20\}r\_\{\\mathrm\{sym\}\}\{\\in\}\\\{0,\{\+\}0\.20\\\}\(\\fracsympifies\),rcons∈\{−0\.25,0\}r\_\{\\mathrm\{cons\}\}\{\\in\}\\\{\-0\.25,0\\\}\(energy/momentum violation; Appendix[C\.2](https://arxiv.org/html/2605.14040#A3.SS2)\)\. Under GSPO with group\-normalized advantages and a difficulty curriculum, four properties land binary as variance\-optimal and Goodhart\-robust \(full derivation in Appendix[C\.1](https://arxiv.org/html/2605.14040#A3.SS1.SSS0.Px1)\)\.\(P1\) Group normalization absorbs reward magnitude:AkA\_\{k\}is invariant to affine rescaling ofrrwithin a group, so dense only matters when it*reorders*rollouts—we measure14\.3%14\.3\\%of within\-group pairs flipped,87%87\\%inside the all\-wrong subgroup\.\(P2\) Wrong\-group reorderings are a Goodhart channel:rewarding well\-formatted\-but\-wrong above poorly\-formatted\-but\-wrong biases the policy toward LaTeX\-format proxies that transfer poorly to the audited held\-out eval\.\(P3\) Variance\-optimal advantage:on a Bernoulli reward,Var\(Abin\)=1\\mathrm\{Var\}\(A^\{\\mathrm\{bin\}\}\)\{=\}1saturates theKK\-sample bound; a bounded shaping termδk∈\[0,Δ\]\\delta\_\{k\}\\in\[0,\\Delta\]inflatesσr\\sigma\_\{r\}byO\(Δ2\)O\(\\Delta^\{2\}\), shrinking\|Acorrectdense\|\|A^\{\\mathrm\{dense\}\}\_\{\\mathrm\{correct\}\}\|below\|Acorrectbin\|\|A^\{\\mathrm\{bin\}\}\_\{\\mathrm\{correct\}\}\|\. Empirical signature at matched step 60 on the seed\-42 ablation \(Table[3](https://arxiv.org/html/2605.14040#S5.T3)\): binary beats dense by\+8\.9\+8\.9/\+4\.9\+4\.9/\+6\.4\+6\.4pp on PhysReason/OlymBench\-Phys\-liberal/PhysOlym\-A\-liberal while tied with dense on PUB\-OE \(−0\.7\-0\.7pp\) and trailing dense by at most0\.60\.6pp on saturated MCQ\. We ship binary as the deployable artifact; the per\-component drop\-out ablation \(Table[11](https://arxiv.org/html/2605.14040#A8.T11)\) is left to follow\-up work\.
## 5Evaluation
We organize this section in two parts\. §[5\.1](https://arxiv.org/html/2605.14040#S5.SS1)characterizesPhysOlym\-Aas a measurement instrument and grounds Findings 2–3 of §[1](https://arxiv.org/html/2605.14040#S1)\(Finding 1, audit\-leakage, is in §[3\.3](https://arxiv.org/html/2605.14040#S3.SS3.SSS0.Px1)\)\. §[5\.2](https://arxiv.org/html/2605.14040#S5.SS2)reports Physics\-R1 results to validatePhysCorp\-Aas trainable\. The Physics\-R1 \(binary, seed 42\) row in Table[3](https://arxiv.org/html/2605.14040#S5.T3)is the headline single\-seed checkpoint; the 3\-seed mean row aggregates seed 42 with two additional seeds \(seed\-17/step\-63 and seed\-23/step\-60\) on the auditedPhysR1Corpcorpus\.
#### Scoring protocol\.
All open\-ended columns of Table[3](https://arxiv.org/html/2605.14040#S5.T3)use*problem\-level*liberal Sonnet\-as\-judge accuracy \(Appendix[D](https://arxiv.org/html/2605.14040#A4)\): for multi\-sub\-part problems on PhysReason and PhysUniBench\-OE,llm\_judge\_v2\_alignment\.pyandllm\_judge\_v3\_pubeo\.pyrespectively call Sonnet 4\.5 once per gold sub\-answer with YES/NO, and the problem is judged correct only if every sub\-part is correct \(AND across sub\-parts\); OlympiadBench\-Physics andPhysOlym\-Ausejudge\_olympiad\.py, which makes a single YES/NO call per problem against the full gold solution\. The unjudgeable rate onPhysOlym\-Ais13\.9%13\.9\\%\(gold solutions consisting of grading rubrics, administrative notes, or figure\-only references\)\. Three layers bound judge optimism: strict vs\. liberal gap on Sonnet \(4\.74\.7pp onPhysOlym\-A\); inter\-judge Cohen’sκ\\kappa\(Cohen,[1960](https://arxiv.org/html/2605.14040#bib.bib20)\)between two Sonnet seeds; and a100100\-problem human\-graded random subset \(Appendix[D](https://arxiv.org/html/2605.14040#A4)\)\. The cross\-vendor judge agreement on a5050\-problem Sonnet/GPT\-4o pair test shows GPT\-4o is*more*lenient than Sonnet \(16%16\\%vs\.8%8\\%positive rate\), bounding self\-grading concern in the opposite direction from naive worry\.
#### Judging concurrency and reproducibility\.
All Sonnet\-judge runs reported in Table[3](https://arxiv.org/html/2605.14040#S5.T3)are executed atworkers=22–44concurrency to stay below Anthropic API rate limits; sub\-judgments that exceed the per\-call timeout are retried at lower concurrency rather than counted as wrong\. Per\-cell judge\-error counts \(typically0–1010out of629629–12001200records, all≤1%\\leq 1\\%\) and per\-record verdicts are released asjudge\_audit\.jsonin the supplementary archive\. The Sonnet 4\.5 PhysReason responses \(Table[3](https://arxiv.org/html/2605.14040#S5.T3),†\) are regenerated withmax\_tokens=16384\\texttt\{max\\\_tokens\}\{=\}16384to match the response budget used by all open\-source baselines and Physics\-R1; intermediate\-length Sonnet responses \(mean<200<200chars under default API settings\) systematically fail to commit a\\boxed\{\}final answer on multi\-sub\-part problems, which v2\_alignment scores as wrong\.
### 5\.1PhysOlym\-Aas a Measurement Instrument
#### Same\-model evaluation reveals a 46\-point format\-and\-novelty gradient\.
On identical Sonnet 4\.5 weights evaluated in the same week, the score sweeps from79\.7%79\.7\\%on PhyX \(4\-way MCQ\) down to50\.4%50\.4\\%liberal on OlympiadBench\-Physics and33\.4%33\.4\\%liberal onPhysOlym\-A—a 46\-point gradient on fixed weights, the strongest evidence the paper has for the central claim that physics evaluation is format\- and novelty\-bound \(Finding 3\)\. Three forces drive the drop: format \(4\-way MCQ vs\. open\-ended\), genre \(PhyX is K\-12 to early\-undergraduate, the bottom two are competition\-grade\), and contamination\-removal \(onlyPhysOlym\-Ais three\-stage audited against the Physics\-R1 training pool\)\. The PhyX→\\toOlympiadBench\-Physics step accounts for∼29\{\\sim\}29pp of the gradient \(dominated by format \+ genre, since both are public and not contamination\-cleaned\), and the OlympiadBench\-Physics→\\toPhysOlym\-Astep adds∼17\{\\sim\}17pp on top \(dominated by audit and novelty since both are open\-ended and competition\-grade\); a controlled 2×\\times2 \(format×\\timesaudit on identical items\) is left to follow\-up work to attribute the residual cleanly\.PhysOlym\-Asits at the bottom of this gradient by construction\. The per\-physics\-category breakdown on OlympiadBench\-Physics \(electromagnetism hardest at38\.4%38\.4\\%, astrophysics easiest at72\.9%72\.9\\%\) and the saturation\-gradient table are in Appendix[H\.7](https://arxiv.org/html/2605.14040#A8.SS7)\.
#### Difficulty\-stratified accuracy from organizer\-issued labels\.
The Estonian Physics Olympiad is the only public physics olympiad whose problems carry organizer\-issued difficulty labels \(1–10\) by construction, eliminating self\-annotation circularity\. On the 131 Estonian problems carrying native annotation, Sonnet 4\.5 strict accuracy decays near\-monotonically from62\.5%62\.5\\%at difficulty 1 to a hard0%0\\%floor at difficulties 3, 6, 8, and 10 \(full table: Appendix[H\.7](https://arxiv.org/html/2605.14040#A8.SS7), Table[12](https://arxiv.org/html/2605.14040#A8.T12)\)\. The trivial end of the Estonian olympiad \(62\.5%62\.5\\%at difficulty 1\) already lies below Sonnet’s PhyX score \(79\.7%79\.7\\%\); the clean zeros at four difficulty bins are the empirical signature of a non\-saturating benchmark, the property required forPhysOlym\-Ato serve as a stopping signal during Physics\-R1 training\.
#### Cross\-lingual ablation: 17\-point translation delta on identical problems\.
The Estonian PhO bilingual subset enables a controlled cross\-lingual experiment on 59 paired problems graded against the same gold by the same Sonnet 4\.5:30\.5%\\mathbf\{30\.5\}\\%strict on Estonian originals vs\.13\.6%\\mathbf\{13\.6\}\\%on English translations \(sign testp=0\.011p\{=\}0\.011; McNemarp=0\.021p\{=\}0\.021; paired bootstrap95%95\\%CI\[\+5\.1,\+28\.9\]\[\+5\.1,\+28\.9\]pp\)\. The per\-problem agreement matrix is asymmetric: 13 correct on Estonian but wrong on English, 3 the reverse\. For weaker open\-source 8B\-class models with limited Estonian training the gap is expected to flip in sign \(pre\-registered as a follow\-up direction; Appendix[H\.4](https://arxiv.org/html/2605.14040#A8.SS4), item \(viii\)\)\.
### 5\.2Physics\-R1
#### Recipe\.
Physics\-R1 is the GSPO\(Zhenget al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib12)\)\+DAPO\(Yuet al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib11)\)recipe of §[4](https://arxiv.org/html/2605.14040#S4)cold\-started from Qwen3\-VL\-8B\-Thinking BASE\(Qwen Team,[2025](https://arxiv.org/html/2605.14040#bib.bib17)\)onPhysR1Corp\(§[3\.1](https://arxiv.org/html/2605.14040#S3.SS1)\) under MM\-Eureka difficulty filtering\(Menget al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib33)\)and a binary correctness reward\.PhyX\-mini\-MC\(1,0001\{,\}000\-problem audit\-clean MCQ subset\(Shenet al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib1)\)\) is held out as the in\-training early\-stop signal: 4\-way MCQ gives a cleaner per\-step trajectory than open\-ended judging, and it is certified disjoint fromPhysR1Corpunder the audit pipeline of §[3\.3](https://arxiv.org/html/2605.14040#S3.SS3)\.
Table 3:Capability across MCQ, numerical, open\-ended, and held\-out olympiad benchmarks\.MCQ random baseline: PhyX\-1k/3k25%25\\%\. All open\-ended columns \(PhysReason, PUB\-OE, OlymBench\-Phys,PhysOlym\-A\) use*problem\-level*liberal Sonnet\-as\-judge accuracy under our v2/v3 judges \(Appendix[D](https://arxiv.org/html/2605.14040#A4)\): every sub\-question of a multi\-part problem must be judged correct for the problem to count\. All Sonnet\-judge runs useworkers=22–44concurrency for rate\-limit safety, with errored sub\-judgments retried at lower concurrency\. Sonnet PhysReason cell \(†\) is generated withmax\_tokens==16384 to match the protocol used by Physics\-R1 and the open\-source baselines\. GPT\-4o PhyX\-1k/3k fromShenet al\.\([2025](https://arxiv.org/html/2605.14040#bib.bib1)\); Gemini PhyX\-1k/3k cells \(∗\) measured here\. All Physics\-R1 evals usemax\_tokens==16384\.MCQOpen\-ended \(problem\-AND aggregation, liberal Sonnet\-judge\)ModelPhyX\-1kPhyX\-3kPhysReasonPUB\-OEOlymBenchPhysOlym\-AMCQ\-exactMCQ\-exactsubpart\-AND \(v2\)subpart\-AND \(v3\)problem\-lvlproblem\-lvl*Closed\-source frontier*Claude Sonnet 4\.579\.780\.649\.1†\\mathbf\{49\.1\}^\{\\dagger\}25\.425\.450\.4\\mathbf\{50\.4\}33\.4\\mathbf\{33\.4\}Gemini 2\.5 Pro75\.1∗49\.8∗38\.833\.437\.412\.2GPT\-4o70\.453\.651\.1¯\\underline\{51\.1\}31\.019\.719\.5*Open\-source bases \(best / second\-best across this block \+ ours:bold/underline\)*Qwen3\-VL\-32B\-Thinking73\.884\.225\.132\.853\.913\.2Qwen3\-VL\-8B\-Thinking \(base\)73\.774\.423\.935\.339\.38\.0InternVL3\-8B46\.843\.113\.323\.510\.44\.0*This work \(subscripts:Δ\\Deltavs\. Qwen3\-VL\-8B\-Thinking base\)*Physics\-R1 \(dense\)78\.3\+4\.6\\mathbf\{78\.3\}\_\{\+4\.6\}77\.5¯\+3\.1\\underline\{77\.5\}\_\{\+3\.1\}23\.3−0\.623\.3\_\{\-0\.6\}37\.7\+2\.4\\mathbf\{37\.7\}\_\{\+2\.4\}40\.5\+1\.240\.5\_\{\+1\.2\}19\.2¯\+11\.2\\underline\{19\.2\}\_\{\+11\.2\}Physics\-R1 \(binary, seed 42\)78\.0¯\+4\.3\\underline\{78\.0\}\_\{\+4\.3\}76\.9\+2\.576\.9\_\{\+2\.5\}32\.2¯\+8\.3\\underline\{32\.2\}\_\{\+8\.3\}37\.0¯\+1\.7\\underline\{37\.0\}\_\{\+1\.7\}45\.4¯\+6\.1\\underline\{45\.4\}\_\{\+6\.1\}25\.6\+17\.625\.6\_\{\+17\.6\}Physics\-R1 \(binary, 3\-seed mean±σ\\pm\\sigma\)‡77\.8±0\.3\+4\.177\.8\_\{\\pm 0\.3\}^\{\+4\.1\}76\.9±0\.3\+2\.576\.9\_\{\\pm 0\.3\}^\{\+2\.5\}39\.6±6\.4\+15\.7\\mathbf\{39\.6\}\_\{\\pm 6\.4\}^\{\+15\.7\}34\.8±3\.3−0\.534\.8\_\{\\pm 3\.3\}^\{\-0\.5\}46\.2±1\.5\+6\.9\\mathbf\{46\.2\}\_\{\\pm 1\.5\}^\{\+6\.9\}26\.3±1\.7\+18\.3\\mathbf\{26\.3\}\_\{\\pm 1\.7\}^\{\+18\.3\}
†Sonnet PhysReason regenerated atmax\_tokens=16384\\texttt\{max\\\_tokens\}=16384,1,192/1,2001\{,\}192/1\{,\}200clean records\.‡3\-seed mean over seeds\{42,17,23\}\\\{42,17,23\\\}on the auditedPhysR1Corp\(2,2682\{,\}268records, binary reward; seed\-42 also the dense\-ablation seed\)\. Per\-seed \(42/17/23\): PR32\.2/43\.1/43\.432\.2/43\.1/43\.4; PUB\-OE37\.0/36\.4/30\.937\.0/36\.4/30\.9; OlymBench\-Phys45\.4/45\.3/48\.045\.4/45\.3/48\.0;PhysOlym\-A25\.6/25\.0/28\.225\.6/25\.0/28\.2; PhyX\-mini78\.0/77\.4/77\.978\.0/77\.4/77\.9; PhyX\-3k76\.9/77\.2/76\.676\.9/77\.2/76\.6\.PhysOlym\-Alift decomposes into∼3\.5\\sim 3\.5pp from\\boxed\{\}\-emission rate \(33\.8%→64\.4%33\.8\\%\{\\to\}64\.4\\%from base to Physics\-R1\) and∼14\.1\\sim 14\.1pp from conditional accuracy \(22\.5→36\.322\.5\{\\to\}36\.3\); details in Appendix[H\.7](https://arxiv.org/html/2605.14040#A8.SS7)\.
#### Capability across formats and the held\-out olympiad column\.
Table[3](https://arxiv.org/html/2605.14040#S5.T3)reports Physics\-R1 alongside closed\-frontier and open\-source bases on three answer formats and the held\-out olympiad split\. The Qwen3\-VL\-8B\-Thinking BASE checkpoint attains73\.7%73\.7\\%on PhyX\-mini\-1k; the 32B sibling is indistinguishable on PhyX \(73\.8%73\.8\\%\), so scale alone in the 8B–32B Thinking band does not move PhyX\. All Physics\-R1 evals usemax\_tokens==16384 to permit full thinking\-mode CoT; eval\-budget sensitivity, harness\-canonical re\-evaluation, and per\-judge gap analysis are in Appendix[H\.7](https://arxiv.org/html/2605.14040#A8.SS7)\.
#### Physics\-R1 \(binary, recommended\): step 60 closes the audited held\-out gap\.
The binary\-reward checkpoint at step 60 lifts the 8B base across all formats \(Table[3](https://arxiv.org/html/2605.14040#S5.T3)\); the largest lift lands onPhysOlym\-Aliberal \(\+18\.3\+18\.3pp at 3\-seed mean\), with lifts on saturated public open\-ended splits substantially smaller—the contamination signal Finding 1 predicts: where the 8B base is already close to ceiling \(PUB\-OE35\.335\.3, OlymBench\-Phys39\.339\.3\), there is little headroom; where the eval is novel\-source and audited \(PhysOlym\-Aliberal8\.08\.0\), the post\-training lift is large\. The numerical/open\-ended jumps from step 40 to step 60 are driven by the additional2020GRPO steps lifting the\\boxed\{\}\-emission rate from46%46\\%to8787–96%96\\%\. The 3\-seed mean \(seed 42 \+ seed\-17/step\-63 \+ seed\-23/step\-60, all on the auditedPhysR1Corp\) is tight across most open\-ended columns: per\-seed PR\{32\.2,43\.1,43\.4\}\\\{32\.2,43\.1,43\.4\\\}\(mean39\.6±6\.439\.6\\pm 6\.4; seed\-42 outlier on PR,∼11\\sim 11pp below seeds 17/23 with otherwise comparable performance on other columns\),PhysOlym\-Aliberal\{25\.6,25\.0,28\.2\}\\\{25\.6,25\.0,28\.2\\\}\(mean26\.3±1\.726\.3\\pm 1\.7\), OlymBench\-Phys\{45\.4,45\.3,48\.0\}\\\{45\.4,45\.3,48\.0\\\}\(mean46\.2±1\.546\.2\\pm 1\.5\), PUB\-OE\{37\.0,36\.4,30\.9\}\\\{37\.0,36\.4,30\.9\\\}\(mean34\.8±3\.334\.8\\pm 3\.3\)\. The audited corpus is trainable and the lift over the 8B base is reproducible across seeds\.
#### Where Physics\-R1 helps: failure modes of the base it mitigates\.
The\+18\.3\+18\.3\-pp 3\-seed\-mean lift onPhysOlym\-Aliberal corresponds to∼92\{\\sim\}92problems flipped from wrong\-on\-base to correct\-on\-Physics\-R1 \(per\-seed range8585–101101\)\. Hand\-inspecting 30 such flips reveals three recurring failure modes of the 8B base, each addressed by a specific recipe lever: \(i\)*reasoning\-without\-committing*\(long correct CoT, no\\boxed\{\}final\) — fixed byrbinr\_\{\\mathrm\{bin\}\}\(§[4](https://arxiv.org/html/2605.14040#S4)\); \(ii\)*unit/dimensional shortcuts*\(dimensionally\-consistent but answer\-wrong\) — fixed by MM\-Eureka curriculum filtering ofN/NN/Nsurface\-heuristic prompts; \(iii\)*multi\-image evidence integration*\(base attends only to the first panel\) — fixed by the cold\-start from Qwen3\-VL\-8B\-Thinking BASE under FSDP1, which preserves the visual encoder\. Physics\-R1 does*not*fix genuine physics\-content gaps \(graduate\-level Tripos\-style perturbation theory remains wrong on both\)\. Full transcripts in Appendix[H\.1](https://arxiv.org/html/2605.14040#A8.SS1)\.
#### PhysOlym\-Agrounds the central training\-utility claim\.
ThePhysOlym\-A\-liberal column of Table[3](https://arxiv.org/html/2605.14040#S5.T3)is the cleanest non\-saturating capability signal in our 7\-axis comparison \(Table[1](https://arxiv.org/html/2605.14040#S2.T1)\)\. Sonnet attains33\.4%33\.4\\%; Physics\-R1 binary at the 3\-seed mean reaches26\.3±1\.7%\\mathbf\{26\.3\\pm 1\.7\}\\%\(per\-seed\{25\.6,25\.0,28\.2\}\\\{25\.6,25\.0,28\.2\\\}across seeds 42/17/23\), exceeding every open\-source baseline \(Qwen3\-VL\-32B13\.2%13\.2\\%, 8B8\.0%8\.0\\%, InternVL34\.0%4\.0\\%\) and the non\-Sonnet closed APIs \(GPT\-4o19\.5%19\.5\\%, Gemini 2\.5 Pro12\.2%12\.2\\%\), trailing only Sonnet by7\.17\.1pp\.
#### Reward\-shape ablation: dense gives a small saturated\-MCQ benefit, binary wins on open\-ended\.
Under problem\-level liberal Sonnet\-judge scoring \(seed 42\), dense at step 60 slightly leads on saturated MCQ \(PhyX\-mini78\.378\.3vs\. binary78\.078\.0; PhyX\-3k77\.577\.5vs\.76\.976\.9\) but trails binary on every non\-MCQ split: PhysReason \(23\.323\.3vs\.32\.2\\mathbf\{32\.2\},\+8\.9\+8\.9pp\), OlympiadBench\-Physics \(40\.540\.5vs\.45\.4\\mathbf\{45\.4\},\+4\.9\+4\.9pp\), andPhysOlym\-Aliberal \(19\.219\.2vs\.25\.6\\mathbf\{25\.6\},\+6\.4\+6\.4pp\)\. On PUB\-OE, dense and binary are within0\.70\.7pp \(37\.737\.7vs\.37\.037\.0\)\. The binary advantage is concentrated on the multi\-sub\-part numerical\-answer column \(PR\) and the held\-out audited olympiad column \(PhysOlym\-A\)—the two columns where the recipe contribution should matter most\. Dense is a reward\-shape ablation; binary is the recommended default\. The five\-component reward drop\-out \(Table[11](https://arxiv.org/html/2605.14040#A8.T11)\), recipe\-flag, and SFT data\-scaling ablations are left to follow\-up work\.
## 6Discussion and Limitations
Physics\-R1 uses unmodified GSPO\+DAPO; the dense reward and recommended baseline \(Algorithm[3](https://arxiv.org/html/2605.14040#alg3)\) are reproducibility artifacts, not method claims\. The audit pipeline catches verbatim and lightly\-paraphrased duplicates and is empirically robust to both embedder choice \(Spearmanρ=0\.78\\rho\{=\}0\.78vs\.text\-embedding\-3\-large; OpenAI candidate set is a strict subset of mxbai’s at every threshold tested; Appendix[A](https://arxiv.org/html/2605.14040#A1)\) and judge choice \(Sonnet 4\.5 vs\. GPT\-4o cross\-judgeκ=0\.44\\kappa\{=\}0\.44on a 50\-problemPhysOlym\-Asubset, with GPT\-4o more lenient—self\-grading direction is opposite to the feared bias; Appendix[D](https://arxiv.org/html/2605.14040#A4)\); the Sonnet\-as\-judge13\.9%13\.9\\%unjudgeable rate is a disclosed noise floor\. The cross\-lingual finding is Sonnet\-4\.5\-specific onn=59n\{=\}59paired items \(65\.7%65\.7\\%MC power, all three tests rejectH0H\_\{0\}\); direction is pre\-registered to reverse for cross\-lingual\-weak models \(Appendix[H\.4](https://arxiv.org/html/2605.14040#A8.SS4), item viii\)\.
## 7Conclusion
Three findings—𝟏𝟑𝟒\\mathbf\{134\}near\-duplicates in SciInstruct surfaced only by three\-stage audit \(Jaccard→\\tocosine→\\toHaiku\-4\.5 judge\), a1717\-pp Estonian–English translation delta on identical olympiad problems, and a4646\-pp format\-and\-novelty gradient on fixed Sonnet 4\.5 weights—motivate four released artifacts:PhysCorp\-A\(6,4326\{,\}432\-record audited corpus, fully Stage\-3 clean against all six public physics evals; Table[2](https://arxiv.org/html/2605.14040#S3.T2)\),PhysR1Corp\(2,2682\{,\}268\-record closed\-form RL pool\),PhysOlym\-A\(500500\-problem held\-out olympiad eval,99\.8%99\.8\\%novel\-source\), and Physics\-R1, a binary\-reward GSPO\+DAPO recipe that liftsPhysOlym\-Aliberal\+18\.3\+18\.3pp over the 8B base at the 3\-seed mean \(8\.0→26\.3±1\.78\.0\{\\to\}\\mathbf\{26\.3\\pm 1\.7\}, still7\.17\.1pp below Sonnet 4\.5; per\-seed\{25\.6,25\.0,28\.2\}\\\{25\.6,25\.0,28\.2\\\}across seeds\{42,17,23\}\\\{42,17,23\\\}on the auditedPhysR1Corp\)\. Audit\-pipeline robustness to embedder and judge choice is established in §[6](https://arxiv.org/html/2605.14040#S6)\(Appendices[A](https://arxiv.org/html/2605.14040#A1),[D](https://arxiv.org/html/2605.14040#A4)\)\. We recommend binary correctness reward as the deployable default \(variance\-optimal under GSPO with group\-normalized advantages, Goodhart\-robust against unit/format proxies; §[4](https://arxiv.org/html/2605.14040#S4)\); reward\-component drop\-out \(Table[11](https://arxiv.org/html/2605.14040#A8.T11)\) is left to follow\-up work; the 3\-seed mean reported in Table[3](https://arxiv.org/html/2605.14040#S5.T3)\(σ≤3\.3\\sigma\\leq 3\.3pp on PUB\-OE, OlymBench\-Phys, andPhysOlym\-A;σ=6\.4\\sigma\{=\}6\.4pp on PhysReason driven by a seed\-42 outlier\) demonstrates that the audited corpus retains training signal across seeds\.
## Acknowledgments
The author thanks Kevin Zhou for granting permission to redistribute his olympiad handouts under CC BY\-NC 4\.0, the Estonian Physics Olympiad committee for making their archived problems and solutions publicly available at[https://fyysika\.ee/](https://fyysika.ee/), the international olympiad committees \(IPhO, NBPhO, EuPhO, APhO, USAPhO, INPhO\) for the public archives that enabled the novel\-source held\-out evaluation, and the maintainers of the public physics\-VL benchmarks \(PhyX, MMMU\-Pro, OlympiadBench, UGPhysics, PhysReason, PhysUniBench\) whose released training pools and evals enabled the contamination audit reported in this work\. Compute support for Physics\-R1 training and evaluation was provided by RunPod\.
## References
- S\. Ahuja, V\. Gumma, and S\. Sitaram \(2024\)Contamination report for multilingual benchmarks\.Note:Tests 7 LLMs on multilingual benchmarks; finds nearly all are contaminated\.External Links:2410\.16186Cited by:[§2](https://arxiv.org/html/2605.14040#S2.SS0.SSS0.Px3.p1.1)\.
- M\. Akhtar, O\. Benjelloun, C\. Conforti,et al\.\(2024\)Croissant: a metadata format for ML\-ready datasets\.InNeurIPS Workshop on Data\-centric Machine Learning Research,External Links:[Link](https://arxiv.org/abs/2403.19546)Cited by:[§G\.7](https://arxiv.org/html/2605.14040#A7.SS7.SSS0.Px1.p1.1)\.
- American Association of Physics Teachers \(2025\)U\.s\. physics olympiad: archived problems and solutions\.Note:[https://www\.aapt\.org/physicsteam/](https://www.aapt.org/physicsteam/)Cited by:[§3\.1](https://arxiv.org/html/2605.14040#S3.SS1.p1.4)\.
- Anthropic \(2025\)Claude sonnet 4\.5\.Note:Model used as judge and frontier\-baseline reference throughout this paper\.External Links:[Link](https://www.anthropic.com/news/claude-sonnet-4-5)Cited by:[§1](https://arxiv.org/html/2605.14040#S1.SS0.SSS0.Px2.p1.6)\.
- Asian Physics Olympiad Committee \(2025\)Asian physics olympiad: archived problems and solutions\.Note:[https://apho2025\.fkfi\.lt/](https://apho2025.fkfi.lt/)Cited by:[§3\.1](https://arxiv.org/html/2605.14040#S3.SS1.p1.4)\.
- W\. Chow, J\. Mao, B\. Li, D\. Seita, V\. Guizilini, and Y\. Wang \(2025\)PhysBench: benchmarking and enhancing vision\-language models for physical world understanding\.InICLR,Note:Oral;[https://openreview\.net/forum?id=Q6a9W6kzv5](https://openreview.net/forum?id=Q6a9W6kzv5)Cited by:[§2](https://arxiv.org/html/2605.14040#S2.SS0.SSS0.Px3.p1.1)\.
- J\. Cohen \(1960\)A coefficient of agreement for nominal scales\.Educational and Psychological Measurement20\(1\),pp\. 37–46\.Cited by:[§5](https://arxiv.org/html/2605.14040#S5.SS0.SSS0.Px1.p1.7)\.
- DeepSeek\-AI \(2025\)DeepSeek\-r1: incentivizing reasoning capability in llms via reinforcement learning\.arXiv preprint arXiv:2501\.12948\.Cited by:[§1](https://arxiv.org/html/2605.14040#S1.p2.1),[§2](https://arxiv.org/html/2605.14040#S2.SS0.SSS0.Px1.p1.1)\.
- J\. Dekoninck, M\. N\. Mueller, and M\. Vechev \(2024\)ConStat: performance\-based contamination detection in large language models\.InNeurIPS,Cited by:[§2](https://arxiv.org/html/2605.14040#S2.SS0.SSS0.Px3.p1.1)\.
- Estonian Physics Olympiad \(2018\)Estonian physics olympiad: problem collection 2004–2018\.Note:[https://www\.fyysika\.ee/](https://www.fyysika.ee/)Cited by:[§3\.1](https://arxiv.org/html/2605.14040#S3.SS1.p1.4)\.
- EuPhO Committee \(2025\)European physics olympiad: archived problems and solutions\.Note:[https://eupho\.ee/](https://eupho.ee/)Cited by:[§3\.1](https://arxiv.org/html/2605.14040#S3.SS1.p1.4)\.
- T\. Gebru, J\. Morgenstern, B\. Vecchione, J\. W\. Vaughan, H\. Wallach, H\. Daume III, and K\. Crawford \(2021\)Datasheets for datasets\.Communications of the ACM\.Cited by:[Appendix G](https://arxiv.org/html/2605.14040#A7.p1.1)\.
- E\. Glazer, E\. Erdil, T\. Besiroglu, D\. Chicharro,et al\.\(2024\)FrontierMath: a benchmark for evaluating advanced mathematical reasoning in AI\.arXiv preprint arXiv:2411\.04872\.Cited by:[§2](https://arxiv.org/html/2605.14040#S2.SS0.SSS0.Px3.p1.1),[Table 1](https://arxiv.org/html/2605.14040#S2.T1.22.14.3)\.
- C\. He, R\. Luo, Y\. Bai, S\. Hu,et al\.\(2024\)OlympiadBench: a challenging benchmark for promoting AGI with olympiad\-level bilingual multimodal scientific problems\.InACL,Cited by:[§1](https://arxiv.org/html/2605.14040#S1.SS0.SSS0.Px3.p1.4),[§2](https://arxiv.org/html/2605.14040#S2.SS0.SSS0.Px2.p1.1),[Table 1](https://arxiv.org/html/2605.14040#S2.T1.16.8.3)\.
- Homi Bhabha Centre for Science Education \(2025\)Indian national physics olympiad: archived problems and solutions\.Note:[https://olympiads\.hbcse\.tifr\.res\.in/](https://olympiads.hbcse.tifr.res.in/)Cited by:[§3\.1](https://arxiv.org/html/2605.14040#S3.SS1.p1.4)\.
- International Physics Olympiad \(2025\)International physics olympiad: archived problems and solutions\.Note:[https://ipho\-unofficial\.org/](https://ipho-unofficial.org/)Cited by:[§3\.1](https://arxiv.org/html/2605.14040#S3.SS1.p1.4)\.
- S\. Lee, A\. Shakir, D\. Koenig, and J\. Lipp \(2024\)Open source strikes bread \- new fluffy embeddings model\.Mixedbread\.Note:mxbai\-embed\-large\-v1: 1024\-dim sentence embedding model used in our Stage\-2 audit\.External Links:[Link](https://www.mixedbread.com/blog/mxbai-embed-large-v1)Cited by:[Appendix A](https://arxiv.org/html/2605.14040#A1.SS0.SSS0.Px2.p2.3)\.
- F\. Meng, L\. Du, Z\. Liu, Z\. Zhou, Q\. Lu, D\. Fu, B\. Han, B\. Shi, W\. Wang, J\. He, K\. Zhang, T\. Zhao, Y\. Qiao, and P\. Luo \(2025\)MM\-eureka: exploring visual aha\-moment with rule\-based large\-scale reinforcement learning\.arXiv preprint arXiv:2503\.07365\.Cited by:[Table 8](https://arxiv.org/html/2605.14040#A8.T8.9.5.3),[§1](https://arxiv.org/html/2605.14040#S1.p2.1),[§2](https://arxiv.org/html/2605.14040#S2.SS0.SSS0.Px1.p1.1),[§2](https://arxiv.org/html/2605.14040#S2.SS0.SSS0.Px2.p1.1),[§4](https://arxiv.org/html/2605.14040#S4.p1.12),[§5\.2](https://arxiv.org/html/2605.14040#S5.SS2.SSS0.Px1.p1.1)\.
- NBPhO Committee \(2025\)Nordic\-baltic physics olympiad: archived problems and solutions\.Note:[https://www\.nbpho\.eu/](https://www.nbpho.eu/)Cited by:[§3\.1](https://arxiv.org/html/2605.14040#S3.SS1.p1.4)\.
- OpenStax \(2024\)College Physics 2e and University Physics Volumes 1\-3\.Note:[https://openstax\.org/subjects/science](https://openstax.org/subjects/science)Cited by:[§3\.1](https://arxiv.org/html/2605.14040#S3.SS1.p1.4)\.
- L\. Phan, A\. Gatti, Z\. Han, N\. Li, J\. Hu, H\. Zhang, C\. B\. C\. Zhang, M\. Shaaban, J\. Ling, S\. Shi,et al\.\(2025\)Humanity’s last exam\.External Links:2501\.14249Cited by:[§2](https://arxiv.org/html/2605.14040#S2.SS0.SSS0.Px3.p1.1),[Table 1](https://arxiv.org/html/2605.14040#S2.T1.24.16.3)\.
- S\. Qiu, S\. Guo, Z\. Song, Y\. Sun, Z\. Cai, J\. Wei, T\. Luo, Y\. Yin, H\. Zhang, Y\. Hu,et al\.\(2025\)PHYBench: holistic evaluation of physical perception and reasoning in LLMs\.InNeurIPS Datasets and Benchmarks Track,Note:[https://openreview\.net/forum?id=brG8FPq1cf](https://openreview.net/forum?id=brG8FPq1cf)Cited by:[§2](https://arxiv.org/html/2605.14040#S2.SS0.SSS0.Px2.p1.1),[Table 1](https://arxiv.org/html/2605.14040#S2.T1.9.1.2)\.
- Qwen Team \(2025\)Qwen3\-VL\.Note:[https://huggingface\.co/Qwen/Qwen3\-VL\-8B\-Thinking](https://huggingface.co/Qwen/Qwen3-VL-8B-Thinking)Vision\-language model releaseCited by:[§1](https://arxiv.org/html/2605.14040#S1.p2.1),[§4](https://arxiv.org/html/2605.14040#S4.p1.12),[§5\.2](https://arxiv.org/html/2605.14040#S5.SS2.SSS0.Px1.p1.1)\.
- R\. Rafailov, A\. Sharma, E\. Mitchell, C\. D\. Manning, S\. Ermon, and C\. Finn \(2023\)Direct preference optimization: your language model is secretly a reward model\.InNeurIPS,Cited by:[Table 8](https://arxiv.org/html/2605.14040#A8.T8.16.14.1)\.
- M\. Ravaut, B\. Ding, F\. Jiao,et al\.\(2024\)A comprehensive survey of contamination detection methods in large language models\.External Links:2404\.00699Cited by:[§2](https://arxiv.org/html/2605.14040#S2.SS0.SSS0.Px3.p1.1)\.
- O\. Sainz, J\. A\. Campos, I\. García\-Ferrero,et al\.\(2023\)NLP evaluation in trouble: on the need to measure LLM data contamination for each benchmark\.InEMNLP Findings,Cited by:[§2](https://arxiv.org/html/2605.14040#S2.SS0.SSS0.Px3.p1.1)\.
- Z\. Shao, P\. Wang, Q\. Zhu,et al\.\(2024\)DeepSeekMath: pushing the limits of mathematical reasoning in open language models\.arXiv preprint arXiv:2402\.03300\.Cited by:[Table 8](https://arxiv.org/html/2605.14040#A8.T8.5.1.2),[§1](https://arxiv.org/html/2605.14040#S1.p2.1),[§2](https://arxiv.org/html/2605.14040#S2.SS0.SSS0.Px1.p1.1)\.
- H\. Shen, T\. Wu, Q\. Han,et al\.\(2025\)PhyX: does your model have the “wits” for physical reasoning?\.arXiv preprint arXiv:2505\.15929\.Cited by:[§H\.3](https://arxiv.org/html/2605.14040#A8.SS3.p1.3),[§1](https://arxiv.org/html/2605.14040#S1.SS0.SSS0.Px3.p1.4),[§2](https://arxiv.org/html/2605.14040#S2.SS0.SSS0.Px2.p1.1),[§5\.2](https://arxiv.org/html/2605.14040#S5.SS2.SSS0.Px1.p1.1),[Table 3](https://arxiv.org/html/2605.14040#S5.T3)\.
- G\. Sheng, C\. Zhang, Z\. Ye, X\. Wu,et al\.\(2024\)HybridFlow: a flexible and efficient RLHF framework \(verl\)\.arXiv preprint arXiv:2409\.19256\.Cited by:[§4](https://arxiv.org/html/2605.14040#S4.p1.12)\.
- A\. K\. Singh, M\. Y\. Kocyigit, A\. Poulton,et al\.\(2024\)Evaluation data contamination in LLMs: how do we measure it and \(when\) does it matter?\.arXiv preprint arXiv:2411\.03923\.Cited by:[§2](https://arxiv.org/html/2605.14040#S2.SS0.SSS0.Px3.p1.1)\.
- Stack Exchange Inc\. \(2024\)Physics stack exchange: question and answer archive\.Note:[https://physics\.stackexchange\.com/](https://physics.stackexchange.com/)Cited by:[§3\.1](https://arxiv.org/html/2605.14040#S3.SS1.p1.4)\.
- G\. Tsoukalas, J\. Lee, J\. Jennings, Y\. Xin,et al\.\(2024\)PutnamBench: evaluating neural theorem\-provers on the putnam mathematical competition\.NeurIPS Datasets & Benchmarks\.Cited by:[§2](https://arxiv.org/html/2605.14040#S2.SS0.SSS0.Px3.p1.1),[Table 1](https://arxiv.org/html/2605.14040#S2.T1.18.10.3)\.
- C\. J\. Wang, D\. Lee, C\. Menghini,et al\.\(2025a\)EnigmaEval: a benchmark of long multimodal reasoning challenges\.Note:Frontier\-model pass\-rate so low that contamination is empirically dismissed\.External Links:2502\.08859Cited by:[§2](https://arxiv.org/html/2605.14040#S2.SS0.SSS0.Px3.p1.1)\.
- L\. Wang, E\. Su, J\. Liu,et al\.\(2025b\)PhysUniBench: a multi\-modal physics reasoning benchmark at undergraduate level\.External Links:2506\.17667Cited by:[§2](https://arxiv.org/html/2605.14040#S2.SS0.SSS0.Px2.p1.1),[Table 1](https://arxiv.org/html/2605.14040#S2.T1.12.4.4)\.
- Y\. Wang, X\. Ma, G\. Zhang, Y\. Ni, A\. Chandra, S\. Guo, W\. Ren, A\. Arulraj, X\. He, Z\. Jiang,et al\.\(2024\)MMLU\-Pro: a more robust and challenging multi\-task language understanding benchmark\.InNeurIPS Datasets and Benchmarks Track,Note:Spotlight;[https://openreview\.net/forum?id=y10DM6R2r3](https://openreview.net/forum?id=y10DM6R2r3)Cited by:[Table 1](https://arxiv.org/html/2605.14040#S2.T1.25.17.2)\.
- M\. Wu, W\. Wang, S\. Liu,et al\.\(2025\)The bitter lesson learned from 2,000\+ multilingual benchmarks\.External Links:2504\.15521Cited by:[§2](https://arxiv.org/html/2605.14040#S2.SS0.SSS0.Px3.p1.1)\.
- X\. Xu, Q\. Xu, T\. Xiao,et al\.\(2025\)UGPhysics: a comprehensive benchmark for undergraduate physics reasoning with large language models\.arXiv preprint arXiv:2502\.00334\.Cited by:[Appendix E](https://arxiv.org/html/2605.14040#A5.SS0.SSS0.Px1.p1.1),[§2](https://arxiv.org/html/2605.14040#S2.SS0.SSS0.Px2.p1.1),[Table 1](https://arxiv.org/html/2605.14040#S2.T1.41.36.1),[§3\.1](https://arxiv.org/html/2605.14040#S3.SS1.p1.4)\.
- W\. Xuan, R\. Yang, H\. Qi, Q\. Zeng, Y\. Xiao, Y\. Feng, J\. Liu, J\. Hou, J\. Zhao, W\. Yu,et al\.\(2025\)MMLU\-ProX: a multilingual benchmark for advanced reasoning across languages\.Note:29 languages, 11,829 identical questions per language\.External Links:2503\.10497Cited by:[§2](https://arxiv.org/html/2605.14040#S2.SS0.SSS0.Px3.p1.1)\.
- S\. Yang, W\. Chiang, L\. Zheng, J\. E\. Gonzalez, and I\. Stoica \(2023\)Rethinking benchmark and contamination for language models with rephrased samples\.Note:Demonstrates that n\-gram contamination audits miss rephrased duplicates; motivates Stage\-2 of our pipeline\.External Links:2311\.04850Cited by:[§2](https://arxiv.org/html/2605.14040#S2.SS0.SSS0.Px3.p1.1)\.
- Q\. Yu, Z\. Zhang, R\. Zhu,et al\.\(2025\)DAPO: an open\-source LLM reinforcement learning system at scale\.arXiv preprint arXiv:2503\.14476\.Cited by:[Table 8](https://arxiv.org/html/2605.14040#A8.T8.7.3.3),[§1](https://arxiv.org/html/2605.14040#S1.p2.1),[§2](https://arxiv.org/html/2605.14040#S2.SS0.SSS0.Px1.p1.1),[§4](https://arxiv.org/html/2605.14040#S4.p1.4),[§5\.2](https://arxiv.org/html/2605.14040#S5.SS2.SSS0.Px1.p1.1)\.
- X\. Yue, Y\. Ni, K\. Zhang, T\. Zheng,et al\.\(2024a\)MMMU: a massive multi\-discipline multimodal understanding and reasoning benchmark for expert AGI\.InCVPR,Cited by:[Appendix E](https://arxiv.org/html/2605.14040#A5.SS0.SSS0.Px4.p1.1),[§2](https://arxiv.org/html/2605.14040#S2.SS0.SSS0.Px2.p1.1),[§3\.1](https://arxiv.org/html/2605.14040#S3.SS1.p1.4)\.
- X\. Yue, T\. Zheng, Y\. Ni, Y\. Wang, K\. Zhang, S\. Tong, Y\. Sun, B\. Yu, G\. Zhang, H\. Sun,et al\.\(2024b\)MMMU\-Pro: a more robust multi\-discipline multimodal understanding benchmark\.InarXiv preprint,External Links:2409\.02813Cited by:[§2](https://arxiv.org/html/2605.14040#S2.SS0.SSS0.Px2.p1.1),[Table 1](https://arxiv.org/html/2605.14040#S2.T1.27.19.3)\.
- D\. Zhang, Z\. Hu, S\. Zhoubian, Z\. Du, K\. Yang, Z\. Wang, Y\. Yue, Y\. Dong, and J\. Tang \(2024\)SciInstruct: a self\-reflective instruction annotated dataset for training scientific language models\.InNeurIPS Datasets and Benchmarks Track,Note:[https://openreview\.net/forum?id=LC1QAqhePv](https://openreview.net/forum?id=LC1QAqhePv)Cited by:[Table 1](https://arxiv.org/html/2605.14040#S2.T1.41.39.1)\.
- X\. Zhang, Y\. Dong, Y\. Wu,et al\.\(2025\)PhysReason: a comprehensive benchmark towards physics\-based reasoning\.arXiv preprint arXiv:2502\.12054\.Cited by:[Appendix E](https://arxiv.org/html/2605.14040#A5.SS0.SSS0.Px5.p1.1),[§2](https://arxiv.org/html/2605.14040#S2.SS0.SSS0.Px2.p1.1),[Table 1](https://arxiv.org/html/2605.14040#S2.T1.14.6.3),[§3\.1](https://arxiv.org/html/2605.14040#S3.SS1.p1.4)\.
- C\. Zheng, S\. Liu, M\. Li,et al\.\(2025\)Group sequence policy optimization\.arXiv preprint arXiv:2507\.18071\.Cited by:[Table 8](https://arxiv.org/html/2605.14040#A8.T8.7.3.3),[§1](https://arxiv.org/html/2605.14040#S1.p2.1),[§2](https://arxiv.org/html/2605.14040#S2.SS0.SSS0.Px1.p1.1),[§4](https://arxiv.org/html/2605.14040#S4.p1.4),[§5\.2](https://arxiv.org/html/2605.14040#S5.SS2.SSS0.Px1.p1.1)\.
- K\. Zhou \(2018\)Olympiad physics handouts\.Note:[https://knzhou\.github\.io/](https://knzhou.github.io/)Cited by:[§3\.1](https://arxiv.org/html/2605.14040#S3.SS1.p1.4)\.
- Y\. Zhu, J\. Wang, Y\. Li,et al\.\(2025\)OIBench: benchmarking strong reasoning models with olympiad in informatics\.InNeurIPS Datasets and Benchmarks Track,Note:arXiv:2506\.10481Cited by:[Table 1](https://arxiv.org/html/2605.14040#S2.T1.20.12.3)\.
## Appendix AAudit pipeline details and worked examples
#### Stage\-3 LLM\-judge: SciInstruct cosine\-bucket near\-duplicate rate\.
For each of the4,8464\{,\}846SciInstruct↔\\leftrightarroweval Stage\-2 candidate pairs \(cos≥0\.85\\geq 0\.85\), the Stage\-3 Haiku\-4\.5 judge receives both problem statements and returns*close duplicate*\(paraphrase / numeric variation of the same problem\) or*same\-topic neighbor*\(related physics, distinct setup\)\. The close\-duplicate share is sharply threshold\-driven: atcos≥0\.95\\cos\{\\geq\}0\.95every flagged pair is a close duplicate; at the threshold edge\[0\.85,0\.87\)\[0\.85,0\.87\)only1\.5%1\.5\\%are\. Stage\-2 still surfaces genuinely close physics content at the low\-cos end—the topic overlap is real, just not strict duplication\.
Cosine bucketN pairsclose duplicatesame\-topic neighbor% close\-dup\[0\.95,0\.99\)\[0\.95,0\.99\)1717𝟏𝟕\\mathbf\{17\}0100\.0%\\mathbf\{100\.0\\%\}\[0\.90,0\.95\)\[0\.90,0\.95\)12712710101171177\.9%7\.9\\%\[0\.87,0\.90\)\[0\.87,0\.90\)1,1591\{,\}15954541,1051\{,\}1054\.7%4\.7\\%\[0\.85,0\.87\)\[0\.85,0\.87\)3,5433\{,\}54353533,4903\{,\}4901\.5%1\.5\\%Total𝟒,𝟖𝟒𝟔\\mathbf\{4\{,\}846\}𝟏𝟑𝟒\\mathbf\{134\}𝟒,𝟕𝟏𝟐\\mathbf\{4\{,\}712\}2\.8%\\mathbf\{2\.8\\%\}The pattern matches the threshold\-sensitivity table \(Table[4](https://arxiv.org/html/2605.14040#A1.T4)\): cosine≥0\.95\\geq 0\.95is precision\-dominant,≥0\.85\\geq 0\.85is recall\-dominant; Stage\-3 is the precision filter that converts a recall\-dominant Stage\-2 candidate set into a high\-precision near\-duplicate set\. Per\-eval Stage\-3 near\-duplicate counts \(out of Stage\-2 raw, matching Table[2](https://arxiv.org/html/2605.14040#S3.T2)\): PhysReason\-full36/2,68736/2\{,\}687\(1\.3%1\.3\\%\), PhysUniBench\-en22/1,02722/1\{,\}027\(2\.1%2\.1\\%\), PhyX\-mini46/70346/703\(6\.5%6\.5\\%\), OlymBench\-Phys15/13015/130\(11\.5%11\.5\\%\),PhysOlym\-A8/1638/163\(4\.9%4\.9\\%\), MMMU\-Pro7/1417/141\(5\.0%5\.0\\%\)\.
#### External\-corpora↔\\leftrightarrowheld\-out pairwise audit \(Stage 2, OlympiadBench shared\-source\)\.
The Stage\-1 / Stage\-2 cells of the main contamination matrix \(Table[2](https://arxiv.org/html/2605.14040#S3.T2)\) cover competitor training pools against the held\-out evals\. The complementary cross\-channel audit below pairs the four public physics\-olympiad benchmarks againstPhysOlym\-Aand PhyX 1000q, exposing shared\-source paraphrase overlap between Olympiad\-style problems composed independently\. The single Stage\-1 hit*OlympiadBench\-Physics*→\\toPhysOlym\-Ais the EuPhO 2020 “Mechanical accelerator” problem that grounds the99\.8%99\.8\\%\(rather than100%100\\%\) novel\-source claim\.
PhysOlym\-APhyX 1000qOlympiadBench\-Physics \(692\)1 / 1360 / 2PhysReason\-mini \(200\)0 / 2—PhysReason\-full \(1,200\)0 / 35—PhysUniBench\-en \(1,022\)0 / 27—
Stage 1 \(5\-gram Jaccard\)\. Tokenize each problem statement with a unicode word tokenizer; build the 5\-gram shingle set; compute Jaccard similarity over shingle sets; flag pairs with similarity≥0\.4\\geq 0\.4\. The threshold is calibrated against worked examples: an OpenStax College Physics↔\\leftrightarrowUniversity Physics duplicate scores Jaccard 1\.0; a UGPhysics paraphrase missed by Stage 1 \(Jaccard 0\.31\) is flagged at Stage 2 \(cosine 0\.91\); plausible misses include numerical\-value substitution \(same setup, different constants\) and translation across languages\. Stage 2 \(embedding cosine\)\. Encode each problem statement withmxbai\-embed\-large\-v1\[Leeet al\.,[2024](https://arxiv.org/html/2605.14040#bib.bib52)\]\(1024\-dim normalized embeddings\); compute pairwise cosine; flag pairs with cosine≥0\.85\\geq 0\.85\.
### A\.1Threshold\-Sensitive Leakage Finding on a Researcher\-Curated Baseline
Table 4:Threshold\-sensitivity of the leakage finding on a researcher\-curated baseline\.A 1,679\-record sample drawn fromPhysCorp\-pre\-audit\(the released 14,294\-record pre\-audit pool\) under conventional 5\-gram\-Jaccard \+ within\-pool embedding dedup is paired against a 500\-record internal analysis eval; each cell reports records flagged atJ≥jthrJ\\geq j\_\{\\text\{thr\}\}*or*cos≥cthr\\cos\\geq c\_\{\\text\{thr\}\}\. The boxed cell \(J≥0\.4J\\geq 0\.4,cos≥0\.85\\cos\\geq 0\.85\) is the operating point referenced throughout the paper\. Stage\-1 alone \(Jaccard\) is bimodal: all 56 leaks are exact atJ=1\.0J\{=\}1\.0\. Stage\-2 \(cosine\) catches an additional∼92\{\\sim\}92paraphrases the n\-gram audit misses at the published threshold; loosening cosine to0\.800\.80pushes the leak rate above27%27\\%\.cosine thresholdJ≥0\.3J\\geq 0\.3J≥0\.4J\\geq 0\.4J≥0\.5J\\geq 0\.5cos≥0\.80\\cos\\geq 0\.80\(lax\)455 \(27\.1%\)455 \(27\.1%\)455 \(27\.1%\)cos≥0\.85\\cos\\geq 0\.85\(paper op\.\)148 \(8\.8%\)148 \(8\.8%\)148 \(8\.8%\)cos≥0\.90\\cos\\geq 0\.90\(strict\)79 \(4\.7%\)79 \(4\.7%\)79 \(4\.7%\)*Single\-stage rates \(no union\):*Stage\-1 only \(J≥jthrJ\\geq j\_\{\\text\{thr\}\}, no cos\)56 \(3\.3%\)56 \(3\.3%\)56 \(3\.3%\)Stage\-2 only \(cos≥cthr\\cos\\geq c\_\{\\text\{thr\}\}, noJJ\)455 \(27\.1%\)148 \(8\.8%\)79 \(4\.7%\)To ground Finding 1 we audit a researcher\-curated baseline—a 1,679\-record sample drawn fromPhysCorp\-pre\-audit\(the released 14,294\-record pre\-audit pool\) under conventional 5\-gram\-Jaccard \+ within\-pool embedding dedup, paired against a 500\-record internal analysis eval\. The 500\-record eval is held internal to ground this finding and is distinct fromPhysOlym\-A\(constructed post\-audit\); the 1,679\-record sample is reproducible from the releasedPhysCorp\-pre\-auditso the audit can be re\-run by downstream users\. Stage\-1 catches 56 records \(3\.3%\) all atJ=1\.0J\{=\}1\.0\(bimodal distribution—verbatim duplication under our normalization\)\. Stage\-2 cosine≥0\.85\\geq 0\.85flags an additional 92 records \(5\.5%\), bringing the joint leak rate to 148 \(8\.8%\); the cosine dimension sweeps4\.74\.7–27\.1%27\.1\\%acrosscos∈\{0\.90,0\.85,0\.80\}\\cos\\in\\\{0\.90,0\.85,0\.80\\\}while Jaccard is flat at 3\.3% \(Table[4](https://arxiv.org/html/2605.14040#A1.T4)\)\. The 148 flagged decompose into 56 exact \(J=1\.0J\{=\}1\.0\), 23 strong paraphrases \(cos≥0\.90\\cos\\geq 0\.90\), 69 weak paraphrases \(0\.85≤cos<0\.900\.85\\leq\\cos<0\.90\)\. The 5\.5\-pp gap is the rephrasing dark\-matter that single\-stage audits miss because the*same*problem appears across upstream aggregations under wording that evadesJ≥0\.4J\\geq 0\.4but trips cosine≥0\.85\\geq 0\.85\.
#### Pipeline reproducibility note\.
Threshold\-sensitivity numbers were produced byaudit\_pipeline/threshold\_sensitivity\.py\. Normalization: lower\-case, strip LaTeX commands \(\\frac,\\sqrt, …\), remove$\{\}\[ \]\(\)delimiters, collapse whitespace, drop<5<5\-word shingles\. Embedding:sentence\-transformers 5\.4\.1, batch 32,L2L\_\{2\}\-normalized\. Best\-overlap via inverted\-index pruning \(Stage 1\) and fullN×MN\{\\times\}Mmatmul \(Stage 2\)\. Saved scores inthreshold\_sensitivity\_scores\.npz\.
#### Bimodal Jaccard distribution\.
Every Stage\-1 leak in our researcher\-curated baseline sits atJ=1\.0J\{=\}1\.0: aggressive normalization collapses surface variance, so shared problem statements yield identical shingle sets while distinct records land belowJ=0\.3J\{=\}0\.3\. Cosine≥0\.85\\geq 0\.85is therefore the sole paraphrase\-class detector here; bimodality is corpus\-specific\.
#### Re\-audit and the cleaned PhysR1Corp\.
Starting from a2,4332\{,\}433\-record candidate closed\-form pool, three sequential cleanup passes against the six paper\-canonical comparison corpora \(PhyX, MMMU\-Pro Physics, OlympiadBench\-Physics, PhysReason\-full, PhysUniBench\-en,PhysOlym\-A\) produced the releasedPhysR1Corp: \(i\) MMMU\-Pro Physics jointJ≥0\.4∨cos≥0\.85J\\geq 0\.4\\vee\\cos\\geq 0\.85re\-audit dropped 87 records \(Stage\-1: 16 records,16/60=26\.7%16/60\{=\}26\.7\\%MMMU\-Pro coverage; Stage\-2 union: 87 records,52/60=86\.7%52/60\{=\}86\.7\\%\); \(ii\) PhyX\-mini cosine≥0\.85\\geq 0\.85Stage\-2 audit dropped 69 records, all real\-near\-duplicate physics\-problem variants confirmed by manual inspection \(e\.g\. MgF2anti\-reflection\-coating problem with slightly tweakednglassn\_\{\\text\{glass\}\}and option labels, top cos0\.930\.93\); \(iii\) PhysUniBench\-en cosine≥0\.85\\geq 0\.85Stage\-2 audit dropped 9 template duplicates \(top cos0\.930\.93, shared problem\-template stems\)\. Total dropped:87\+69\+9=16587\+69\+9=165records \(with 3 records flagged in multiple channels netting to87\+78=16587\+78=165unique\)\. The released𝟐,𝟐𝟔𝟖\\mathbf\{2\{,\}268\}\-recordPhysR1Corpretains 19 Stage\-2 hits across PhysOlym\-A \(3\), MMMU\-Pro \(1\), OlymBench\-Phys \(4\), PhysReason\-full \(11\) classified as same\-topic neighbors on manual inspection \(top cos≤0\.87\\leq 0\.87, different problem setups, dimensionalities, or geometries; Table[2](https://arxiv.org/html/2605.14040#S3.T2)\)\. MMMU\-Pro Physics remains excluded from the headline Table[3](https://arxiv.org/html/2605.14040#S5.T3), the saturation\-gradient narrative in §[5\.1](https://arxiv.org/html/2605.14040#S5.SS1.SSS0.Px1), and the dense\-vs\-binary ablation in §[5\.2](https://arxiv.org/html/2605.14040#S5.SS2.SSS0.Px6), because the eval is small \(60 records\) and prior work has flagged it as contaminated against multiple frontier training corpora; Sonnet’s MMMU\-Pro Physics number is reported only as a frontier\-model reference \(Sonnet was not trained onPhysR1Corp\)\.
#### Embedding\-model and threshold rationale\.
We chosemxbai\-embed\-large\-v1for Stage 2 overbge\-large\-en\-v1\.5,e5\-large\-v2, OpenAItext\-embedding\-3\-large, and Voyagevoyage\-3because it \(i\) is permissively licensed \(Apache 2\.0, no API dependency\); \(ii\) ships 1024\-dim normalized vectors with cosine tuning; \(iii\) scored highest on MTEB physics\-adjacent retrieval at audit time\. Thecos≥0\.85\\cos\\geq 0\.85threshold was calibrated against worked examples \(UGPhysics paraphrase:J=0\.31J\{=\}0\.31Stage\-1\-miss,cos=0\.91\\cos\{=\}0\.91Stage\-2\-catch\); thecos∈\{0\.80,0\.85,0\.90\}\\cos\\in\\\{0\.80,0\.85,0\.90\\\}sensitivity grid brackets the operating point\.
#### Embedder\-sensitivity ablation \(mxbai vs\.text\-embedding\-3\-large\)\.
We re\-encode the releasedPhysCorp\-pre\-auditpool \(14,29414\{,\}294records\) andPhysOlym\-A\(500500records\) under OpenAItext\-embedding\-3\-largeand compute pairwise cosines, then compare per\-train\-record max\-cosine rankings against the mxbai baseline\.Spearmanρ=0\.78\\rho\{=\}0\.78\(p≈0p\{\\approx\}0\); the candidate\-set relationship at the operating threshold is summarized below\.
Thresholdcos≥\\cos\\geqmxbai cand\.text\-embedding\-3\-largecand\.bothonly mxbaionly OpenAI0\.850\.85\(paper op\.\)763763514514514514249249𝟎\\mathbf\{0\}0\.870\.87631631512512512512119119𝟎\\mathbf\{0\}0\.900\.905515515115115115114040𝟎\\mathbf\{0\}0\.950\.955205205065065065061414𝟎\\mathbf\{0\}
text\-embedding\-3\-largeflags a*strict subset*of mxbai’s candidates at every threshold \(only\-OpenAI count is0at all four levels\): every recordtext\-embedding\-3\-largewould catch is also caught by mxbai\. mxbai is therefore the more conservative \(higher\-recall\) Stage\-2 embedder, and the audit cannot have missed any contamination thattext\-embedding\-3\-largewould have surfaced\. The candidate\-set Jaccard at the operating threshold is0\.670\.67\.*Caveat\.*This ablation is on the user\-facing audit case \(released pool↔\\leftrightarrowreleased eval\), not the SciInstruct competitor\-pool case from Table[2](https://arxiv.org/html/2605.14040#S3.T2); the strict\-subset direction does not formally transfer, but theρ=0\.78\\rho\{=\}0\.78rank correlation suggests the SciInstruct134134\-near\-duplicate count is robust in direction under embedder change\. Sensitivity ablation againstvoyage\-3is left to follow\-up work\.
#### Stage\-3 judge model dependence and reproducibility\.
Stage\-3 introduces a dependence on Anthropic’s Haiku\-4\.5 \(the close\-duplicate vs\. same\-topic\-neighbor classifier\)\. To ensure long\-term reproducibility we \(i\) pin the exact model identifier \(claude\-haiku\-4\-5as of 2026\-05\) in the releasedaudit\_three\_stage\.py; \(ii\) release the full per\-pair judge prompts and per\-pair verdict labels \(judge\_labelarrays\) alongside the cosine scores inthreshold\_sensitivity\_scores\.npz, so the contamination flag set is reproducible without re\-querying the API; \(iii\) document a fallback protocol \(Sonnet 4\.5 or GPT\-4o on the same prompt template\) for cases where Haiku\-4\.5 access is no longer available\. Because Stage\-3 is a precision filter applied only to Stage\-2 candidates \(≤0\.5%\\leq 0\.5\\%of the train pool\), its labels are also the most amenable to manual re\-verification by a downstream auditor; the cosine\-bucketed precision profile \(Table[A](https://arxiv.org/html/2605.14040#A1.SS0.SSS0.Px1)\) gives a cheap calibration signal against any future judge\.
## Appendix BHeld\-out splits and corpus annotation schema
The released splits arePhysCorp\-A\(the 6,432\-record audited corpus\) with its closed\-form RL carve\-outPhysR1Corp\(2,268 records\) andPhysOlym\-A\(the 500\-problem held\-out olympiad eval\)\. The annotation schema below applies uniformly across all three\.
#### Eight\-field annotation schema\.
Each record carries the following annotations, generated by Sonnet 4\.5 batch annotation \(3,900 records with full annotation; the remainder carry source\-native labels merged into the same schema for∼\\sim31,000 total label values\):
- •difficulty∈\\in\{1,2,3,4,5\} \(Sonnet\-aggregated\), plus optional source\-native: Estonian organizer\-issued 1–10, Zhou pedagogical 1–5, Zhou advanced \[A\] flag\.
- •concept∈\\in\{Mechanics, Electromagnetism, Quantum, Thermodynamics, Waves, Optics, Modern, Relativity, Particle\}\.
- •problem\_type∈\\in\{Conceptual, Computational, Proof\-based, Experimental\}\.
- •expected\_solution\_length∈\\in\{S, M, L\} \(short / medium / long\)\.
- •math\_level∈\\in\{Algebra, Calculus, Vector, LinearAlgebra, DiffEq\}\.
- •modality∈\\in\{text, multimodal\}\.
- •language\(BCP\-47\):en,et,en\-et\(bilingual paired\)\.
- •license\(SPDX\-style\):CC\-BY\-4\.0,CC\-BY\-SA\-4\.0,CC\-BY\-NC\-4\.0,Public\-Domain\.
#### Worked example record \(JSONL\)\.
A typical record fromPhysOlym\-A:
> \{ "index": "estonian\_2017\_lahtine\_2", "source": "estonian\_olympiad", "license": "CC\-BY\-NC\-4\.0", "language": "en\-et", "messages": \[\{"role": "user", "content": "An ideal gas…"\}\], "solution": "By the first law…\\boxed\{1\.5\}\\backslash boxed\\\{1\.5\\\}", "images": \[\], "concept": "Thermodynamics", "difficulty": 4, "native\_difficulty": \{"scale": "1\-\-10", "value": 7\}, "problem\_type": "Computational", "expected\_solution\_length": "M", "math\_level": "Calculus", "modality": "text", "audit\_passed": true \}
#### Inter\-annotator agreement\.
For the 3,900\-record fully\-annotated subset, Sonnet 4\.5 is run with two seeds on the same 100 records \(random sample, seed 42\); per\-field Cohen’sκ\\kappabetween the two annotation runs is reported in the released dataset card\. Preliminary inspection showsκ≥0\.85\\kappa\\geq 0\.85onconceptandproblem\_type\(categorical with sharp boundaries\),κ∼0\.7\\kappa\\sim 0\.7ondifficulty\(ordinal with neighboring\-class confusion\), andκ∼0\.6\\kappa\\sim 0\.6onexpected\_solution\_length\. The lowerκ\\kappaonexpected\_solution\_lengthreflects genuine ambiguity in the medium\-vs\-long boundary; downstream users requiring stable solution\-length labels should treat the field as a noisy proxy\.
#### Native difficulty preserved\.
For the 27% ofPhysOlym\-Arecords with Estonian native difficulty 1–10 and the 38% with Zhou pedagogical 1–5 point values, the*Sonnet\-aggregated*difficulty field is reported for cross\-source consistency, but the*native*difficulty is preserved in a separatenative\_difficultyfield with explicitscaleandvaluekeys\. The Sonnet difficulty curve in Table[12](https://arxiv.org/html/2605.14040#A8.T12)uses native Estonian 1–10, not the aggregated 1–5, because native labels avoid the self\-annotation circularity in which the difficulty estimate depends on the same model whose accuracy is being measured\.
## Appendix CReward function and full hyperparameter table
This appendix specifies both reward shapes referenced from §[4](https://arxiv.org/html/2605.14040#S4): the recommended binary correctness rewardrbinr\_\{\\mathrm\{bin\}\}defined inline in §[4](https://arxiv.org/html/2605.14040#S4)\(§[C\.1](https://arxiv.org/html/2605.14040#A3.SS1)\) and the dense five\-component reward of Equation[2](https://arxiv.org/html/2605.14040#S4.E2)reported as an ablation \(§[C\.2](https://arxiv.org/html/2605.14040#A3.SS2)\)\. Both share the matching/extraction primitives below; binary usesransr\_\{\\mathrm\{ans\}\}alone, dense composes all five components and clips\.
### C\.1Binary correctness reward \(recommended\)
The recommended reward is
rbin\(y,x\)=1\[Match\(ExtractBoxed\(y\),g\(x\)\)\]∈\{0,1\},r\_\{\\mathrm\{bin\}\}\(y,x\)\\;=\\;\\mathbb\{1\}\\\!\\big\[\\,\\textsc\{Match\}\\\!\\big\(\\textsc\{ExtractBoxed\}\(y\),\\,g\(x\)\\big\)\\,\\big\]\\in\\\{0,1\\\},whereExtractBoxedparses the last\\boxed\{\}via brace\-counting \(handles unlimited nesting, e\.g\.\\sqrt\{\\frac\{T\_0\}\{\\eta\}\}\) andMatchaccepts: \(i\) MCQ\-letter equality on\{\\\{A,B,C,D,…\}\\\}gold; \(ii\) multi\-part numeric agreement within±1%\\pm 1\\%relative tolerance, afterlatex\_to\_plainnormalization \(\\text\{\}/\\mathrm\{\}stripped,\\frac\{a\}\{b\}→\\to\(a\)/\(b\),\\times/\\cdot→\\to\*,\\pi→\\topi\),float\(\), theneval\(\)on expression\-like strings, then prefix\-numeric extraction; \(iii\) symbolic equivalence viasympy\.simplify\(expr\_pred \- expr\_gold\) == 0for symbolic gold\. Released asreward\_physics\.py; selected by env varDENSE\_REWARD=0\(the default\), which returnsransr\_\{\\mathrm\{ans\}\}as a clean 0/1 binary\.
#### Why binary is the deployable default \(full theoretical analysis\)\.
The \(P1\)/\(P2\) intuitions are stated inline in §[4](https://arxiv.org/html/2605.14040#S4); we add the \(P3\) derivation here\. \(P1\) Group\-normalization rendersAkA\_\{k\}invariant to affine rescaling ofrrwithin a group, so dense shaping only matters when it*reorders*rollouts; we measure14\.3%14\.3\\%of within\-group pairs flipped by dense,87%87\\%inside the all\-wrong subgroup\. \(P2\) Those wrong\-group flips reward LaTeX\-format proxies satisfiable without solving the physics—a Goodhart channel that hurts prose\-and\-equation open\-ended evaluation; the matched\-step\-60 binary\-vs\-dense gap is largest on the auditedPhysOlym\-A\-liberal split\.
#### \(P3\) Binary reward maximizes per\-prompt advantage variance after the difficulty curriculum\.
The MM\-Eureka curriculum drops prompts where allKKrollouts are correct or all are wrong, so on every surviving prompt the rollout\-correct rate isp∈\(0,1\)p\\in\(0,1\)\. Binary reward is then a Bernoulli over rollouts:r¯=p\\bar\{r\}=p,σr2=p\(1−p\)\\sigma\_\{r\}^\{2\}=p\(1\-p\), and the group\-normalized advantages take exactly two values
Akbin=\{\(1−p\)/prk=1−p/\(1−p\)rk=0,Var\(Abin\)=1\.A^\{\\mathrm\{bin\}\}\_\{k\}\\;=\\;\\begin\{cases\}\\;\\sqrt\{\(1\-p\)/p\}&r\_\{k\}=1\\\\\[2\.0pt\] \-\\sqrt\{p/\(1\-p\)\}&r\_\{k\}=0\\end\{cases\},\\qquad\\mathrm\{Var\}\(A^\{\\mathrm\{bin\}\}\)\\;=\\;1\.\(3\)The advantage variance is exactly11, the maximum aKK\-sample group\-normalized estimator can carry on a Bernoulli reward\. Adding a bounded shaping termδk∈\[0,Δ\]\\delta\_\{k\}\\in\[0,\\Delta\]withΔ=0\.45\\Delta\{=\}0\.45\(therfmt\+rdim\+rsymr\_\{\\mathrm\{fmt\}\}\{\+\}r\_\{\\mathrm\{dim\}\}\{\+\}r\_\{\\mathrm\{sym\}\}budget of the dense reward\) inflates the within\-group standard deviationσr\\sigma\_\{r\}by an O\(Δ2\)\(\\Delta^\{2\}\)term while leaving the between\-correctness mean separation roughly unchanged, so the magnitude of the average correct\-vs\-wrong advantage*shrinks*:
\|Acorrectdense\|≈1−pp\(1−p\)\+Δ2/12<1−pp\(1−p\)=\|Acorrectbin\|\.\\big\|A^\{\\mathrm\{dense\}\}\_\{\\mathrm\{correct\}\}\\big\|\\;\\approx\\;\\frac\{1\-p\}\{\\sqrt\{p\(1\-p\)\+\\Delta^\{2\}/12\}\}\\;<\\;\\frac\{1\-p\}\{\\sqrt\{p\(1\-p\)\}\}\\;=\\;\\big\|A^\{\\mathrm\{bin\}\}\_\{\\mathrm\{correct\}\}\\big\|\.\(4\)Dense reward thus trades*between*\-correctness gradient signal\-to\-noise for*within*\-correctness rank information, but the within\-correctness flips are exactly the Goodhart channel of \(P2\)\.
### C\.2Dense five\-component physics\-native reward \(ablation\)
The dense five\-component Physics\-R1 reward \(Equation[2](https://arxiv.org/html/2605.14040#S4.E2), Section[4](https://arxiv.org/html/2605.14040#S4)\) is implemented asr=rans\+rfmt\+rdim\+rsym\+rconsr=r\_\{\\mathrm\{ans\}\}\+r\_\{\\mathrm\{fmt\}\}\+r\_\{\\mathrm\{dim\}\}\+r\_\{\\mathrm\{sym\}\}\+r\_\{\\mathrm\{cons\}\}, clipped to\[−1,1\]\[\-1,1\]\.
#### Per\-component implementation\.
Full code inreward\_physics\.py\.*rans∈\{0,\+1\}r\_\{\\mathrm\{ans\}\}\\in\\\{0,\+1\\\}:*brace\-counting\\boxedparser; MCQ\-letter equality; numeric/symbolic vialatex\_to\_plainnormalization \(\\frac,\\text,\\pi,\\times/\\cdot\) thenfloat/eval/prefix\-numeric; multi\-part requires all parts±1%\\pm 1\\%\.*rfmt∈\{0,\+0\.1\}r\_\{\\mathrm\{fmt\}\}\\in\\\{0,\+0\.1\\\}:*non\-empty\\boxed\{\}\.*rdim∈\{0,\+0\.15\}r\_\{\\mathrm\{dim\}\}\\in\\\{0,\+0\.15\\\}:*number\-prefix\-guarded regex extracts unit tokens, mapped tosympy\.physics\.units\(32 tokens\); fired only when every detected unit resolves\.*rsym∈\{0,\+0\.20\}r\_\{\\mathrm\{sym\}\}\\in\\\{0,\+0\.20\\\}:*first\-successsympy\.sympifyon LaTeX\-cleaned\\frac\{NUM\}\{DEN\}intermediates\.*rcons∈\{−0\.25,0\}r\_\{\\mathrm\{cons\}\}\\in\\\{\-0\.25,0\\\}:*negative\-only penalty when energy/momentum balance from corpus annotation is violated by\>5%\>5\\%relative\.
#### Mode and version pins\.
Released asreward\_physics\.py\. The mode is selected by env varDENSE\_REWARD:0\(default, recommended\) returnsransr\_\{\\mathrm\{ans\}\}only as 0/1 binary, recoveringrbinr\_\{\\mathrm\{bin\}\}of §[C\.1](https://arxiv.org/html/2605.14040#A3.SS1);1returns the full clipped dense sum \(ablation\)\. The audit factor \(envAUDIT\_LAMBDA\) zeros out reward on contaminated training items when set\. Version pins:sympy==1\.13\.3,transformers==4\.57\.0,vllm==0\.11\.0,verl==0\.6\.1\.
#### Worked example \(rollout group; binary vs\. dense advantages\)\.
A concrete realization of the \(P1\)/\(P2\) Goodhart channel of §[4](https://arxiv.org/html/2605.14040#S4)on one prompt withK=8K\{=\}8rollouts\. The prompt is a two\-step kinematics problem with gold19\.6\\boxed\{19\.6\}N \(problem idefo\-2014\-3a\); rolloutsy1,…,y4y\_\{1\},\\dots,y\_\{4\}commit to the correct answer with varying CoT quality,y5,…,y8y\_\{5\},\\dots,y\_\{8\}commit to wrong answers ranging from arithmetic\-slip \(19\.8\\boxed\{19\.8\},\>1%\{\>\}1\\%\) to off\-by\-physics \(42\.0\\boxed\{42\.0\}\)\.
kkFinal⋅\\boxed\{\\cdot\}CoT shaperansr\_\{\\mathrm\{ans\}\}rfmtr\_\{\\mathrm\{fmt\}\}rdimr\_\{\\mathrm\{dim\}\}rsymr\_\{\\mathrm\{sym\}\}rconsr\_\{\\mathrm\{cons\}\}rbinr\_\{\\mathrm\{bin\}\}rdenser\_\{\\mathrm\{dense\}\}119\.619\.6full units \+\\frac110\.100\.100\.150\.150\.200\.2001\.00\\mathbf\{1\.00\}1\.00\\mathbf\{1\.00\}219\.619\.6units only, no\\frac110\.100\.100\.150\.15001\.00\\mathbf\{1\.00\}1\.00\\mathbf\{1\.00\}319\.619\.6\\fraconly, no units110\.100\.1000\.200\.2001\.00\\mathbf\{1\.00\}1\.00\\mathbf\{1\.00\}419\.619\.6sparse CoT \(“answer is19\.619\.6”\)110\.100\.100001\.00\\mathbf\{1\.00\}1\.00\\mathbf\{1\.00\}542\.042\.0full units \+\\frac00\.100\.100\.150\.150\.200\.2000\.00\\mathbf\{0\.00\}0\.450\.45619\.819\.8units only, no\\frac00\.100\.100\.150\.15000\.00\\mathbf\{0\.00\}0\.250\.25742\.042\.0\\fraconly, no units00\.100\.1000\.200\.2000\.00\\mathbf\{0\.00\}0\.300\.308no⋅\\boxed\{\\cdot\}rambling, no commit000000\.00\\mathbf\{0\.00\}0\.000\.00
After clipping \(r≤1r\{\\leq\}1\) the four correct rollouts collapse to identical reward under both shapes\. After group\-normalization \(Eq\.[1](https://arxiv.org/html/2605.14040#S4.E1)\) the binary advantage vector isAbin=\(\+1,\+1,\+1,\+1,−1,−1,−1,−1\)A^\{\\mathrm\{bin\}\}\{=\}\(\+1,\+1,\+1,\+1,\-1,\-1,\-1,\-1\), every correct rollout receives equal positive gradient and every wrong rollout equal negative gradient\. The dense advantage vector isAdense≈\(\+1\.04,\+1\.04,\+1\.04,\+1\.04,−0\.55,−1\.00,−0\.93,−1\.69\)A^\{\\mathrm\{dense\}\}\\\!\\approx\\\!\(\+1\.04,\+1\.04,\+1\.04,\+1\.04,\-0\.55,\-1\.00,\-0\.93,\-1\.69\)\(computed exactly:r¯=0\.625\\bar\{r\}\{=\}0\.625,σr≈0\.428\\sigma\_\{r\}\{\\approx\}0\.428\)\.*Three observations land the \(P1\)–\(P3\) theory at the sample level:*
- •\(P1\) realized\. Among the four correct rollouts \(k=1,2,3,4k\{=\}1,2,3,4\), dense and binary assign*the same*advantage to every rollout—rank\-equivalent in the correct subgroup, even though rollouts 1–3 “deserve” more credit by the*a priori*physics\-native intuition\. Clipping at11erases the dense\-side variation among correct rollouts\.
- •\(P2\) realized\. Among the four wrong rollouts \(k=5,6,7,8k\{=\}5,6,7,8\), dense reorders them by LaTeX surface form:k=5k\{=\}5\(well\-formatted, units,\\frac,42\.0\\boxed\{42\.0\}way wrong\) gets the*smallest*negative advantage \(−0\.55\-0\.55\),k=8k\{=\}8\(no boxed commit\) gets the most negative \(−1\.69\-1\.69\)\. The policy gradient is therefore pushed toward producing well\-formatted wrong reasoning over poorly\-formatted wrong reasoning—the canonical Goodhart channel\. Notek=5k\{=\}5has dense advantage−0\.55\-0\.55vs\. binary−1\.00\-1\.00: the wrong\-answer gradient is*weakened*for the format\-compliant rollout, exactly the bias toward proxy satisfaction\.
- •\(P3\) realized\. The magnitude of the average correct\-rollout advantage is\|Acorrectbin\|=1\.0\|A^\{\\mathrm\{bin\}\}\_\{\\mathrm\{correct\}\}\|\{=\}1\.0vs\.\|Acorrectdense\|≈1\.04\|A^\{\\mathrm\{dense\}\}\_\{\\mathrm\{correct\}\}\|\{\\approx\}1\.04\(very close because clipping caps both atr=1r\{=\}1\)\. The magnitude of the average wrong\-rollout advantage is\|Awrongbin\|=1\.0\|A^\{\\mathrm\{bin\}\}\_\{\\mathrm\{wrong\}\}\|\{=\}1\.0vs\.\|Awrongdense\|≈1\.04\|A^\{\\mathrm\{dense\}\}\_\{\\mathrm\{wrong\}\}\|\{\\approx\}1\.04on the dense side as well, but the within\-wrong spread ofσ=0\.42\\sigma\{=\}0\.42across\{−0\.55,−0\.93,−1\.00,−1\.69\}\\\{\-0\.55,\-0\.93,\-1\.00,\-1\.69\\\}is the within\-group rank\-flipping variance that absorbs gradient capacity into the Goodhart direction\. Binary spends zero capacity on within\-correctness ranking and all of it on the correct\-vs\-wrong axis\.
The aggregate effect of running this calculus across∼\\sim1,024 prompts and 60 GRPO steps is the matched\-step\-60 binary\-vs\-dense gap of Table[3](https://arxiv.org/html/2605.14040#S5.T3)\(problem\-level liberal Sonnet\-judge accuracy across all open\-ended columns; bug\-corrected per §[5](https://arxiv.org/html/2605.14040#S5.SS0.SSS0.Px2)\): PhysReason32\.232\.2vs\.23\.323\.3\(\+8\.9\+8\.9pp\), PUB\-OE37\.037\.0vs\.37\.737\.7\(−0\.7\-0\.7pp, tied\), OlymBench\-Phys liberal45\.445\.4vs\.40\.540\.5\(\+4\.9\+4\.9pp\),PhysOlym\-Aliberal25\.625\.6vs\.19\.219\.2\(\+6\.4\+6\.4pp\)\.
#### Hyperparameter and framework details\.
The full GSPO\+DAPO configuration with all flags, lambdas, clip ranges, dynamic\-sampling settings, and the difficulty\-curriculum thresholds is in Table[10](https://arxiv.org/html/2605.14040#A8.T10)\(main text\) and the released YAML\. Theverl0\.6\.1 FSDP1 reproducibility note \(Section[6](https://arxiv.org/html/2605.14040#S6)\) and the upstream GitHub issue link are tracked in the released README\.
Algorithm 2Dense five\-component physics\-native reward for one rollout\.1:Solution string
ss, gold answer
gg, optional conservation flag
cons\\mathrm\{cons\}fromextra\_info
2:
r∈\[−1,1\]r\\in\[\-1,1\]
3:
ans←ExtractBoxed\(s\)\\mathrm\{ans\}\\leftarrow\\textsc\{ExtractBoxed\}\(s\)⊳\\trianglerightbrace\-counting parser; nested\\boxedOK
4:
rans←\+1r\_\{\\mathrm\{ans\}\}\\leftarrow\+1if
Match\(ans,g\)\\textsc\{Match\}\(\\mathrm\{ans\},g\)under MCQ\-letter / multi\-part /
±1%\\pm 1\\%numeric tolerance, else
0
5:
rfmt←\+0\.1r\_\{\\mathrm\{fmt\}\}\\leftarrow\+0\.1if
ans\\mathrm\{ans\}is non\-empty \(well\-formed\\boxed\{…\}\), else
0
6:
U←ExtractUnits\(s\)U\\leftarrow\\textsc\{ExtractUnits\}\(s\)⊳\\trianglerightregex\(number\)\(unit\)\(exponent\)with number\-prefix guard
7:
rdim←\+0\.15r\_\{\\mathrm\{dim\}\}\\leftarrow\+0\.15if
\|U\|≥1\|U\|\\geq 1and
∀u∈U:ResolvesInSymPy\(u\)\\forall u\\in U:\\textsc\{ResolvesInSymPy\}\(u\), else
0
8:
F←ExtractFracs\(s\)F\\leftarrow\\textsc\{ExtractFracs\}\(s\)⊳\\trianglerightfind all\\frac\{NUM\}\{DEN\}
9:
rsym←\+0\.20r\_\{\\mathrm\{sym\}\}\\leftarrow\+0\.20if
∃\(NUM,DEN\)∈F:Sympifies\(NUM\)∧Sympifies\(DEN\)\\exists\(\\mathrm\{NUM\},\\mathrm\{DEN\}\)\\in F:\\textsc\{Sympifies\}\(\\mathrm\{NUM\}\)\\land\\textsc\{Sympifies\}\(\\mathrm\{DEN\}\), else
0
10:if
cons\\mathrm\{cons\}is provided \(
∼\\sim3,100\-record subset\)then
11:
p^←TryFloat\(ans\)\\hat\{p\}\\leftarrow\\textsc\{TryFloat\}\(\\mathrm\{ans\}\);
p∗←∑cons\.in−∑cons\.out∖ansp^\{\*\}\\leftarrow\\sum\\\!\\mathrm\{cons\.in\}\-\\sum\\\!\\mathrm\{cons\.out\}\\setminus\\mathrm\{ans\}
12:
rcons←−0\.25r\_\{\\mathrm\{cons\}\}\\leftarrow\-0\.25if
\|p^−p∗\|/max\(\|p∗\|,\|p^\|,ε\)\>0\.05\|\\hat\{p\}\-p^\{\*\}\|/\\max\(\|p^\{\*\}\|,\|\\hat\{p\}\|,\\varepsilon\)\>0\.05, else
0
13:else
14:
rcons←0r\_\{\\mathrm\{cons\}\}\\leftarrow 0⊳\\trianglerightnegative\-only; no positive reward for satisfaction
15:endif
16:return
clip\(rans\+rfmt\+rdim\+rsym\+rcons,−1,1\)\\mathrm\{clip\}\(r\_\{\\mathrm\{ans\}\}\+r\_\{\\mathrm\{fmt\}\}\+r\_\{\\mathrm\{dim\}\}\+r\_\{\\mathrm\{sym\}\}\+r\_\{\\mathrm\{cons\}\},\-1,1\)
## Appendix DLLM\-judge details: three judges, scoring conventions, and reproducibility
This appendix documents the three Sonnet\-4\.5\-as\-judge variants used in Table[3](https://arxiv.org/html/2605.14040#S5.T3)and the scoring conventions adopted across columns\. Source code for all three judges is released in thejudge/directory of the code repository \([github\.com/shanyang\-me/physics\-r1\-neurips2026](https://arxiv.org/html/2605.14040v1/github.com/shanyang-me/physics-r1-neurips2026)\)\.
#### Why three judges\.
Open\-ended physics olympiad problems differ structurally: PhysReason and PhysUniBench\-OE problems are explicitly multi\-sub\-part \(often22–55sub\-questions per record, each with its own gold answer\), whilePhysOlym\-Aand OlympiadBench\-Physics problems are graded at the problem level \(a single gold solution document, with the model’s final answer compared against it\)\. A single judge prompt cannot serve both\. We therefore use three judges, each tuned to its eval’s structure:
- •judge\_olympiad\.py\(problem\-level, used forPhysOlym\-Aand OlympiadBench\-Physics\)\. One Sonnet 4\.5 call per problem; the prompt provides the full gold solution paragraph and a\\boxed\{\}\-extracted candidate answer \(with a600600\-char response\-tail fallback if no boxed is emitted\), and asks for a single YES/NO verdict on whether the candidate’s final answer is mathematically/physically equivalent to the gold’s\. Tolerance is2%2\\%relative\.
- •llm\_judge\_v2\_alignment\.py\(per\-subpart, AND across sub\-parts, used for PhysReason\)\. For each gold sub\-answergig\_\{i\}, a separate Sonnet 4\.5 call asks:*does ANY of the candidate’s predictions equalgig\_\{i\}?*Per\-subpart verdict is YES/NO\.judge\_problem\_correctis the AND across all sub\-parts\. Tolerance:1%1\\%\.
- •llm\_judge\_v3\_pubeo\.py\(per\-subpart, AND across sub\-parts, with cached clean gold \+ tail fallback, used for PhysUniBench\-OE\)\. Same per\-subpart structure as v2, but with a pre\-extracted*clean*per\-subpart gold list \(e\.g\.,\["2\.68nC", "7853\.1W"\]\) that bypasses PUB\-OE’s verbose paragraph\-form gold\. If the candidate’s\\boxed\{\}list is empty, a regex fallback scans the last300300chars of the response for likely numeric/symbolic answers\. Per\-subpart verdict is YES/NO;judge\_problem\_correct\_v3is the AND across all sub\-parts\. Tolerance:2%2\\%\.
#### Scoring conventions in Table[3](https://arxiv.org/html/2605.14040#S5.T3)\.
All open\-ended cells use*problem\-level*accuracy: for multi\-sub\-part problems \(PhysReason, PhysUniBench\-OE\) every sub\-part must be judged correct for the problem to count, via thejudge\_problem\_correctfield of the v2/v3 judges \(AND across sub\-parts\); for problem\-level evals \(PhysOlym\-A, OlympiadBench\-Physics\)judge\_olympiad\.pyreturns one YES/NO per problem\. We also computed a softer*per\-subpart*variant \(partial credit,∑i𝟙\[subicorrect\]/∑i1\\sum\_\{i\}\\mathbb\{1\}\[\\text\{sub\}\_\{i\}\\text\{ correct\}\]/\\sum\_\{i\}1\) for the multi\-sub\-part columns and verified it is uniformly44–1717pp higher across rows; per\-subpart values are released alongside the dataset for users who want to study the partial\-credit lens but are not reported as headline in Table[3](https://arxiv.org/html/2605.14040#S5.T3)\.
#### Reproducibility note\.
All Sonnet\-judge runs in Table[3](https://arxiv.org/html/2605.14040#S5.T3)useworkers=22–44concurrency; errored sub\-judgments are filtered and re\-judged at lower concurrency rather than counted as wrong\. Per\-cell judge\-error counts \(typically≤1%\\leq 1\\%of records\) and per\-record verdicts are released in the supplementary archive atjudge\_audit\.json, so downstream auditors can verify each cell independently\.
#### Verbatim judge prompt \(problem\-level,PhysOlym\-A\+ OlymBench\)\.
The Sonnet 4\.5 judge is invoked with the following template:
> You are grading a physics olympiad answer\. GOLD \(full reference solution; the final numeric/symbolic answer is what matters\): \{gold\} CANDIDATE answers \(extracted from the model’s \\boxed\{\} markers\): \{preds\} \(If the candidate emitted no \\boxed\{\}, the candidate is the last 600 chars of its full response:\) \{tail\} Task: decide whether the candidate’s final answer is mathematically/physically equivalent to the gold’s final answer\. Allow: different but equivalent algebraic forms; trivial unit/format differences; rounding within 2% relative tolerance; trailing prose\. Reject: different magnitude, different sign, different functional form, missing or wrong physical content, no answer\. Respond with EXACTLY one word: YES or NO\.
#### Verbatim judge prompt \(per\-subpart, PhysReason \+ PhysUniBench\-OE\)\.
For each gold sub\-answergig\_\{i\}, the judge is invoked with:
> You are grading a physics olympiad answer\. GOLD answer: \{g\_i\} CANDIDATE predictions \(one or more, separated by ===\): \{preds\} Task: decide whether ANY of the candidate predictions is mathematically/physically equivalent to the gold answer\. Allow: different but equivalent algebraic forms; trivial unit/format differences \("450 N" == "450 \\text\{ N\}"\); rounding within 1\-\-2% relative tolerance; trailing prose \("approximately", "to the right"\); different variable names mapping cleanly\. Reject: different magnitude, different sign, different functional form, missing or wrong physical content\. Respond with EXACTLY one word: YES \(if any candidate matches\) or NO\. No other text\.
#### Per\-chunk verdict breakdown forPhysOlym\-A\(Sonnet 4\.5\)\.
The 500 problems are partitioned into 5 chunks of 100 \(random seed 42, source\-stratified\)\. Per\-chunk strict\-correct counts: 38, 20, 35, 30, 40 \(judgeable\-only denominators 98, 100, 99, 98, 96 after removing13\.9%13\.9\\%unjudgeables\)\. The chunk\-1 outlier at20%20\\%accuracy is concentrated in reference\-pointer Zhou problems whose gold solutions cite external olympiad\-handbook material rather than providing self\-contained answers; these are the unjudgeable category we surface as a known noise floor\.
#### Inter\-judge agreement\.
The Sonnet judge is run with two seeds on the same 500 problems forPhysOlym\-A\. Cohen’sκ\\kappaon the YES/NO label between the two passes is reported alongside the released dataset; preliminary inspection showsκ≥0\.8\\kappa\\geq 0\.8on thePhysOlym\-Acorpus\.
#### Human\-graded subset\.
A 100\-problem random subsample ofPhysOlym\-ASonnet predictions is graded by a single physics\-trained annotator using the same YES/NO rubric\. Per\-record human\-vs\-LLM agreement and Cohen’sκ\\kappaare released as a JSON alongside the dataset\. We report thisκ\\kappaas a calibration check on the LLM judge, not as a calibrated human\-baseline ground\-truth \(cf\. Section[6](https://arxiv.org/html/2605.14040#S6)\)\.
#### Released artifacts\.
The verbatim judge prompts \(problem\-level \+ per\-subpart\), the YES/NO rubric, the per\-chunk verdict breakdown, the 100\-problem human\-graded subset, and the per\-record agreement matrices are released asjudge\_prompts\.txt,rubric\.json,per\_chunk\_verdicts\.json, andhuman\_graded\_subset\.json\. Thepub\_oe\_gold\_cache\.json\(per\-id clean per\-subpart gold list\) is released alongsidellm\_judge\_v3\_pubeo\.pyand is reproducible from the original PhysUniBench\-OE source via the cache\-builder script in the same directory\.
#### Self\-grading concern\.
Sonnet 4\.5 is both the highest\-scoring frontier baseline onPhysOlym\-Aand the judge used for liberal accuracy\. Three checks bound any self\-favoring bias: \(i\) the strict \(numeric/symbolic match, judge\-independent\) Sonnet score is reported alongside liberal—the4\.74\.7\-pp Sonnet strict\-vs\-liberal gap \(28\.7%28\.7\\%vs\.33\.4%33\.4\\%\) bounds maximum leniency; \(ii\) on the failure\-mode taxonomy \(Appendix[H\.1](https://arxiv.org/html/2605.14040#A8.SS1)\) the dominant Sonnet error category iswrong\_subpart\(structural mismatch any judge marks wrong\), notvalid\_partial; \(iii\) we report a cross\-vendor judge agreement against GPT\-4o on a5050\-problem random subsample ofPhysOlym\-A\(Qwen3\-VL\-8B\-Thinking responses, seed4242\)\. Sonnet 4\.5 and GPT\-4o under an identical prompt template show88%88\\%raw agreement\(44/5044/50\) andCohen’sκ=0\.44\\kappa\{=\}0\.44\(moderate, per Landis & Koch\)\. The disagreement is asymmetric: GPT\-4o flips55Sonnet\-NO records to YES, while only11Sonnet\-YES record is flipped by GPT\-4o to NO \(McNemar exact on66discordant pairs\{5:Sonnet\-NO→GPT\-YES,1:Sonnet\-YES→GPT\-NO\}\\\{5\{:\}\\text\{Sonnet\-NO\}\\\!\\to\\\!\\text\{GPT\-YES\},\\,1\{:\}\\text\{Sonnet\-YES\}\\\!\\to\\\!\\text\{GPT\-NO\}\\\}, two\-sidedp=0\.219p\{=\}0\.219; the asymmetry is not significant atn=6n\{=\}6, but the direction is preserved on a larger200200\-problem ablation left to follow\-up work\)\. GPT\-4o’s positive rate \(16%16\\%,8/508/50\) is roughly*twice*Sonnet’s \(8%8\\%,4/504/50\), which means the cross\-vendor judge would assign Physics\-R1*higher*numbers than the Sonnet\-judge headlines, not lower—the self\-grading direction is the opposite of what the self\-favoring concern would predict\. The full per\-pair verdicts are released ascross\_judge\_50\.jsonlalongside the dataset\.
## Appendix EPer\-source license and provenance log
Per\-source provenance for each of the nine source families: full name, scrape URL, scrape date, original license string, redistribution tier, and outreach log\.
#### UGPhysics\.
Source:Xuet al\.\[[2025](https://arxiv.org/html/2605.14040#bib.bib3)\]\(ICML 2025\)\. 5,520 EN/ZH undergraduate physics problems\. Scrape URL:[https://huggingface\.co/datasets/UGPhysics/ugphysics\-bench](https://huggingface.co/datasets/UGPhysics/ugphysics-bench)\. Scrape date: 2026\-03\-15\. Original license: CC BY\-NC\-SA 4\.0\. Redistribution: CC BY\-NC\-SA 4\.0 carried through\. Outreach: not contacted \(license is permissive for academic redistribution under same\-license sharing\)\.
#### OpenStax College \+ University Physics\.
#### Physics Stack Exchange\.
Source: Physics Stack Exchange Q&A archive\. 2,291 problem\-and\-accepted\-answer records filtered for olympiad\-style physics\. Scrape URL:[https://physics\.stackexchange\.com/](https://physics.stackexchange.com/)via Stack Exchange data dump\. Scrape date: 2026\-01\-08\. Original license: CC BY\-SA 4\.0 \(Stack Exchange contributor agreement\)\. Redistribution: CC BY\-SA 4\.0 carried through\.
#### MMMU \+ o1\-CoT seed \(RL\-SFT seed\)\.
Source: 1,293\-record pool of MMMU physics problems augmented with o1\-style CoT solutions generated by Sonnet 4\.5\. MMMU base license: MIT\[Yueet al\.,[2024a](https://arxiv.org/html/2605.14040#bib.bib5)\]; generated CoT is our contribution\. Scrape date: 2026\-02\-10\. Redistribution: MIT \(MMMU base\) with generated CoT released under CC BY 4\.0\.
#### PhysReason\.
Source:Zhanget al\.\[[2025](https://arxiv.org/html/2605.14040#bib.bib4)\]\(ACL 2025\)\. 1,200 step\-graded physics reasoning problems\. Scrape URL:[https://huggingface\.co/datasets/PhysReason](https://huggingface.co/datasets/PhysReason)\. Scrape date: 2026\-02\-22\. Original license: CC BY 4\.0\. Redistribution: CC BY 4\.0 carried through\.
#### Estonian Physics Olympiad collection\.
Source: Estonian Physics Olympiad \(EFO\), 2004–2018\. 418 problems with organizer\-issued 1–10 difficulty labels and a 201\-problem bilingual EN\+ET subset\. Scrape URL:[https://fyysika\.ee/](https://fyysika.ee/)\(rounds: lahtine, koolivoor, vabariikilik\)\. Scrape date: 2026\-03\-30\. Provenance: publicly archived problems and solutions on the official EFO portal, released by the Estonian Physics Olympiad committee for educational use under competition policy \(consistent with international physics\-olympiad practice for IPhO, NBPhO, EuPhO, APhO, USAPhO, INPhO\)\. Redistribution: public\-domain by competition policy, used for non\-commercial research evaluation; downstream users redistributing for commercial purposes or in materially modified form should consult[https://fyysika\.ee/](https://fyysika.ee/)directly\.
#### Kevin Zhou’s olympiad handouts\.
Source: Kevin Zhou’s olympiad training documents \(USAPhO/IPhO/APhO/EuPhO content with Cambridge Tripos and graduate\-qualifier extensions\),[https://knzhou\.github\.io/](https://knzhou.github.io/)\. 692 problems with native point values 1–5 and a 3\.2% advanced \[A\] flag\. Scrape date: 2026\-02\-01\. Original license: CC BY\-NC 4\.0 \(written confirmation from Kevin Zhou \(kzhou7@gmail\.com\), Date headerSun, 3 May 2026 17:53:30 \+0800, archived in supplementary aszhou\_license\_2026\-05\-02\.eml, SHA\-2567f25c859f5ae1d790e45dbdfd23ab6be27aa1814a76de10f9bdbffb67088aba4\)\. Redistribution: the692692redistributed problems plus their reference solutions \(∼\\sim1,600 problem\-solution items in total counting solutions as separate documents\) are released under CC BY\-NC 4\.0 with attribution link to[https://knzhou\.github\.io/](https://knzhou.github.io/)preserved per record\. Third\-party content disclosure \(per Zhou’s reply\): some problems in the handouts are drawn from books or other olympiad archives, with the original source attributed inline by Zhou\. We preserve all such per\-problem internal attributions verbatim in the released records; downstream users should treat any in\-record secondary\-source attribution as binding under the original source’s terms, which may be more restrictive than CC BY\-NC 4\.0\. Outreach log: initial inquiry 2026\-04\-10 \(proposal:∼\\sim1,600 problems \+ solutions, CC BY\-NC 4\.0, attribution to source site\); written agreement to the proposed terms 2026\-05\-03 with the third\-party\-content caveat noted by Zhou\.
#### IPhO \+ NBPhO \+ EuPhO scrape\.
Sources: International Physics Olympiad \([ipho\-new\.org](https://arxiv.org/html/2605.14040v1/ipho-new.org)\), Nordic Baltic Physics Olympiad, European Physics Olympiad official archives\. 258 problems \(IPhO 66, NBPhO 165, EuPhO 27\)\. Scrape date: 2026\-03\-05\. Original license: public\-domain by competition policy \(problems published for educational use without restriction; we preserve attribution per record\)\. Redistribution: public\-domain carried through with per\-record attribution\.
#### APhO \+ USAPhO \+ INPhO scrape\.
Sources: Asian Physics Olympiad, USA Physics Olympiad \(Physics Olympiad Foundation\), Indian National Physics Olympiad\. 241 problems recovered after a multi\-pattern splitter fix that recognizesA1\.\-style markers \(USAPhO 1997–2006\) and1\. \(a\)solution markers \(INPhO\)\. The INPhO recovery alone added \+33 records relative to a single\-regex baseline\. Scrape date: 2026\-03\-12\. Original license: public\-domain by competition policy\. Redistribution: public\-domain carried through with per\-record attribution\.
Table 5:Per\-source license and provenance\. Each released record carries its source license through; non\-commercial sources \(UGPhysics, Zhou\) restrict downstream use to academic research, and the public\-domain\-by\-competition\-policy olympiad scrapes \(EFO, IPhO, NBPhO, EuPhO, APhO, USAPhO, INPhO\) preserve per\-record source attribution\.Source familyRecordsOriginal licenseOpenStax College \+ University Physics2,381CC BY 4\.0Physics Stack Exchange2,291CC BY\-SA 4\.0PhysReason1,200CC BY 4\.0MMMU \+ o1\-CoT seed \(RL\-SFT seed\)1,293MIT \(MMMU\); generated CoTIPhO \+ NBPhO \+ EuPhO scrape258Public\-domain \(competition policy\)APhO \+ USAPhO \+ INPhO scrape241Public\-domain \(competition policy\)UGPhysics5,520CC BY\-NC\-SA 4\.0Estonian Physics Olympiad418Public\-domain \(competition policy\)Kevin Zhou’s olympiad handouts692CC BY\-NC 4\.0 \(Zhou confirmed 2026\-05\-03\)Total pre\-audit14,294
## Appendix FReproducibility checklist
#### Compute budget per experiment\.
ExperimentWall\-clockHardwareCostNotesSonnet baselines \(PhyX / OlymBench /PhysOlym\-A\)∼5\{\\sim\}5hAPI∼$80\\sim\\mathdollar 80Two\-judgeOpen\-source baseline sweep on PhyX\-mini\-MC 1000q∼90\{\\sim\}90min1×\{\\times\}H200∼$5\{\\sim\}\\mathdollar 5vLLM 0\.11\.0 bf16Two\-stage contamination audit∼3\{\\sim\}3minMPS/CPUfree5 k records pairwiseThreshold\-sensitivity analysis∼3\{\\sim\}3minMPSfreesentence\-transformersPhysics\-R1 RL training to step 60\+∼30\{\\sim\}30h4×\{\\times\}H200∼$120\{\\sim\}\\mathdollar 120verl 0\.6\.1, FSDP1Reward\-component drop\-out \(follow\-up work\)∼40\{\\sim\}40h4×\{\\times\}H200∼$700\{\\sim\}\\mathdollar 700Table[11](https://arxiv.org/html/2605.14040#A8.T11)3\-seed Physics\-R1 sensitivity \(seeds 17 \+ 23 retrainings\)∼60\{\\sim\}60h4×\{\\times\}H200∼$200\{\\sim\}\\mathdollar 200added to seed\-42 in Table[3](https://arxiv.org/html/2605.14040#S5.T3)Total \(this paper\)∼$700\\sim\\mathdollar 700seeds 42/17/23 \+ Sonnet baselines \+ Sonnet\-judge runsFollow\-up budget \(reward drop\-out\)∼$1,000\\sim\\mathdollar 1\{,\}000Table[11](https://arxiv.org/html/2605.14040#A8.T11)
#### Random seeds\.
The headline single\-seed Physics\-R1 binary checkpoint uses seed4242\. The 3\-seed mean reported in Table[3](https://arxiv.org/html/2605.14040#S5.T3)aggregates seeds\{42,17,23\}\\\{42,17,23\\\}, all on the auditedPhysR1Corpcorpus under the binary correctness reward; checkpoint selection per seed uses MM\-Eureka difficulty\-curriculum saturation on the held\-out PhyX\-mini\-MC \(1,0001\{,\}000\-problem\) early\-stop signal\. The data\-build pipeline \(PhysOlym\-Asampling, train/val splits, audit pass\) usesnumpy\.random\.default\_rng\(42\)\.
#### Version pins\.
transformers==4\.57\.0\(re\-evaluation with4\.57\.6drifts results by∼1\.5\{\\sim\}1\.5points; we pin to4\.57\.0for reproducibility\),vllm==0\.11\.0,verl==0\.6\.1,sympy==1\.13\.3,sentence\-transformers==5\.4\.1,torch==2\.8\.0\+cu128\. Embedding model:mxbai\-embed\-large\-v1from MixedBread AI\.
#### Code and data hosting\.
#### Quick\-start audit\.
```
python audit/audit_two_stage.py \
--train_jsonl your_pool.jsonl --eval_jsonl data/physolym_a.jsonl \
--jaccard_thr 0.4 --cosine_thr 0.85 --emit report.json
```
Writes per\-record audit \+ aggregate3×33\{\\times\}3threshold\-sensitivity table \(Table[4](https://arxiv.org/html/2605.14040#A1.T4)\);∼3\{\\sim\}3min on MPS or a CUDA GPU for a5,0005\{,\}000\-record pool against44held\-out splits, with shingle/embedding caches inaudit\_cache/\.
#### Supplementary materials index\.
\(i\) Croissant 1\.0 \+ RAI JSON\-LD per artifact \(croissant\_rai\_<artifact\>\.json\); \(ii\) the four released datasets \(physcorp\_a\.jsonl,physr1corp\.jsonl,physolym\_a\.jsonl, plusPhysCorp\-pre\-audit\); \(iii\) audit pipeline \(audit\_two\_stage\.py\+threshold\_sensitivity\_scores\.npz\); \(iv\) reward implementation \(reward\_physics\.py\); \(v\) LLM\-judge artifacts \(judge\_prompt\.txt,rubric\.json,per\_chunk\_verdicts\.json,human\_graded\_subset\.json\); \(vi\) archived license confirmation \(zhou\_license\_2026\-05\-02\.eml, the Kevin Zhou CC BY\-NC 4\.0 grant\); \(vii\) training config \(configs/physics\-r1\.yaml\); \(viii\) checkpoints at steps\{20,40,60,80\}\\\{20,40,60,80\\\}\.
#### Reproducibility checklist\.
Code released withrequirements\.txtand deterministic build script; data released with per\-source license provenance \(Table[5](https://arxiv.org/html/2605.14040#A5.T5)\) and Croissant 1\.0\+RAI metadata; datasheet in Appendix[G](https://arxiv.org/html/2605.14040#A7); compute, seeds, and hyperparameters above \+ Table[10](https://arxiv.org/html/2605.14040#A8.T10)\. Hosting: HuggingFace \(≥\\geq5 yr\) \+ GitHub \+ Zenodo\. Maintenance: quarterly contamination audit\. Follow\-up commitments: reward\-component drop\-out ablation \(Table[11](https://arxiv.org/html/2605.14040#A8.T11)\), embedder\-sensitivity audit againstvoyage\-3andtext\-embedding\-3\-large, paraphrase/translation\-aware audit pass\.
#### Recommended baseline configuration\.
Algorithm[3](https://arxiv.org/html/2605.14040#alg3)captures the joint setting; each individual choice is small in magnitude\. Binary correctness is the recommended default; dense \(Algorithm[2](https://arxiv.org/html/2605.14040#alg2)\) is reported as an ablation\.
Algorithm 3Recommended baseline configuration for physics\-VL RL post\-training\.1:Base thinking\-mode VLM
πbase\\pi\_\{\\mathrm\{base\}\}, audited training pool
T′T^\{\\prime\}\(Algorithm[1](https://arxiv.org/html/2605.14040#alg1)\), held\-out MCQ early\-stop signal
HMCH\_\{\\mathrm\{MC\}\}, dense reward
rr\(Algorithm[2](https://arxiv.org/html/2605.14040#alg2)\)
2:Init:
πθ←πbase\\pi\_\{\\theta\}\\leftarrow\\pi\_\{\\mathrm\{base\}\}⊳\\trianglerightcold\-start from base, no SFT pass \(MM\-Eureka thesis\)
3:Optimizer: GSPO\+DAPO \(sequence\-level importance, decoupled clip\), unmodified
4:KL anchor: add
βKLDKL\(πθ∥πbase\)\\beta\_\{\\mathrm\{KL\}\}\\,D\_\{\\mathrm\{KL\}\}\(\\pi\_\{\\theta\}\\\|\\pi\_\{\\mathrm\{base\}\}\)with
βKL=10−3\\beta\_\{\\mathrm\{KL\}\}=10^\{\-3\}⊳\\trianglerightbounds drift
5:Entropy bonus: add
βHℋ\(πθ\)\\beta\_\{H\}\\,\\mathcal\{H\}\(\\pi\_\{\\theta\}\)with
βH=10−3\\beta\_\{H\}=10^\{\-3\}⊳\\trianglerightprevents entropy collapse
6:Difficulty curriculum: drop train items where the base model gets
0/N0/Nor
N/NN/Nrollouts⊳\\trianglerightMM\-Eureka\-style; preserves learning signal
7:LR schedule:
1×10−61\{\\times\}10^\{\-6\}initial, step\-cosine decay; halve LR after the first plateau on
HMCH\_\{\\mathrm\{MC\}\}⊳\\trianglerightcounters drift and length collapse
8:Response budget: fixed long\-CoT budget \(12,288 tokens for thinking\-mode\),*not adaptive*⊳\\trianglerightstable per\-step cost
9:Early stopping: stop at the saturation peak on held\-out MCQ
HMCH\_\{\\mathrm\{MC\}\},*not*on the train\-distribution validation set⊳\\trianglerightaddresses train\-up/eval\-down divergence
10:Reward:*recommended*binary correctness reward
rans∈\{0,\+1\}r\_\{\\mathrm\{ans\}\}\\in\\\{0,\+1\\\}\(simpler, fully reproducible onvllm== 0\.11\.0 multi\-image eval\); dense five\-component physics\-native reward
rr\(Algorithm[2](https://arxiv.org/html/2605.14040#alg2)\) is reported as an ablation
11:Audit gate: train pool must be Algorithm[1](https://arxiv.org/html/2605.14040#alg1)\-audited against held\-out splits*and*external benchmarks before training begins
12:Reproducibility pins:transformers== 4\.57\.0,vllm== 0\.11\.0,verl== 0\.6\.1, FSDP1 sharding for Qwen3\-VL \(§[6](https://arxiv.org/html/2605.14040#S6)\)
## Appendix GDatasheet for Physics\-R1
We followGebruet al\.\[[2021](https://arxiv.org/html/2605.14040#bib.bib21)\]with seven sections: Motivation, Composition, Collection process, Preprocessing/cleaning/labeling, Uses, Distribution, and Maintenance\. The full per\-source provenance log is in Appendix[E](https://arxiv.org/html/2605.14040#A5); the audit pipeline in Appendix[A](https://arxiv.org/html/2605.14040#A1); the LLM\-judge protocol in Appendix[D](https://arxiv.org/html/2605.14040#A4)\.
### G\.1Motivation
For what purpose was the dataset created? To support contamination\-audited evaluation and post\-training of multimodal vision\-language models on visual physics reasoning, with three specific gaps the field had not closed: \(i\) no public physics\-VL training pool was audited under a three\-stage \(n\-gram, embedding, LLM\-judge\) protocol that catches paraphrase\-class duplicates and recovers threshold\-edge topic\-similarity false positives; \(ii\) no public physics\-olympiad eval was both novel\-source and contamination\-clean against the major training\-side aggregations \(PhyX, MMMU\-Pro Physics, OlympiadBench\-Physics, UGPhysics\); \(iii\) no public physics\-VL benchmark exposed the format\-and\-novelty saturation gradient at frontier\-model scale\. Who created the dataset and on whose behalf? Shan Yang\. Who funded the creation? Self\-funded; no third\-party sponsor\. Other comments\. The corpus aggregates and audits material that already existed in scattered formats; the Estonian Physics Olympiad, Kevin Zhou’s olympiad handouts, and seven international olympiads are first\-time ML\-format releases\.
### G\.2Composition
*Instances:*one physics problem per record \(statement, optional images PNG/JPEG, gold MCQ\-letter / numeric / symbolic / multi\-part answer, optional reference solution, 14\-field annotation; schema in §[B](https://arxiv.org/html/2605.14040#A2)\)\.*Counts:*PhysCorp\-A6,432 audited;PhysR1Corp2,268 closed\-form RL pool;PhysOlym\-A500 held\-out \(stratified sample, seed 42, source\-family\-stratified\);PhysCorp\-pre\-audit14,294 raw\.*Labels:*gold answer \+ schema labels;∼\\sim3,900 records carry full Sonnet\-4\.5 annotation, rest from source\-native labels;native\_difficultypresent where organizers publish \(Estonian 27%, Zhou 38%\)\.*Splits:*see §[B](https://arxiv.org/html/2605.14040#A2)\.*Noise:*LLM\-judge unjudgeable rate13\.9%13\.9\\%onPhysOlym\-A; Stage\-1 audit misses paraphrase/translation, Stage\-2 misses numerical substitution—both reported as floors\.*Self\-contained:*yes for problems and solutions; some Zhou records carry inline secondary\-source attribution \(Appendix[E](https://arxiv.org/html/2605.14040#A5)\)\.*No PII or sensitive content\.*
### G\.3Collection process
*Acquisition:*5 repackaged benchmark releases \(UGPhysics, OpenStax, Physics Stack Exchange, MMMU\+o1\-CoT, PhysReason\) \+ 4 first\-ML scrapes by authors \(Estonian PhO, Kevin Zhou’s handouts, 7 international olympiads\); per\-source URLs and dates in Appendix[E](https://arxiv.org/html/2605.14040#A5)\.*Sampling:*per\-source complete enumeration over public archive date ranges\.*Authorship:*the author \(Shan Yang\) handled scraping/parsing/audit; Sonnet 4\.5 batch produced annotation labels\.*Time frame:*scrapes 2025\-12 to 2026\-04, audit/curation 2026\-02 to 2026\-04\.*Ethics:*no human subjects, no PII, no IRB\.*Consent:*Kevin Zhou confirmed CC BY\-NC 4\.0 redistribution of his olympiad handouts in writing 2026\-05\-03; the Estonian Physics Olympiad and other olympiad sources \(IPhO, NBPhO, EuPhO, APhO, USAPhO, INPhO\) are released under public\-domain competition policy with per\-record source attribution; repackaged sources are redistributed under their original CC/MIT licenses with attribution preserved\.
### G\.4Preprocessing, cleaning, labeling
Pre\-tokenization normalization for Stage\-1 audit \(Appendix[A](https://arxiv.org/html/2605.14040#A1)\); three\-stage audit \(Stage\-1J≥0\.4J\\geq 0\.4, Stage\-2cos≥0\.85\\cos\\geq 0\.85, Stage\-3 Haiku\-4\.5 LLM\-judge close\-duplicate vs\. same\-topic\-neighbor classification\) pairwise across PhyX, MMMU\-Pro Physics, OlympiadBench\-Physics, UGPhysics\-Train,PhysOlym\-A, with Stage\-3 close\-duplicate records removed; 14\-field schema annotation by Sonnet 4\.5 batch \(3,900 records\) \+ source\-native labels\. Raw pre\-audit pool \(PhysCorp\-pre\-audit, 14,294 records\) released alongside so users can re\-run audits\. Audit pipeline released asaudit\_three\_stage\.pywith savedbest\_jaccard/best\_cosine/judge\_labelarrays\.
### G\.5Uses
*Used for:*Physics\-R1 RL recipe \(§[4](https://arxiv.org/html/2605.14040#S4)\) trains on the audited pool and evaluates onPhysOlym\-A\+ auxiliary splits\. Supplementary archive ships paper, datasets, Croissant\+RAI metadata, \.eml license confirmations, checkpoints\.*Other use cases:*visual physics reasoning eval, contamination\-audit methodology, cross\-lingual studies \(EN/ET subset\), native\-difficulty calibration, VLM RL post\-training\.*Caveats:*cross\-lingual finding is Sonnet\-specific; dense reward is an ablation not a tuned standard\.*Out of scope:*general physics ability beyond visual reasoning, human olympiad grading substitution, experimental physics, research capability; held\-out splits must not enter pretraining/fine\-tuning\.
### G\.6Distribution
*Distribution:*HuggingFace \(≥5\\geq 5yr\) \+ GitHub \+ Zenodo DOI\. URLs:[huggingface\.co/datasets/shanyangmie/physolym\-a](https://arxiv.org/html/2605.14040v1/huggingface.co/datasets/shanyangmie/physolym-a)and[huggingface\.co/datasets/shanyangmie/physics\-r1\-corpus](https://arxiv.org/html/2605.14040v1/huggingface.co/datasets/shanyangmie/physics-r1-corpus); code at[github\.com/shanyang\-me/physics\-r1\-neurips2026](https://arxiv.org/html/2605.14040v1/github.com/shanyang-me/physics-r1-neurips2026)\. All four artifacts released alongside this paper with per\-source licenses documented in Appendix[E](https://arxiv.org/html/2605.14040#A5); the Kevin Zhou CC BY\-NC 4\.0 grant is preserved as a written \.eml in the supplementary archive, and the remaining olympiad sources are released under public\-domain competition policy\.*Licenses:*CC BY 4\.0, CC BY\-SA 4\.0, public\-domain by competition policy \(EFO, IPhO, NBPhO, EuPhO, APhO, USAPhO, INPhO\), MIT, CC BY\-NC 4\.0 \(Zhou\), CC BY\-NC\-SA 4\.0 \(UGPhysics\)—each record carries its source license; Table[5](https://arxiv.org/html/2605.14040#A5.T5)\. CC BY\-NC sources are non\-commercial; Zhou records honor inline secondary\-source attribution\. No export controls\.
### G\.7Maintenance
Maintained by the author \(Shan Yang,alexyangshan@gmail\.com\); contact via email or GitHub Issues at[github\.com/shanyang\-me/physics\-r1\-neurips2026](https://arxiv.org/html/2605.14040v1/github.com/shanyang-me/physics-r1-neurips2026)\. Versioned CHANGELOG \(v1\.0\.0initial release,v1\.1\.xadditive,v2\.0\.0schema\-breaking\); per\-record diffs per release\. Quarterly contamination audit against new physics\-VL benchmarks;≥1%\\geq 1\\%leakage triggers documented\-diff removal\. Planned follow\-ups: paraphrase/translation\-aware audit, embedder\-sensitivity ablation againstvoyage\-3/text\-embedding\-3\-large\. Hosting≥5\\geq 5yr on HF \+ Zenodo DOI per release; all versions tagged and accessible\. Contributions via GitHub Issues / PR \+ per\-recorderratum\.json, reviewed within 60 days\.
#### Machine\-readable metadata: Croissant \+ RAI JSON\-LD\.
The release ships a Croissant 1\.0\[Akhtaret al\.,[2024](https://arxiv.org/html/2605.14040#bib.bib35)\]JSON\-LD descriptor \(croissant\.json\) declaring distribution objects forPhysR1Corp,PhysOlym\-A, and the audit\-pipeline source archive, plus a 14\-fieldproblem\_recordschema and the full RAI extension \(rai:dataCollection,rai:dataAnnotationProtocol,rai:dataPreprocessingProtocol,rai:personalSensitiveInformation,rai:dataLimitations,rai:dataReleaseMaintenancePlan,rai:dataUseCases,rai:dataBiases,rai:dataSocialImpact\) covering the audit methodology, the Kevin Zhou CC BY\-NC 4\.0 written grant \(Appendix[E](https://arxiv.org/html/2605.14040#A5)\), the public\-domain\-by\-competition\-policy basis for the olympiad scrapes, three documented distributional biases \(EN/ET asymmetry, per\-physics\-category variance, difficulty\-stratified decay\), and the Stage\-1/Stage\-2 thresholds\. Passesmlcroissant validateand the MLCommons RAI checker\.
## Appendix HExtended discussion: failure modes, ethics, methodological notes, and future work
This appendix collects extended\-discussion content that was trimmed from the main body to fit the page limit\. Each subsection below corresponds to a one\-line pointer in the analysis section \(§[6](https://arxiv.org/html/2605.14040#S6)\)\.
### H\.1Failure\-mode taxonomy \(extended\)
A manual taxonomy of 100 randomly\-sampled wrong or partial Sonnet predictions on OlympiadBench\-Physics was completed post\-evaluation\. Methodology\. The 100 cases were drawn \(random seed 42\) from the 354 total wrong predictions on OlympiadBench\-Physics; each case was categorized by a single physics\-trained annotator \(graduate\-level physics background,∼\\sim3 hours of total annotation time,∼\\sim1\.8 min/case median\) using a nine\-category mutually\-exclusive rubric \(wrong\_subpart,missing\_physics,calc\_error,different\_question,valid\_partial,symbolic\_vs\_numeric,magnitude\_error,sign\_error,diagram\_misread\)\. The annotator saw the problem statement, the gold solution, the gold final answer, and the Sonnet response; they did not see Sonnet’s confidence or the verdict from the LLM judge\. We report this single\-annotator taxonomy as a methodological diagnostic,*not*as a calibrated human grade; a second\-annotator pass on a 50\-problem random subsample with inter\-annotatorκ\\kappaper category is left to follow\-up work \(distinct from the audit\-flag inter\-annotatorκ\\kappapre\-registered in Appendix[A](https://arxiv.org/html/2605.14040#A1), which targets contamination\-flag agreement, not failure\-category agreement\)\.
Table 6:Failure\-mode taxonomy of 100 randomly\-sampled Sonnet wrong/partial predictions on OlympiadBench\-Physics\. Categories are mutually exclusive; each prediction received one label\.Categoryn%Descriptionwrong\_subpart3030%Answered a neighboring sub\-question instead of the one askedmissing\_physics2222%Wrong physical model, law, or geometry; crucial effect absentcalc\_error1616%Correct approach; small numerical slip \(factor 2–4 off\)different\_question1010%Reasoning chain for an unrelated problem in the same promptvalid\_partial88%Right approach with arithmetic slip; partial\-credit territorysymbolic\_vs\_numeric77%Formula returned where a number was required, or vice versamagnitude\_error55%Correct expression, wrong order of magnitude \(≥10×\\geq 10\\times\)sign\_error00%None observeddiagram\_misread00%None observed \(text\-only evaluation; figures unavailable\)Total100100%#### Reading the taxonomy\.
wrong\_subpartdominates \(30%\): multi\-part olympiad prompts containing \(a\)/\(b\)/\(c\) sub\-questions where the model fluently answers a neighboring sub\-question instead of the asked one—a 5×\\timesunderestimate by regex\-only flagging \(∼\\sim6% in §[5](https://arxiv.org/html/2605.14040#S5.SS0.SSS0.Px1)\)\.missing\_physics\(22%\) anddifferent\_question\(10%\) together account for nearly a third of failures and represent a real reasoning floor unlikely to be format\-induced\.calc\_error\(16%\) andvalid\_partial\(8%\) are correct\-approach/failed\-execution, matching the 4\.7\-pp strict\-vs\-liberal gap \(§[5\.1](https://arxiv.org/html/2605.14040#S5.SS1.SSS0.Px1)\)\.sign\_erroranddiagram\_misreadare zero \(the latter because OlympiadBench\-Physics is text\-only in the public release\)\.
#### Per\-physics\-category breakdown\.
Table 7:Per\-physics\-category accuracy on OlympiadBench\-Physics \(Sonnet 4\.5, strict\)\. The\+34\.5\+34\.5\-pp EM \(38\.4%38\.4\\%\) vs\. astrophysics \(72\.9%72\.9\\%\) gap on identical weights motivates per\-category reporting\.CategorynAcc\.CategorynAcc\.CategorynAcc\.Electromagnetism23738\.4%Relativity8951\.7%Other1060\.0%Quantum12841\.4%Classical mech\.1957\.9%Astrophysics7072\.9%Waves3946\.2%Thermodynamics8759\.8%All692—Fluid mechanics1346\.2%
### H\.2Ethical considerations and intended use
#### Intended and out\-of\-scope use\.
The released artifacts support research on visual physics reasoning, contamination\-audit methodology, cross\-lingual LLM evaluation, and RL post\-training of VLMs\.PhysOlym\-Ais held\-out: please treat it as test\-only, not for pretraining or fine\-tuning\. Out of scope: general physics ability beyond visual reasoning, human olympiad grading substitution, experimental\-physics evaluation, paper\-writing, derivation novelty, or open\-ended hypothesis generation\.
#### Source provenance, consent, and misuse risk\.
Kevin Zhou confirmed CC BY\-NC 4\.0 redistribution of his olympiad handouts in writing on 2026\-05\-03 \(∼\\sim1,600 problem\-solution items, Appendix[E](https://arxiv.org/html/2605.14040#A5)\); the Estonian Physics Olympiad collection and the six international olympiad scrapes \(IPhO, NBPhO, EuPhO, APhO, USAPhO, INPhO\) are redistributed under public\-domain competition policy with per\-record source attribution\. No PII; problems are olympiad/textbook content\. Public\-domain olympiad scrapes are released by competition policy\. The dense reward and Physics\-R1 recipe ship for reproducibility, not as a tuned gold standard; the audit pipeline is a measurement tool, not a certification authority, and we disclose threshold sensitivity \(Table[4](https://arxiv.org/html/2605.14040#A1.T4)\) so users do not over\-claim “contamination\-free\.”
#### Caveats and compute\.
The 17\-pp EN/ET cross\-lingual delta is Sonnet\-4\.5\-specific onn=59n\{=\}59paired items; direction may flip on low\-resource\-language\-weak open\-source models—treat as model\-specific\. Total compute∼\\sim360 H200\-GPU\-hours across the 3 seeds in Table[3](https://arxiv.org/html/2605.14040#S5.T3)\(seed\-42 at3030h×4\\times 4H200; seed\-17 and seed\-23 retrainings∼60\\sim 60h×4\\times 4H200 combined; Appendix[F](https://arxiv.org/html/2605.14040#A6)\), plus frontier\-API inference for baselines and Sonnet\-judge runs\. Per\-experiment carbon estimates are omitted because cloud\-provider electricity mix is not consistently disclosed\.
### H\.3Construction\-process disclosures
\(i\) The initialPhysOlym\-Adescription claimed 100% novel\-source; the Stage\-1 audit surfaced one EuPhO 2020 overlap with OlympiadBench\-Physics \(J=0\.91J\{=\}0\.91\), so the honest claim is99\.8%99\.8\\%\(499/500499/500\)—disclosed, not dropped\. \(ii\) Our initial header describedPhyX\-mini\-MCas 500 problems; the canonical MC subset\[Shenet al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib1)\]is 1,000\.
### H\.4Future work and named follow\-ups
Ten follow\-ups consolidated\.*Eval refinements:*\(i\) post\-hoc MCQ\-ification ofPhysOlym\-Ato isolate the format axis \(§[5\.1](https://arxiv.org/html/2605.14040#S5.SS1.SSS0.Px1)\); \(ii\) paraphrase\- and translation\-aware audit pass\.*Cross\-corpus:*\(iii\) audit OlympiadBench\-Physics against our pool, then add the Physics\-R1 row; \(iv\) frontier cross\-evaluation against PHYBench, PhysUniBench, HLE\-physics,PhysOlym\-A\.*Recipe and scale:*\(v\) transfer to InternVL3\-8B and LLaVA\-OneVision\-7B; \(vi\) 32B \+ full\-RL comparator; \(vii\) SFT\-only data\-scaling curve at500/1,293/5,000/9,575500/1\{,\}293/5\{,\}000/9\{,\}575audited prompts\.*Cross\-lingual:*\(viii\) confirm/refute EN/ET sign\-flip on low\-resource open\-source models on the same 59\-pair Estonian Physics Olympiad subset;*pre\-registered*—if 8B\-class open\-source models with documented Estonian\-weak training \(e\.g\. Qwen2\.5\-VL\-7B, LLaVA\-OV\-7B\) show ET<<EN by≥5\\geq 5pp under paired sign test, F2’s Sonnet\-4\.5 ET\>\>EN direction is model\-specific \(confirmed\); a≥5\\geq 5pp result in the same direction \(ET\>\>EN\) on those same models would replicate F2 across model families; results within±5\\pm 5pp are inconclusive atn=59n\{=\}59\.*Methodology:*\(ix\) 50\-problem inter\-annotatorκ\\kappaon the failure\-mode taxonomy; \(x\) embedder\-sensitivity ablation againstvoyage\-3andtext\-embedding\-3\-large\.
### H\.5Versioning and maintenance commitment
Versioned releases with semantic\-version tags \(v1\.0\.0initial,v1\.1\.xadditive,v2\.0\.0schema\-breaking\)\. Each release ships Croissant 1\.0 \+ RAI JSON\-LD, per\-source license matrix \(Table[5](https://arxiv.org/html/2605.14040#A5.T5)\), threshold\-sensitivity grid \(Table[4](https://arxiv.org/html/2605.14040#A1.T4)\) recomputed against newly\-released physics\-VL benchmarks, and the audit\-pipeline source archive with saved best\-overlap scores\. Quarterly contamination audit against new benchmarks; benchmarks introducing≥1%\\geq 1\\%leakage to any held\-out split trigger a documented\-diff removal in the next minor release\. Hosting: HuggingFace \(≥5\\geq 5yr\), GitHub, Zenodo DOI per release\.
### H\.6Method comparison and benchmark comparison tables \(extended\)
Table 8:Physics\-R1 vs\. rule\-based RL recipes for thinking\-mode VLMs\.Init: base or SFT cold\-start\. Reward signals:*Ans*\(answer\),*Fmt*\(format\),*Dim*\(units\),*Sym*\(symbolic\),*Cons*\(conservation\)\. Audit and Filter as defined in §[3\.3](https://arxiv.org/html/2605.14040#S3.SS3)\.✓\\checkmarkpresent; — absent;∘\\circpartial\.RecipeInitAnsFmtDimSymConsAuditFilterDPO\[Rafailovet al\.,[2023](https://arxiv.org/html/2605.14040#bib.bib32)\]SFT———————GRPO\[Shaoet al\.,[2024](https://arxiv.org/html/2605.14040#bib.bib10)\]SFT✓\\checkmark——————GSPO\+DAPO\[Zhenget al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib12), Yuet al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib11)\]SFT✓\\checkmark∘\\circ—————MM\-Eureka\[Menget al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib33)\]base✓\\checkmark✓\\checkmark—————Physics\-R1base✓\\checkmark✓\\checkmark✓\\checkmark✓\\checkmark✓\\checkmark✓\\checkmark✓\\checkmark
### H\.7Corpus composition, hyperparameters, and reward ablation \(extended\)
Table 9:PhysCorp\-pre\-auditcomposition by source family \(14,29414\{,\}294records total before two\-stage audit\)\.The audited releasePhysCorp\-A\(6,4326\{,\}432records\) is the subset surviving Algorithm[1](https://arxiv.org/html/2605.14040#alg1)after a re\-audit pass against PhysReason\-full and PhysUniBench\-en dropped an additional 804 records from a7,2367\{,\}236\-record candidate \(the construction audit excluded those two corpora\)\. The 804\-record re\-audit drop concentrates on PhysReason\-cousin records \(540\) and PhysUniBench\-en\-cousin records \(186\), drawn predominantly from the*repackaged\-benchmark*source families \(UGPhysics / OpenStax / Physics SE / MMMU\+o1\-CoT seed / PhysReason\); the four*first\-ML\-format*source families \(Estonian PhO, Zhou’s handouts, the 7\-international\-olympiad scrapes\) are minimally affected and retain their full1,6091\{,\}609\-record contribution toPhysCorp\-A\. All 9 source families remain represented in the released pool\. Per\-source licenses are carried through to each released record; full provenance and outreach log in Appendix[E](https://arxiv.org/html/2605.14040#A5)\.Source familyCountLicenseAnswer typeUGPhysics5,520CC BY\-NC\-SAOpen\-ended numericOpenStax Physics2,381CC BY 4\.0NumericPhysics Stack Exchange2,291CC BY\-SA 4\.0EquationsRL\-SFT seed \(MMMU \+ o1 CoT\)1,293MIT \(MMMU\); generated CoTVariedPhysReason1,200CC BY 4\.0CoT \+ open\-endedZhou Olympiad Handouts692CC BY\-NC 4\.0†MixedEstonian Physics Olympiad418Public\-domain \(competition policy\)Open\-endedIPhO \+ NBPhO \+ EuPhO scrape258Public domainOpen\-endedAPhO \+ USAPhO \+ INPhO scrape241Public domainOpen\-endedTotal novel1,609——Total corpus14,294——
†\\daggerAuthor granted explicit redistribution permission with attribution \(email confirmation, May 2026\); we redistribute under CC BY\-NC 4\.0\.
Table 10:Physics\-R1 training configuration\.GSPO\+DAPO viaverl0\.6\.1 with vLLM 0\.11\.0 \(TP=4\) for rollouts\. FSDP1 is required for Qwen3\-VL underverl0\.6\.1 \(FSDP2 fails on the multimodal projector path; reproducibility note in §[6](https://arxiv.org/html/2605.14040#S6)\)\. Cells marked*\(recipe\)*differ from a default unconstrained GSPO\+DAPO configuration; together with the dense physics\-native reward \(Section[4](https://arxiv.org/html/2605.14040#S4)\), the audited training pool, and held\-out early stopping, they constitute the Physics\-R1 reference recipe\.ParameterValueRole in the recipeBase / init*\(recipe\)*Qwen3\-VL\-8B\-Thinking BASE \(*no SFT*\)Cold\-start RL, mirroring MM\-Eureka thesisAlgorithmGSPO \+ DAPOStable group\-policy backboneImportance sampling*\(recipe\)*truncated, sequence\-levelRollout\-vs\-train policy correctionLearning rate1×10−61\\\!\\times\\\!10^\{\-6\}Standard for 8B\-class RLLR decay*\(recipe\)*step\-cosine;0\.5×0\.5\\\!\\timesafter step 60Counters drift and length collapseBatch size96Effective batch via gradient accumulationRollouts per prompt16DAPO group sizeMax response length12,288 tokensLong\-CoT thinking budgetKL anchor*\(recipe\)*1×10−31\\\!\\times\\\!10^\{\-3\}to baseAnchors policy to base; bounds driftEntropy bonus*\(recipe\)*1×10−31\\\!\\times\\\!10^\{\-3\}Prevents entropy collapseClip range \(decoupled\)0\.2 / 0\.28DAPO decoupled clipDifficulty curriculum*\(recipe\)*drop 0/16 and 16/16 rollout promptsMM\-Eureka\-style; preserves learning signalReward*\(recipe\)*binary correctness \(§[4](https://arxiv.org/html/2605.14040#S4)\)Recommended default; dense \(Algo\.[2](https://arxiv.org/html/2605.14040#alg2)\) is the shape ablationTrain pool*\(recipe\)*2,268PhysR1CorppromptsClosed\-form carve\-out ofPhysCorp\-A\(§[3](https://arxiv.org/html/2605.14040#S3)\)Early stop*\(recipe\)*held\-out PhyX\-mini\-MCCatches the saturation peakSharding*\(recipe\)*FSDP1 \(not FSDP2\)Required for Qwen3\-VL underverl0\.6\.1Hardware4×\\timesH200 \(single node\)—Step time∼40\{\\sim\}40minutes/stepRollout\-bound \(gen∼65%\{\\sim\}65\\%, update∼23%\{\\sim\}23\\%\)Wall\-clock to saturation peak∼30\{\\sim\}30–4545hours \(steps 60–80\)Training cost∼$120\{\\sim\}\\mathdollar 120–$180\\mathdollar 180transformers4\.57\.0 \(pinned\)4\.57\.6 yields different forward pathvllm0\.11\.0TP=4 rollout backendverl0\.6\.1Training frameworkRandom seed \(headline\)423\-seed sweep\{17,23,42\}\\\{17,23,42\\\}reported as the headline binary row in Table[3](https://arxiv.org/html/2605.14040#S5.T3)Table 11:Physics\-R1 reward\-shape ablation on PhyX\-mini\-MC\.Each row toggles one of the five reward components on or off, training Qwen3\-VL\-8B\-Thinking\-base under the same GSPO\+DAPO\+KL\-anchor recipe \(Equation[1](https://arxiv.org/html/2605.14040#S4.E1)\) for the same step budget\. Ans = answer\-correctness binary \(\+1\+1,≡rbin\\equiv r\_\{\\mathrm\{bin\}\}of §[4](https://arxiv.org/html/2605.14040#S4)\); Fmt =\\boxed\{\}format \(\+0\.1\+0\.1\); Dim = dimensional consistency from regex\-detected units \+ sympy unit\-system \(\+0\.15\+0\.15\); Sym = symbolic equation verification of intermediate\\fracexpressions via sympy \(\+0\.20\+0\.20\); Cons = conservation\-law penalty \(energy/momentum\) when applicable \(−0\.25\-0\.25\)\. Composed reward is clipped to\[−1,1\]\[\-1,1\]\. Init in every cell is the Qwen3\-VL\-8B\-Thinking BASE checkpoint \(*no SFT*; cold\-start, mirroring the MM\-Eureka thesis\)\. The Ans\-only row is the recommended Physics\-R1 recipe \(binary correctness reward, §[4](https://arxiv.org/html/2605.14040#S4)\); the all\-on row is the dense ablation of §[4](https://arxiv.org/html/2605.14040#S4); the intermediate rows isolate the marginal effect of each physics\-native shaping component\. Drop\-out cells \(—\) are intermediate\-component runs left to follow\-up work \(compute estimate in Appendix[F](https://arxiv.org/html/2605.14040#A6)\)\.ConfigurationAnsFmtDimSymConsPhyX\-mini\-MCBase \(no RL\)—————73\.7%Physics\-R1 \(binary, recommended\)≡\\equivAns\-only✓————78\.0\+ Format✓✓————\+ Dim✓✓✓———\+ Sym✓✓✓✓——Physics\-R1 \(dense, ablation, all\-on\)✓✓✓✓✓78\.3*Single\-component drop\-outs \(dense ablation minus one\)*−\-Dim✓✓—✓✓—−\-Sym✓✓✓—✓—−\-Cons✓✓✓✓——Table 12:Sonnet 4\.5 strict accuracy on Estonian Physics Olympiad problems by organizer\-issued native difficulty \(n=131\)\. The curve is near\-monotonically decreasing in difficulty and hits a hard floor of 0% at difficulties 3, 6, 8, and 10\. Non\-monotone bumps at 4 and 5 are within sampling noise on small per\-bin counts\.DifficultynCorrectAcc\.1171062\.5%215320\.0%3800\.0%412325\.0%5271037\.0%6900\.0%714321\.4%8900\.0%915321\.4%10500\.0%All1313224\.4%Table 13:Cross\-lingual Sonnet performance on the Estonian Physics Olympiad bilingual subset \(n=59, identical problems, same judge protocol\)\. Strict% = numeric/symbolic match; Liberal% = LLM\-judge score≥0\.5\\geq 0\.5\.LanguagenCorrectPartialIncorrectUnjudgeableStrict%Liberal%English \(translated\)598843013\.6%20\.3%Estonian \(original\)5918927530\.5%38\.1%Table 14:Per\-problem agreement matrix for the EN/ET cross\-lingual ablation \(n=59\)\. The 4\.3:1 asymmetry in the off\-diagonal cells \(ET\-correct/EN\-wrong vs\. EN\-correct/ET\-wrong\) rules out a noise explanation\.ET correctET wrongEN correct5 \(both correct\)3 \(EN only\)EN wrong13 \(ET only\)38 \(both wrong\)Similar Articles
Physics-IQ Verified
This paper presents a systematic audit of the Physics-IQ benchmark for evaluating physical understanding in video generative models, proposing improvements to prompts and scoring to enhance reliability.
OmniPhys: A Unified Multimodal Benchmark for Physics Understanding and Generation from Chinese Educational Corpora
OmniPhys is a large-scale multimodal benchmark for physics understanding and generation, covering middle school to university-level problems from Chinese educational corpora, aimed at evaluating and advancing multimodal large language models in scientific domains.
SeePhys Pro: Diagnosing Modality Transfer and Blind-Training Effects in Multimodal RLVR for Physics Reasoning
The paper introduces SeePhys Pro, a benchmark to diagnose modality transfer issues in multimodal RL for physics reasoning, revealing that models struggle with representation-invariant reasoning and often rely on residual textual cues rather than visual evidence.
Physics Question Scene Graph: Fine-grained Evaluation of Physical Plausibility in Text-to-Video Generation
Physics Question Scene Graph (PQSG) is a hierarchical question-based pipeline using VLMs to evaluate video generation models' physical plausibility with fine-grained violation detection. It introduces the FinePhyEval dataset and shows higher correlation with human judgments than prior work.
BilliardPhys-Bench: Benchmarking Physical Reasoning and Visual Dynamics of Multimodal LLMs
BilliardPhys-Bench is a new benchmark that tests multimodal LLMs on physical reasoning using synthetic billiards scenarios, requiring predictions of collisions and final ball positions. The paper finds that current models struggle with longer simulations and exhibit a 'stasis bias' of predicting no interaction when uncertain.