Physics-R1: An Audited Olympiad Corpus and Recipe for Visual Physics Reasoning

arXiv cs.CL Papers

Summary

This paper audits multimodal physics evaluation pipelines, revealing issues like train-eval contamination, translation drift, and MCQ saturation. It releases new datasets (PhysCorp-A, PhysR1Corp, PhysOlym-A) and a training recipe (Physics-R1) that significantly improves performance on held-out olympiad problems.

arXiv:2605.14040v1 Announce Type: new Abstract: We audit the multimodal-physics evaluation pipeline end-to-end and document three undetected construction practices that distort how the field measures vision-language reasoning: train-eval contamination, translation drift, and MCQ saturation. (1) Public training pools (UGPhysics-Train, SciInstruct, MMK12) pass single-stage 5-gram-Jaccard audits with zero hits across all six public physics evals; a three-stage audit (Jaccard -> mxbai-embed-large cosine -> Haiku-4.5 LLM-judge) surfaces 134 near-duplicates and 4,846 paraphrase candidates in SciInstruct alone. (2) A 17-pp Sonnet 4.5 delta on 59 paired Estonian-English olympiad problems (30.5% vs. 13.6%; sign test p=0.011, McNemar p=0.021, paired bootstrap 95% CI [+5.1, +28.9] pp). (3) A 46-pp format-and-novelty gradient on identical Sonnet weights between MCQ (79.7% on PhyX) and open-ended olympiad evaluation (33.4% on PhysOlym-A). We release four artifacts addressing these gaps: PhysCorp-A (6,432-record three-stage-audited multimodal corpus), PhysR1Corp (2,268-record closed-form RL pool), PhysOlym-A (500-problem, 99.8% novel-source held-out olympiad eval with native difficulty labels and an EN/ET bilingual subset), and Physics-R1, a reference GSPO+DAPO recipe cold-started from Qwen3-VL-8B-Thinking. Across 3 seeds, Physics-R1 lifts the audited corpus over the 8B base by +18.3 pp on PhysOlym-A liberal (8.0 -> 26.3 +/- 1.7; 7.1 pp behind Sonnet 4.5), +15.7 pp on PhysReason (23.9 -> 39.6 +/- 6.4; ahead of Qwen3-VL-32B and Gemini 2.5 Pro), +6.9 pp on OlympiadBench-Physics (46.2 +/- 1.5), and +4.1 pp on PhyX MCQ (77.8 +/- 0.3).
Original Article
View Cached Full Text

Cached at: 05/15/26, 06:18 AM

# An Audited Olympiad Corpus and Recipe for Visual Physics Reasoning
Source: [https://arxiv.org/html/2605.14040](https://arxiv.org/html/2605.14040)
###### Abstract

We audit the multimodal\-physics evaluation pipeline end\-to\-end and document three undetected construction practices that distort how the field measures vision\-language reasoning: train–eval contamination, translation drift, and MCQ saturation\. \(1\) Public training pools \(UGPhysics\-Train, SciInstruct, MMK12\) pass single\-stage 5\-gram\-Jaccard audits with zero hits across all six public physics evals; a three\-stage audit \(Jaccard→\\tomxbai\-embed\-largecosine→\\toHaiku\-4\.5 LLM\-judge\) surfaces𝟏𝟑𝟒\\mathbf\{134\}near\-duplicates and𝟒,𝟖𝟒𝟔\\mathbf\{4\{,\}846\}paraphrase candidates in SciInstruct alone\. \(2\) A 17\-pp Sonnet\-4\.5\(Anthropic,[2025](https://arxiv.org/html/2605.14040#bib.bib53)\)delta on 59 paired Estonian\-English olympiad problems \(30\.5%30\.5\\%vs\.13\.6%13\.6\\%; sign testp=0\.011p\{=\}0\.011, McNemarp=0\.021p\{=\}0\.021, paired bootstrap95%95\\%CI\[\+5\.1,\+28\.9\]\[\+5\.1,\+28\.9\]pp\)\. \(3\) A 46\-pp format\-and\-novelty gradient on identical Sonnet weights between MCQ \(79\.7%79\.7\\%on PhyX\) and open\-ended olympiad evaluation \(33\.4%33\.4\\%onPhysOlym\-A\)\. We release four artifacts addressing these gaps:PhysCorp\-A\(6,4326\{,\}432\-record three\-stage\-audited multimodal corpus\),PhysR1Corp\(2,2682\{,\}268\-record closed\-form RL pool\),PhysOlym\-A\(500500\-problem,99\.8%99\.8\\%novel\-source held\-out olympiad eval with native difficulty labels and an EN/ET bilingual subset\), and Physics\-R1, a reference GSPO\+DAPO recipe cold\-started from Qwen3\-VL\-8B\-Thinking\. Across33seeds \(§[5](https://arxiv.org/html/2605.14040#S5)\), Physics\-R1 lifts the audited corpus over the 8B base by\+18\.3\+18\.3pp onPhysOlym\-Aliberal \(8\.0→26\.3±1\.78\.0\{\\to\}\\mathbf\{26\.3\}\{\\pm\}1\.7;7\.17\.1pp behind Sonnet 4\.5\),\+15\.7\+15\.7pp on PhysReason \(23\.9→39\.6±6\.423\.9\{\\to\}\\mathbf\{39\.6\}\{\\pm\}6\.4; ahead of Qwen3\-VL\-32B and Gemini 2\.5 Pro\),\+6\.9\+6\.9pp on OlympiadBench\-Physics \(46\.2±1\.5\\mathbf\{46\.2\}\{\\pm\}1\.5\), and\+4\.1\+4\.1pp on PhyX MCQ \(77\.8±0\.3\\mathbf\{77\.8\}\{\\pm\}0\.3\)\.

## 1Introduction

Multimodal physics reasoning is increasingly tracked via vision\-language benchmarks, but how those benchmarks are constructed is rarely audited\. Researcher\-curated training pools aggregate physics problems from publicly available sources whose paraphrase relationships evade conventional n\-gram dedup; multilingual benchmarks distribute English translations of problems first composed in another language; MCQ\-format splits saturate against the closed\-frontier ceiling\. Each represents a methodological gap in how the field constructs benchmarks, and together they distort cross\-model comparisons, inflate frontier\-model rankings on public leaderboards, and obscure the format\-and\-novelty axis along which capability actually diverges\.

We argue that defensible measurement of multimodal physics reasoning requires an end\-to\-end audit of the evaluation pipeline\. This paper performs that audit, surfaces three measurement findings, and constructs released artifacts directly against the gap each finding identifies\. Physics\-R1, a reference GSPO\+DAPO recipe\(Zhenget al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib12); Yuet al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib11)\)cold\-started from Qwen3\-VL\-8B\-Thinking\(Qwen Team,[2025](https://arxiv.org/html/2605.14040#bib.bib17)\)and building on MM\-Eureka\(Menget al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib33)\)and DeepSeek\-R1’s binary correctness signal\(DeepSeek\-AI,[2025](https://arxiv.org/html/2605.14040#bib.bib34); Shaoet al\.,[2024](https://arxiv.org/html/2605.14040#bib.bib10)\), accompanies the corpus as evidence\-of\-trainability rather than as the primary contribution: it lifts the audited held\-out eval over the 8B base while still trailing the closed frontier \(§[5\.2](https://arxiv.org/html/2605.14040#S5.SS2.SSS0.Px5)\)\.

#### Finding 1: single\-stage 5\-gram\-Jaccard audit reports public physics\-VL training pools as clean, but a three\-stage audit \(Jaccard→\\tomxbai cosine→\\toLLM\-judge\) surfaces𝟏𝟑𝟒\\mathbf\{134\}near\-duplicates among4,8464\{,\}846Stage\-2 candidates in SciInstruct alone\.

Across the three published physics\-VL training pools we re\-audit against six public evals \(UGPhysics\-Train, SciInstruct’s 42K\-record en\_phy\_chem split, MMK12’s 15K\-record train pool\), conventional 5\-gram\-Jaccard atJ≥0\.4J\\geq 0\.4\(Stage\-1\) reports*zero*hits for every pool against all six evals—a single\-stage audit calls them all clean\. Stage\-2mxbai\-embed\-largecosine at≥0\.85\\geq 0\.85then surfaces𝟒,𝟖𝟒𝟔\\mathbf\{4\{,\}846\}paraphrase\-class candidate pairs from SciInstruct alone \(PhysReason\-full2,6872\{,\}687, PhysUniBench\-en1,0271\{,\}027dominant\),99from UGPhysics\-Train, and6666from MMK12 \(Table[2](https://arxiv.org/html/2605.14040#S3.T2)\)\. Stage\-3, a Haiku\-4\.5 LLM\-judge, classifies each Stage\-2 candidate as a*close duplicate*or a*same\-topic neighbor*: of the4,8464\{,\}846SciInstruct candidates,𝟏𝟑𝟒\\mathbf\{134\}\(2\.8%2\.8\\%\) are close duplicates and the duplicate fraction is sharply cosine\-driven \(100%100\\%atcos≥0\.95\\cos\\geq 0\.95,1\.5%1\.5\\%atcos∈\[0\.85,0\.87\)\\cos\\in\[0\.85,0\.87\)\)\. On a1,6791\{,\}679\-record researcher\-curated sample ofPhysCorp\-pre\-audit\(14,29414\{,\}294records\) under the field\-default within\-pool dedup workflow,345345records \(20\.5%\\mathbf\{20\.5\\%\}\) leak at Stage\-1 alone against the six public evals \(concentrated in PhysUniBench\-en,339339, and MMMU\-Pro Physics,2020\); the joint Stage\-1∨\\veeStage\-2 sweep on this same sample against an internal analysis eval reaches8\.8%\\mathbf\{8\.8\\%\}at the published operating point and27\.1%27\.1\\%atcos≥0\.80\\cos\\geq 0\.80\(Table[4](https://arxiv.org/html/2605.14040#A1.T4)\)\.

#### Finding 2: translation introduces a measurable score delta on identical physics problems\.

On 59 paired Estonian/English Physics Olympiad problems, Sonnet 4\.5\(Anthropic,[2025](https://arxiv.org/html/2605.14040#bib.bib53)\)attains30\.5%\\mathbf\{30\.5\\%\}strict on Estonian originals against only13\.6%\\mathbf\{13\.6\\%\}on English translations of the same problems \(sign test on 16 discordant pairsp=0\.011p\{=\}0\.011; McNemar exactp=0\.021p\{=\}0\.021; bootstrap95%95\\%CI\[\+5\.1,\+28\.9\]\[\+5\.1,\+28\.9\]pp\)\. Estonian PhO problems were composed in Estonian first; English versions are translations whose physics vocabulary, grammatical case mapping, and subtlety of scope degrade information content\. For Sonnet 4\.5, whose cross\-lingual transfer covers Estonian, published numbers on the English\-translation benchmark systematically*underestimate*model ability relative to original\-language gold; for models with weaker training in the original language, the relationship is expected to reverse \(App\.[H\.4](https://arxiv.org/html/2605.14040#A8.SS4)\(viii\), pre\-registered\) \(§[3\.2](https://arxiv.org/html/2605.14040#S3.SS2), §[5\.1](https://arxiv.org/html/2605.14040#S5.SS1.SSS0.Px3)\)\.

#### Finding 3: same\-model evaluation across three physics benchmarks reveals a 46\-point format\-and\-novelty gradient\.

Evaluated in the same week on identical Sonnet 4\.5 weights, the score sweeps from79\.7%\\mathbf\{79\.7\\%\}on PhyX\(Shenet al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib1)\)\(4\-way MCQ\) down to50\.4%\\mathbf\{50\.4\\%\}liberal on OlympiadBench\-Physics\(Heet al\.,[2024](https://arxiv.org/html/2605.14040#bib.bib2)\)and33\.4%\\mathbf\{33\.4\\%\}liberal on our held\-out audited eval—format\-and\-novelty alone move the score by4646points on fixed weights \(§[3\.2](https://arxiv.org/html/2605.14040#S3.SS2); scoring in §[5](https://arxiv.org/html/2605.14040#S5)\)\.

Together the three findings imply that defensible physics\-VL measurement requires three properties at construction time: a three\-stage audit \(n\-gram Jaccard→\\toembedding cosine→\\toLLM\-judge precision filter\), original\-language gold, and open\-ended novel\-source evaluation\. Four released artifacts instantiate this protocol: \(a\)PhysCorp\-A, the audited multimodal physics corpus produced by the three\-stage pipeline \(Algorithm[1](https://arxiv.org/html/2605.14040#alg1)\), and the closed\-form RL training poolPhysR1Corpon which Physics\-R1 is trained \(§[3](https://arxiv.org/html/2605.14040#S3)\); \(b\)PhysOlym\-A, the open\-ended held\-out olympiad benchmark with native difficulty calibration, an EN/ET bilingual subset, and a Sonnet\-as\-judge protocol whose unjudgeable rate \(13\.9%13\.9\\%\) we disclose \(§[3\.2](https://arxiv.org/html/2605.14040#S3.SS2), §[5\.1](https://arxiv.org/html/2605.14040#S5.SS1)\); \(c\) Physics\-R1, a reference RL recipe whose audited held\-out lift onPhysOlym\-Avalidates the corpus as trainable rather than memorized \(Table[3](https://arxiv.org/html/2605.14040#S5.T3)\); we recommend a binary correctness reward as the default—variance\-optimal under GSPO with group\-normalized advantages, Goodhart\-robust against unit/conservation/format proxies, and harness\-portable \(§[4](https://arxiv.org/html/2605.14040#S4), properties P1–P4\)—and report the dense five\-component physics\-native reward as a shape ablation; and \(d\) the audit protocol itself, released asaudit\_three\_stage\.pywith saved best\-overlap scores and Stage\-3 judge labels \(Appendix[A](https://arxiv.org/html/2605.14040#A1)\)\. The 3\-seed sensitivity sweep \(seeds\{42,17,23\}\\\{42,17,23\\\}on the auditedPhysR1Corp\) is reported in Table[3](https://arxiv.org/html/2605.14040#S5.T3)withσ≤3\.3\\sigma\\leq 3\.3pp on PUB\-OE, OlymBench\-Phys, andPhysOlym\-A, andσ=6\.4\\sigma\{=\}6\.4pp on PhysReason \(seed\-42 outlier\); the reward\-component drop\-out ablation \(Table[11](https://arxiv.org/html/2605.14040#A8.T11)\) is left to follow\-up work\.

## 2Related Work

#### Rule\-based RL for reasoning\.

DeepSeek\-R1\(DeepSeek\-AI,[2025](https://arxiv.org/html/2605.14040#bib.bib34)\)established that simple rule\-based rewards \(binary correctness \+ format\) suffice to train competitive math reasoners directly from a base model without SFT, using GRPO\(Shaoet al\.,[2024](https://arxiv.org/html/2605.14040#bib.bib10)\)\. MM\-Eureka\(Menget al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib33)\)extended the recipe to VLMs with a difficulty curriculum; DAPO\(Yuet al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib11)\)added decoupled clipping and dynamic sampling; GSPO\(Zhenget al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib12)\)replaced token\-level with sequence\-level importance weighting\. Physics\-R1 inherits MM\-Eureka’s structural choices and the binary correctness reward unchanged: although physics intermediate steps carry units, conservation laws, and symbolic equations that*a priori*admit per\-step verification, we find that under GSPO with group\-normalized advantages a binary reward is variance\-optimal and robust to the within\-wrong\-group Goodhart channel that physics\-native shaping opens \(§[4](https://arxiv.org/html/2605.14040#S4)\); the dense physics\-native reward is reported as an ablation\.

#### Physics QA benchmarks\.

PhyX\(Shenet al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib1)\), OlympiadBench\-Physics\(Heet al\.,[2024](https://arxiv.org/html/2605.14040#bib.bib2)\), UGPhysics\(Xuet al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib3)\), PhysReason\(Zhanget al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib4)\), MMMU/MMMU\-Pro\(Yueet al\.,[2024a](https://arxiv.org/html/2605.14040#bib.bib5),[b](https://arxiv.org/html/2605.14040#bib.bib41)\), MMK12\(Menget al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib33)\), PHYBench\(Qiuet al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib36)\), and PhysUniBench\(Wanget al\.,[2025b](https://arxiv.org/html/2605.14040#bib.bib37)\)are the canonical references\. Top entries cluster within ten points of the closed\-frontier ceiling on MCQ formats; only PHYBench, OIBench, and PutnamBench publish a contamination protocol, and none publish the three\-stage \(n\-gram, embedding, LLM\-judge\) pairwise audit we introduce in §[3\.3](https://arxiv.org/html/2605.14040#S3.SS3)\. Table[1](https://arxiv.org/html/2605.14040#S2.T1)maps our released audited corpus andPhysOlym\-Aagainst related benchmarks on seven axes\.

Table 1:Released artifacts vs related benchmarks across eight axes\.*Audit:*2\-stage \(n\-gram\+embedding\) / 1\-stage / orig\. \(constructed\-novel\) / none\.*T/T leak:*train→\\totest joint\-stage \(J≥0\.4∨cos≥0\.85J\\geq 0\.4\\vee\\cos\\geq 0\.85\) audit against six public physics evals;✓\\checkmarkall 6 = clean\.*Diff:*organizer difficulty\.*X\-L:*paired cross\-lingual\.*Use:*E/T = eval/train\.*RL\-ready:*closed\-form gold \+ audit\-clean \+ RL recipe\. “⋅\\cdot” = eval\-only; “n/r” = train pool, no cross\-corpus audit\. Only this work reports train/test contamination: after re\-audit cleanup,PhysCorp\-A\(6,432\) andPhysR1Corp\(2,268\) are clean againstall sixevals \(Table[2](https://arxiv.org/html/2605.14040#S3.T2)\)\.BenchmarkSizeFormatMMAuditT/T leakDiffX\-LUseRL\-ready*Physics\-domain benchmarks*PHYBench\(Qiuet al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib36)\)500open\+EED–orig\.⋅\\cdot––E–PhysUniBench\(Wanget al\.,[2025b](https://arxiv.org/html/2605.14040#bib.bib37)\)3,304open MM✓\\checkmark1\-stage⋅\\cdot✓\\checkmark–E–UGPhysics\(Xuet al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib3)\)5,520open text–1\-stagen/r–EN/ZHT–PhysReason\(Zhanget al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib4)\)1,200step open MM✓\\checkmarknone⋅\\cdot––E–OlympiadBench\(Heet al\.,[2024](https://arxiv.org/html/2605.14040#bib.bib2)\)8,952open MM✓\\checkmarknone⋅\\cdot–EN/ZHE–*Olympiad / formal / contamination\-by\-design*PutnamBench\(Tsoukalaset al\.,[2024](https://arxiv.org/html/2605.14040#bib.bib6)\)1,692Lean/Isab\.–orig\.⋅\\cdot✓\\checkmark–E–OIBench\(Zhuet al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib38)\)250open code–2\-stage⋅\\cdot✓\\checkmarkEN/ZHE–FrontierMath\(Glazeret al\.,[2024](https://arxiv.org/html/2605.14040#bib.bib7)\)290open math–orig\.⋅\\cdot✓\\checkmark–E–HLE\(Phanet al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib39)\)2,500expert exam✓\\checkmarkorig\.⋅\\cdot––E–*Multimodal / multi\-domain*MMLU\-Pro\(Wanget al\.,[2024](https://arxiv.org/html/2605.14040#bib.bib40)\)12,03210\-MCQ–none⋅\\cdot––E–MMMU\-Pro Phys\(Yueet al\.,[2024b](https://arxiv.org/html/2605.14040#bib.bib41)\)6010\-MCQ MM✓\\checkmarknone⋅\\cdot––E–SciInstruct\(Zhanget al\.,[2024](https://arxiv.org/html/2605.14040#bib.bib42)\)254,KSFT instr\.–1\-stagen/r––T–*This work \(train→\\totest cross\-corpus audit reported; Table[2](https://arxiv.org/html/2605.14040#S3.T2)\)*PhysCorp\-A\(ours\)6,432open\+MCQ MM✓\\checkmark2\-stage✓\\checkmarkall 6✓\\checkmark✓\\checkmarkT✓\\checkmarkPhysR1Corp\(ours\)2,268MCQ \+ num MM✓\\checkmark2\-stage✓\\checkmarkall 6✓\\checkmark✓\\checkmarkT✓\\checkmarkPhysOlym\-A\(ours\)500open MM novel✓\\checkmark2\-stage⋅\\cdot\(eval; clean\)✓\\checkmarkEN/ETE–
✓\\checkmark= present; – = absent or not reported\. “2\-stage” audit = pairwise 5\-gram\-Jaccard*and*embedding\-cosine against external corpora and held\-out splits\.

#### Contamination audits and other prior work\.

PutnamBench\(Tsoukalaset al\.,[2024](https://arxiv.org/html/2605.14040#bib.bib6)\), FrontierMath\(Glazeret al\.,[2024](https://arxiv.org/html/2605.14040#bib.bib7)\), HLE\(Phanet al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib39)\), and EnigmaEval\(Wanget al\.,[2025a](https://arxiv.org/html/2605.14040#bib.bib51)\)provide release\-policy templates and dismissal grounds; methodological work spans n\-gram audits\(Sainzet al\.,[2023](https://arxiv.org/html/2605.14040#bib.bib8)\), the rephrased\-samples failure mode\(Yanget al\.,[2023](https://arxiv.org/html/2605.14040#bib.bib43)\)\(which our Stage 2 catches\), embedding\-based detection\(Singhet al\.,[2024](https://arxiv.org/html/2605.14040#bib.bib9)\), and performance\-based detection\(Dekonincket al\.,[2024](https://arxiv.org/html/2605.14040#bib.bib49)\); the survey ofRavautet al\.\([2024](https://arxiv.org/html/2605.14040#bib.bib48)\)consolidates these\. We import the math template, adding the embedding\-cosine pass because physics statements \(units, vectors, figure references\) are more paraphrase\-sensitive than typical math problems—a sensitivity Table[4](https://arxiv.org/html/2605.14040#A1.T4)quantifies\. PhysBench\(Chowet al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib47)\)evaluates intuitive\-physics dynamics from video, orthogonal scope\. Multilingual benchmarks have proliferated\(Xuanet al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib46); Ahujaet al\.,[2024](https://arxiv.org/html/2605.14040#bib.bib44); Wuet al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib50)\); our cross\-lingual finding \(§[5\.1](https://arxiv.org/html/2605.14040#S5.SS1.SSS0.Px3)\) differs methodologically by evaluating identical 59 problems in original Estonian and English translation on the same closed model with paired tests, isolating a within\-problem effect aggregate benchmarks cannot\.

## 3Data: The Audited Corpus and Held\-Out Olympiad Eval

Released artifacts:PhysCorp\-A\(6,432\-record audited corpus, including 1,609 first\-ML\-format olympiad problems—Estonian PhO with native 1–10 difficulty \+ 201 EN/ET bilingual, Kevin Zhou’s handouts, 7 international olympiads\);PhysR1Corp\(2,268\-record closed\-form RL pool, MCQ and numerical only\); the held\-outPhysOlym\-Aeval \(§[3\.2](https://arxiv.org/html/2605.14040#S3.SS2)\); the Physics\-R1 recipe \(Algorithms[2](https://arxiv.org/html/2605.14040#alg2),[3](https://arxiv.org/html/2605.14040#alg3)\); and the audit pipeline \(Algorithm[1](https://arxiv.org/html/2605.14040#alg1), Table[4](https://arxiv.org/html/2605.14040#A1.T4)\)\. All ship under per\-source licenses \(Table[5](https://arxiv.org/html/2605.14040#A5.T5)\) on HuggingFace\+GitHub\+Zenodo with Croissant 1\.0 metadata\.

### 3\.1Training Corpus Composition

The corpus is drawn from nine source families \(Table[9](https://arxiv.org/html/2605.14040#A8.T9)\)\. Five are repackaged from existing benchmarks under documented licenses \(UGPhysics\(Xuet al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib3)\), OpenStax College and University Physics\(OpenStax,[2024](https://arxiv.org/html/2605.14040#bib.bib30)\), Physics Stack Exchange\(Stack Exchange Inc\.,[2024](https://arxiv.org/html/2605.14040#bib.bib31)\), an MMMU\+o1\-CoT seed\(Yueet al\.,[2024a](https://arxiv.org/html/2605.14040#bib.bib5)\), PhysReason\(Zhanget al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib4)\)\); four contribute first\-ML\-format material: the Estonian Physics Olympiad collection\(Estonian Physics Olympiad,[2018](https://arxiv.org/html/2605.14040#bib.bib22)\)\(418 problems, 2004–2018, with organizer\-issued 1–10 difficulty labels and a 201\-problem bilingual EN\+ET subset\), Kevin Zhou’s olympiad handouts\(Zhou,[2018](https://arxiv.org/html/2605.14040#bib.bib23)\)\(692 problems, with native point values 1–5 and a 3\.2% advanced flag; some problems are drawn from books or other olympiad archives with inline attribution preserved per record, see Appendix[E](https://arxiv.org/html/2605.14040#A5)\), and refreshed scrapes of seven international olympiads \(IPhO\(International Physics Olympiad,[2025](https://arxiv.org/html/2605.14040#bib.bib24)\), NBPhO\(NBPhO Committee,[2025](https://arxiv.org/html/2605.14040#bib.bib28)\), EuPhO\(EuPhO Committee,[2025](https://arxiv.org/html/2605.14040#bib.bib29)\), APhO\(Asian Physics Olympiad Committee,[2025](https://arxiv.org/html/2605.14040#bib.bib26)\), USAPhO\(American Association of Physics Teachers,[2025](https://arxiv.org/html/2605.14040#bib.bib25)\), INPhO\(Homi Bhabha Centre for Science Education,[2025](https://arxiv.org/html/2605.14040#bib.bib27)\), IYPT\)\. Source families ship under a mix of CC BY 4\.0, CC BY\-SA 4\.0, public\-domain by competition policy \(Estonian PhO, IPhO, NBPhO, EuPhO, APhO, USAPhO, INPhO\), CC BY\-NC 4\.0 \(Kevin Zhou’s handouts; written grant 2026\-05\-03\), and CC BY\-NC\-SA 4\.0 \(UGPhysics\); per\-source licenses are listed in Appendix[E](https://arxiv.org/html/2605.14040#A5)\(Table[5](https://arxiv.org/html/2605.14040#A5.T5)\) and carried through to each released record\. The full14,29414\{,\}294\-record pre\-audit pool is released asPhysCorp\-pre\-auditso that downstream users can reproduce the audit;PhysCorp\-Ais the6,4326\{,\}432\-record subset that survives all three stages plus a re\-audit against PhysReason\-full and PhysUniBench\-en \(804 records dropped, dominated by PhysReason\-full540540and PhysUniBench\-en186186\)\. The released pool is disjoint from PhyX, MMMU\-Pro Physics, OlympiadBench\-Physics, UGPhysics\-Train, PhysReason\-full, PhysUniBench\-en, andPhysOlym\-Aat the joint operating thresholds\. The candidate\-to\-release cleanup forPhysR1Corpis detailed in §[3\.3](https://arxiv.org/html/2605.14040#S3.SS3)\.

#### LLM\-touched\-statement subset disclosure\.

Of the2,2682\{,\}268records inPhysR1Corp, approximately7373\(3\.2%3\.2\\%\) have LLM\-touched problem statements:∼11\\sim 11are derived from a8585\-record Claude\-generated synthetic\-MCQ augmentation pool \(3 verbatim, 8 numeric paraphrases\), and∼62\\sim 62are numeric\-variation paraphrases of realPhysCorp\-Arecords \(e\.g\., variant problem constants\)\. The remaining∼2,195\\sim 2\{,\}195records have unmodified problem statements from the nine source families\. LLM augmentation is documented per\-distribution in the Croissant metadata’ssyntheticDataDescriptionfield; the held\-outPhysOlym\-Aeval contains no synthetic problem content\.

### 3\.2PhysOlym\-A: Held\-Out Olympiad Eval

Standard physics\-VL benchmarks no longer resolve frontier\-class differences: PhyX clusters top entries within ten points of the80%80\\%ceiling; OlympiadBench\-Physics predates the contamination\-audit discipline; UGPhysics is itself a candidate for audited training data, not held\-out evaluation\. Physics\-R1’s stopping rule and reward\-component ablation depend on a held\-out signal that is non\-saturating and contamination\-clean against the training pool\.

PhysOlym\-A\(Physics Olympiad, Audited\) is composed of 200 problems from Kevin Zhou’s olympiad handouts, 136 from the Estonian PhO collection, 85 from an IPhO/NBPhO/EuPhO scrape, and 79 from an APhO/USAPhO/INPhO scrape \(500 total,𝟒𝟗𝟗\\mathbf\{499\}novel\-source under our four\-corpus audit\)\. Native difficulty signals:27%27\\%of records carry Estonian organizer\-issued 1–10 difficulty;38%38\\%carry Zhou’s pedagogical 1–5 point values;2%2\\%carry Zhou’s advanced \[A\] flag\. The three\-stage audit \(§[3\.3](https://arxiv.org/html/2605.14040#S3.SS3)\) certifies𝟎\\mathbf\{0\}Stage\-3 near\-duplicate overlaps between the audited training pool andPhysOlym\-A, and𝟎\\mathbf\{0\}overlaps between the novel pool and PhyX 1000q\. The single non\-novel record is an EuPhO 2020 problem also present in OlympiadBench\-Physics atJ=0\.91J\{=\}0\.91; we disclose this in Appendix[A](https://arxiv.org/html/2605.14040#A1)rather than silently drop it\. The scoring protocol \(LLM\-judge with strict/liberal accuracy,κ\\kappainter\-judge agreement, and the auxiliary held\-out splits used during training\) is described in §[5](https://arxiv.org/html/2605.14040#S5)\.

### 3\.3The Three\-Stage Audit Pipeline

Table 2:Train/test contamination across released physics\-VL training pools, three\-stage audit\.Rows: public physics eval splits; columns: training pools \(three competitor, two cleaned ours\)\. Cells:Stage\-1 / Stage\-2 raw / Stage\-3 near\-duppair counts \(Algorithm[1](https://arxiv.org/html/2605.14040#alg1)\)\. Stage\-1 = 5\-gram Jaccard≥0\.4\\geq 0\.4, Stage\-2 =mxbai\-embed\-large\-v1cosine≥0\.85\\geq 0\.85\(high recall over close\-content pairs\), Stage\-3 = Haiku\-4\.5 LLM judge separating each Stage\-2 candidate into*close duplicate*\(paraphrase / numeric variation of the same problem\) vs\.*same\-topic neighbor*\(related physics, distinct setup\)\. Competitor pools \(UGPhysics\-Train: 200\-record annotated subset; SciInstruct: en\_phy\_chem42,35242\{,\}352\-record subset of 254 K; MMK12:15,60815\{,\}608\-record MM\-Eureka train pool\) report𝟎/𝟔\\mathbf\{0/6\}Stage\-1 hits; Stage\-2 surfaces𝟒,𝟖𝟒𝟔\\mathbf\{4\{,\}846\}close\-content pairs in SciInstruct,99in UGPhysics\-Train, and6666in MMK12\.Stage\-3 LLM\-judge separates close duplicates from same\-topic neighbors:SciInstruct4,846→𝟏𝟑𝟒4\{,\}846\{\\to\}\\mathbf\{134\}near\-duplicates \(PhysReason\-full2,687→362\{,\}687\{\\to\}36, PhysUniBench\-en1,027→221\{,\}027\{\\to\}22, PhyX\-mini703→46703\{\\to\}46dominant\); UGPhysics\-Train9→𝟎9\{\\to\}\\mathbf\{0\}; MMK1266→𝟎66\{\\to\}\\mathbf\{0\}\. The close\-duplicate share is sharply cosine\-driven:100%100\\%atcos≥0\.95\\cos\\geq 0\.95vs\.1\.5%1\.5\\%atcos∈\[0\.85,0\.87\)\\cos\\in\[0\.85,0\.87\)\(Appendix[A](https://arxiv.org/html/2605.14040#A1)\)\.Both released pools are fully Stage\-3 clean against all six evals:PhysCorp\-A\(6,4326\{,\}432\) after dropping804804from a7,2367\{,\}236candidate, with𝟎/𝟎\\mathbf\{0/0\}Stage\-3 close\-duplicates against all six evals \(no S2 candidates surviving the joint S1∨\\veeS2 cleanup\);PhysR1Corp\(2,2682\{,\}268\) after dropping8787MMMU\-Pro \+7878PhyX\-mini/PhysUniBench\-en hits from a2,4332\{,\}433candidate, with all1919remaining S2 candidates classified as same\-topic neighbors by Stage\-3 \(𝟎/𝟏𝟗\\mathbf\{0/19\}near\-duplicates\), agreeing100%100\\%with manual inspection\.Eval↓\\downarrow/ Train pool→\\to*Other published pools \(we re\-audit\)**This work \(cleaned\)*UGPhysics\-Train\(200 sub\)SciInstruct\(en\_phy\_chem; 42 K\)MMK12\(MM\-Eureka; 15 K\)PhysCorp\-A\(6,432\)PhysR1Corp\(2,268\)PhysOlym\-A\(500\)0/0/00\\,/\\,0\\,/\\,\\mathbf\{0\}0/163/80\\,/\\,163\\,/\\,\\mathbf\{8\}0/0/00\\,/\\,0\\,/\\,\\mathbf\{0\}0/0/00\\,/\\,0\\,/\\,\\mathbf\{0\}0/3/00\\,/\\,3\\,/\\,\\mathbf\{0\}PhyX\-mini \(1,000\)0/1/00\\,/\\,1\\,/\\,\\mathbf\{0\}0/703/460\\,/\\,703\\,/\\,\\mathbf\{46\}0/0/00\\,/\\,0\\,/\\,\\mathbf\{0\}0/0/00\\,/\\,0\\,/\\,\\mathbf\{0\}0/0/00\\,/\\,0\\,/\\,\\mathbf\{0\}MMMU\-Pro Phys \(60\)0/0/00\\,/\\,0\\,/\\,\\mathbf\{0\}0/141/70\\,/\\,141\\,/\\,\\mathbf\{7\}0/0/00\\,/\\,0\\,/\\,\\mathbf\{0\}0/0/00\\,/\\,0\\,/\\,\\mathbf\{0\}0/1/00\\,/\\,1\\,/\\,\\mathbf\{0\}OlymBench\-Phys \(692\)0/2/00\\,/\\,2\\,/\\,\\mathbf\{0\}0/130/150\\,/\\,130\\,/\\,\\mathbf\{15\}0/0/00\\,/\\,0\\,/\\,\\mathbf\{0\}0/0/00\\,/\\,0\\,/\\,\\mathbf\{0\}0/4/00\\,/\\,4\\,/\\,\\mathbf\{0\}PhysReason\-full \(1,200\)0/4/00\\,/\\,4\\,/\\,\\mathbf\{0\}0/2,687/360\\,/\\,2\{,\}687\\,/\\,\\mathbf\{36\}0/62/00\\,/\\,62\\,/\\,\\mathbf\{0\}0/0/00\\,/\\,0\\,/\\,\\mathbf\{0\}0/11/00\\,/\\,11\\,/\\,\\mathbf\{0\}PhysUniBench\-en \(1,022\)0/2/00\\,/\\,2\\,/\\,\\mathbf\{0\}0/1,027/220\\,/\\,1\{,\}027\\,/\\,\\mathbf\{22\}0/4/00\\,/\\,4\\,/\\,\\mathbf\{0\}0/0/00\\,/\\,0\\,/\\,\\mathbf\{0\}0/0/00\\,/\\,0\\,/\\,\\mathbf\{0\}*Total S2 / S3 real*99/𝟎\\mathbf\{0\}𝟒,𝟖𝟒𝟔\\mathbf\{4\{,\}846\}/𝟏𝟑𝟒\\mathbf\{134\}6666/𝟎\\mathbf\{0\}0/𝟎\\mathbf\{0\}𝟏𝟗\\mathbf\{19\}/𝟎\\mathbf\{0\}
Format:Stage\-1 / Stage\-2 / Stage\-3pair counts \(Stage\-3 = Haiku\-4\.5 LLM\-judge classifying each Stage\-2 candidate as close duplicate vs\. same\-topic neighbor\)\. SciInstruct’s S3\-near\-dup cells reveal the close\-duplicate share is threshold\-driven:17/1717/17atcos≥0\.95\\cos\{\\geq\}0\.95,54/1,15954/1\{,\}159at\[0\.87,0\.90\)\[0\.87,0\.90\),53/3,54353/3\{,\}543at\[0\.85,0\.87\)\[0\.85,0\.87\)\(Appendix[A](https://arxiv.org/html/2605.14040#A1), Table[A](https://arxiv.org/html/2605.14040#A1.SS0.SSS0.Px1)\)\.

The pipeline constructs both the audited training pool and the held\-outPhysOlym\-Aeval under the same definition of contamination, applied pairwise across the training pool, four external corpora \(PhyX, MMMU\-Pro Physics, OlympiadBench\-Physics, UGPhysics\-Train\), and the held\-out splits\.*Stage 1 \(n\-gram\)\.*Tokenize each problem statement with a unicode word tokenizer, build the 5\-gram shingle set, and flag pairs with Jaccard≥0\.4\\geq 0\.4\.*Stage 2 \(embedding\)\.*Encode each statement withmxbai\-embed\-large\(1024\-dim,L2L\_\{2\}\-normalized\) and flag pairs with cosine≥0\.85\\geq 0\.85\. Stage\-2 has high recall on close\-content pairs, including the rephrasing\-class duplicates Stage\-1 misses, but its single\-threshold operating point also flags same\-topic\-but\-distinct\-problem pairs\.*Stage 3 \(LLM\-judge precision filter\)\.*For each Stage\-2 candidate, a Haiku\-4\.5 judge receives both problem statements and classifies the pair as a*close duplicate*\(paraphrase or numeric variation of the same problem\) or a*same\-topic neighbor*\(related physics, distinct setup\)\. Only Stage\-3 close\-duplicate records are removed from the training pool\. Pseudocode is in Algorithm[1](https://arxiv.org/html/2605.14040#alg1); worked examples in Appendix[H\.5](https://arxiv.org/html/2605.14040#A8.SS5); calibration of the embedder \+ thresholds in Appendix[A](https://arxiv.org/html/2605.14040#A1)\.

On the train/test contamination matrix of Table[2](https://arxiv.org/html/2605.14040#S3.T2), the cosine\-bucketed precision pattern \(100%100\\%close\-duplicates atcos≥0\.95\\cos\\geq 0\.95vs\.1\.5%1\.5\\%atcos∈\[0\.85,0\.87\)\\cos\\in\[0\.85,0\.87\); Appendix[A](https://arxiv.org/html/2605.14040#A1), Table[A](https://arxiv.org/html/2605.14040#A1.SS0.SSS0.Px1)\) confirms the protocol’s design hypothesis: embedding cosine alone is recall\-dominant and an LLM judge is the appropriate precision filter\.Both released training pools are fully Stage\-3 clean against all six public evals \(Table[2](https://arxiv.org/html/2605.14040#S3.T2)\):PhysCorp\-A\(6,4326\{,\}432\), built via Stage\-1∨\\veeStage\-2 audit dropping804804of a7,2367\{,\}236candidate, surfaces0Stage\-2 candidates and hence𝟎/𝟎\\mathbf\{0/0\}Stage\-3 close\-duplicates by construction;PhysR1Corp\(2,2682\{,\}268\), additionally dropping8787MMMU\-Pro and7878PhyX\-mini/PhysUniBench near\-duplicates from a2,4332\{,\}433\-record candidate \(Appendix[A\.1](https://arxiv.org/html/2605.14040#A1.SS1)\), retains1919Stage\-2 candidates classified as same\-topic neighbors by Stage\-3 with100%100\\%manual\-inspection agreement \(𝟎/𝟏𝟗\\mathbf\{0/19\}close\-duplicates\)\.

Algorithm 1Three\-stage contamination audit\.1:Train pool

TT, external corpora

\{Ek\}k=1K\\\{E\_\{k\}\\\}\_\{k=1\}^\{K\}, held\-out splits

\{Hj\}j=1J\\\{H\_\{j\}\\\}\_\{j=1\}^\{J\}, normalize fn

norm​\(⋅\)\\mathrm\{norm\}\(\\cdot\), embedder

enc​\(⋅\)\\mathrm\{enc\}\(\\cdot\), LLM judge

Judge​\(⋅,⋅\)∈\{close\-dup,topic\-neighbor\}\\textsc\{Judge\}\(\\cdot,\\cdot\)\\in\\\{\\text\{close\-dup\},\\text\{topic\-neighbor\}\\\}, thresholds

τJ=0\.4\\tau\_\{J\}\{=\}0\.4,

τC=0\.85\\tau\_\{C\}\{=\}0\.85
2:Audited pool

T′T^\{\\prime\}disjoint from

⋃kEk∪⋃jHj\\bigcup\_\{k\}E\_\{k\}\\cup\\bigcup\_\{j\}H\_\{j\}at the joint thresholds\.

3:*Stage 1: 5\-gram Jaccard \(n\-gram audit\)\.*

4:

St←\{5\-gram shingle set of​norm​\(t\)\}S\_\{t\}\\leftarrow\\\{\\text\{5\-gram shingle set of \}\\mathrm\{norm\}\(t\)\\\}for each

t∈T∪⋃Ek∪⋃Hjt\\in T\\cup\\bigcup E\_\{k\}\\cup\\bigcup H\_\{j\}
5:for

t∈Tt\\in Tdo

6:

Jmax​\(t\)←maxx∈⋃Ek∪⋃Hj⁡\|St∩Sx\|/\|St∪Sx\|J\_\{\\max\}\(t\)\\leftarrow\\max\_\{x\\in\\bigcup E\_\{k\}\\cup\\bigcup H\_\{j\}\}\\,\|S\_\{t\}\\cap S\_\{x\}\|/\|S\_\{t\}\\cup S\_\{x\}\|
7:endfor

8:*Stage 2:mxbai\-embed\-largecosine \(paraphrase recall\)\.*

9:

𝐞t←enc​\(norm​\(t\)\)/‖enc​\(norm​\(t\)\)‖\\mathbf\{e\}\_\{t\}\\leftarrow\\mathrm\{enc\}\(\\mathrm\{norm\}\(t\)\)/\\\|\\mathrm\{enc\}\(\\mathrm\{norm\}\(t\)\)\\\|for each

tt
10:for

t∈Tt\\in Tdo

11:

Cmax​\(t\)←maxx∈⋃Ek∪⋃Hj⁡𝐞t⊤​𝐞xC\_\{\\max\}\(t\)\\leftarrow\\max\_\{x\\in\\bigcup E\_\{k\}\\cup\\bigcup H\_\{j\}\}\\mathbf\{e\}\_\{t\}^\{\\top\}\\mathbf\{e\}\_\{x\}
12:endfor

13:

C​\(T\)←\{t:Jmax​\(t\)≥τJ​or​Cmax​\(t\)≥τC\}C\(T\)\\leftarrow\\\{t:J\_\{\\max\}\(t\)\\geq\\tau\_\{J\}\\,\\textsc\{or\}\\,C\_\{\\max\}\(t\)\\geq\\tau\_\{C\}\\\}⊳\\trianglerightcandidate set, high\-recall union

14:*Stage 3: Haiku\-4\.5 LLM\-judge \(precision filter\)\.*For each

t∈C​\(T\)t\\in C\(T\)with top\-matching

x∗​\(t\)←arg⁡maxx⁡𝐞t⊤​𝐞xx^\{\*\}\(t\)\\leftarrow\\arg\\max\_\{x\}\\mathbf\{e\}\_\{t\}^\{\\top\}\\mathbf\{e\}\_\{x\}, queryJudgeto classify the pair as a close duplicate \(paraphrase / numeric variation of the same problem\) or a same\-topic neighbor \(related physics, distinct setup\)\.

15:

R​\(T\)←\{t∈C​\(T\):Judge​\(t,x∗​\(t\)\)=close\-dup\}R\(T\)\\leftarrow\\\{t\\in C\(T\):\\textsc\{Judge\}\(t,x^\{\*\}\(t\)\)=\\text\{close\-dup\}\\\}
16:

T′←T∖R​\(T\)T^\{\\prime\}\\leftarrow T\\setminus R\(T\)
17:return

T′T^\{\\prime\}and per\-stage counts

\|Jmax≥τJ\|\|J\_\{\\max\}\\geq\\tau\_\{J\}\|,

\|Cmax≥τC\|\|C\_\{\\max\}\\geq\\tau\_\{C\}\|,

\|R​\(T\)\|\|R\(T\)\|\(Tables[2](https://arxiv.org/html/2605.14040#S3.T2),[4](https://arxiv.org/html/2605.14040#A1.T4)\)\.

#### Threshold\-sensitive leakage on a researcher\-curated baseline \(Finding 1\)\.

On a 1,679\-record sample drawn fromPhysCorp\-pre\-auditunder conventional 5\-gram\-Jaccard \+ within\-pool embedding dedup, audited against a 500\-record internal analysis eval \(distinct fromPhysOlym\-A, constructed post\-audit\), the joint Stage\-1∨\\veeStage\-2 audit raises the detected leak rate from3\.3%3\.3\\%\(Stage\-1 alone, all exact matches atJ=1\.0J\{=\}1\.0\) to8\.8%\\mathbf\{8\.8\\%\}, sweeping4\.74\.7–27\.1%27\.1\\%as the cosine threshold moves between0\.900\.90and0\.800\.80\(Appendix[A\.1](https://arxiv.org/html/2605.14040#A1.SS1), Table[4](https://arxiv.org/html/2605.14040#A1.T4)\)\. The5\.55\.5\-pp gap is the rephrasing dark\-matter that justifies the audited release as a measurement intervention\.

## 4Physics\-R1: A Multi\-Model RL Recipe

Physics\-R1 is reported as evidence the audited corpus has training utility under standard rule\-based RL, not as an algorithmic contribution\. The optimizer is GSPO\(Zhenget al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib12)\)\+DAPO\(Yuet al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib11)\), unmodified\. For each promptxx, sampleK=16K\{=\}16rollouts\{yk\}∼πθold\(⋅∣x\)\\\{y\_\{k\}\\\}\\sim\\pi\_\{\\theta\_\{\\mathrm\{old\}\}\}\(\\cdot\\mid x\), score with rewardr​\(yk,x\)r\(y\_\{k\},x\), form group\-normalized advantages and the clipped sequence\-level GSPO objective

Ak\\displaystyle A\_\{k\}=r​\(yk,x\)−r¯σr\+ε,wk​\(θ\)=\(πθ​\(yk∣x\)πθold​\(yk∣x\)\)1/\|yk\|,\\displaystyle=\\frac\{r\(y\_\{k\},x\)\-\\bar\{r\}\}\{\\sigma\_\{r\}\+\\varepsilon\},\\qquad w\_\{k\}\(\\theta\)=\\left\(\\frac\{\\pi\_\{\\theta\}\(y\_\{k\}\\mid x\)\}\{\\pi\_\{\\theta\_\{\\mathrm\{old\}\}\}\(y\_\{k\}\\mid x\)\}\\right\)^\{\\\!1/\|y\_\{k\}\|\},\(1\)ℒGSPO\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{GSPO\}\}=−𝔼​\[1K​∑kmin⁡\(wk​Ak,clip​\(wk,1±ϵ\)​Ak\)\]\+βKL​DKL​\(πθ∥πbase\),\\displaystyle=\-\\mathbb\{E\}\\\!\\left\[\\tfrac\{1\}\{K\}\\\!\\sum\_\{k\}\\\!\\min\\\!\\big\(w\_\{k\}A\_\{k\},\\ \\mathrm\{clip\}\(w\_\{k\},\\,1\{\\pm\}\\epsilon\)A\_\{k\}\\big\)\\right\]\+\\beta\_\{\\mathrm\{KL\}\}D\_\{\\mathrm\{KL\}\}\(\\pi\_\{\\theta\}\\\|\\pi\_\{\\mathrm\{base\}\}\),with\(r¯,σr\)\(\\bar\{r\},\\sigma\_\{r\}\)the group mean/std,\(ϵlo,ϵhi\)=\(0\.20,0\.28\)\(\\epsilon\_\{\\mathrm\{lo\}\},\\epsilon\_\{\\mathrm\{hi\}\}\)\{=\}\(0\.20,0\.28\),βKL=10−3\\beta\_\{\\mathrm\{KL\}\}\{=\}10^\{\-3\},πbase=\\pi\_\{\\mathrm\{base\}\}\{=\}Qwen3\-VL\-8B\-Thinking BASE\. Cold\-start from base, KL anchor, MM\-Eureka\(Menget al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib33)\)difficulty curriculum \(drop0/N0/NandN/NN/Nprompts,∼22%\{\\sim\}22\\%filtered\),12,28812\{,\}288\-token CoT budget, and held\-out PhyX\-mini\-MC early stopping fix the joint setting \(Algorithm[3](https://arxiv.org/html/2605.14040#alg3), Table[10](https://arxiv.org/html/2605.14040#A8.T10)\); implementation usesverl0\.6\.1\(Shenget al\.,[2024](https://arxiv.org/html/2605.14040#bib.bib16)\)on Qwen3\-VL\-8B\-Thinking\(Qwen Team,[2025](https://arxiv.org/html/2605.14040#bib.bib17)\)with FSDP1 sharding \(§[6](https://arxiv.org/html/2605.14040#S6)\)\.

#### Two reward shapes: binary \(recommended\) vs\. dense \(ablation\)\.

Physics rollouts admit physics\-native per\-step signals—units, conservation, symbolic form—so a denser reward looks free\. We compare:

\(binary, recommended\)rbin​\(y,x\)=𝟙​\[Match​\(ExtractBoxed​\(y\),g​\(x\)\)\]∈\{0,1\},\\displaystyle r\_\{\\mathrm\{bin\}\}\(y,x\)=\\mathbb\{1\}\\\!\\big\[\\textsc\{Match\}\(\\textsc\{ExtractBoxed\}\(y\),\\,g\(x\)\)\\big\]\\in\\\{0,1\\\},\(2\)\(dense, ablation\)rdense=clip​\(rans\+rfmt\+rdim\+rsym\+rcons,−1,1\)\.\\displaystyle r\_\{\\mathrm\{dense\}\}=\\mathrm\{clip\}\\\!\\big\(r\_\{\\mathrm\{ans\}\}\+r\_\{\\mathrm\{fmt\}\}\+r\_\{\\mathrm\{dim\}\}\+r\_\{\\mathrm\{sym\}\}\+r\_\{\\mathrm\{cons\}\},\\,\-1,1\\big\)\.whereMatchaccepts MCQ\-letter equality,±1%\\pm 1\\%numeric tolerance, or symbolic equivalence \(Appendix[C\.1](https://arxiv.org/html/2605.14040#A3.SS1)\); the dense components arerans≡rbinr\_\{\\mathrm\{ans\}\}\{\\equiv\}r\_\{\\mathrm\{bin\}\},rfmt∈\{0,\+0\.1\}r\_\{\\mathrm\{fmt\}\}\{\\in\}\\\{0,\{\+\}0\.1\\\}\(\\boxed\{\}present\),rdim∈\{0,\+0\.15\}r\_\{\\mathrm\{dim\}\}\{\\in\}\\\{0,\{\+\}0\.15\\\}\(sympy\.physics\.units\),rsym∈\{0,\+0\.20\}r\_\{\\mathrm\{sym\}\}\{\\in\}\\\{0,\{\+\}0\.20\\\}\(\\fracsympifies\),rcons∈\{−0\.25,0\}r\_\{\\mathrm\{cons\}\}\{\\in\}\\\{\-0\.25,0\\\}\(energy/momentum violation; Appendix[C\.2](https://arxiv.org/html/2605.14040#A3.SS2)\)\. Under GSPO with group\-normalized advantages and a difficulty curriculum, four properties land binary as variance\-optimal and Goodhart\-robust \(full derivation in Appendix[C\.1](https://arxiv.org/html/2605.14040#A3.SS1.SSS0.Px1)\)\.\(P1\) Group normalization absorbs reward magnitude:AkA\_\{k\}is invariant to affine rescaling ofrrwithin a group, so dense only matters when it*reorders*rollouts—we measure14\.3%14\.3\\%of within\-group pairs flipped,87%87\\%inside the all\-wrong subgroup\.\(P2\) Wrong\-group reorderings are a Goodhart channel:rewarding well\-formatted\-but\-wrong above poorly\-formatted\-but\-wrong biases the policy toward LaTeX\-format proxies that transfer poorly to the audited held\-out eval\.\(P3\) Variance\-optimal advantage:on a Bernoulli reward,Var​\(Abin\)=1\\mathrm\{Var\}\(A^\{\\mathrm\{bin\}\}\)\{=\}1saturates theKK\-sample bound; a bounded shaping termδk∈\[0,Δ\]\\delta\_\{k\}\\in\[0,\\Delta\]inflatesσr\\sigma\_\{r\}byO​\(Δ2\)O\(\\Delta^\{2\}\), shrinking\|Acorrectdense\|\|A^\{\\mathrm\{dense\}\}\_\{\\mathrm\{correct\}\}\|below\|Acorrectbin\|\|A^\{\\mathrm\{bin\}\}\_\{\\mathrm\{correct\}\}\|\. Empirical signature at matched step 60 on the seed\-42 ablation \(Table[3](https://arxiv.org/html/2605.14040#S5.T3)\): binary beats dense by\+8\.9\+8\.9/\+4\.9\+4\.9/\+6\.4\+6\.4pp on PhysReason/OlymBench\-Phys\-liberal/PhysOlym\-A\-liberal while tied with dense on PUB\-OE \(−0\.7\-0\.7pp\) and trailing dense by at most0\.60\.6pp on saturated MCQ\. We ship binary as the deployable artifact; the per\-component drop\-out ablation \(Table[11](https://arxiv.org/html/2605.14040#A8.T11)\) is left to follow\-up work\.

## 5Evaluation

We organize this section in two parts\. §[5\.1](https://arxiv.org/html/2605.14040#S5.SS1)characterizesPhysOlym\-Aas a measurement instrument and grounds Findings 2–3 of §[1](https://arxiv.org/html/2605.14040#S1)\(Finding 1, audit\-leakage, is in §[3\.3](https://arxiv.org/html/2605.14040#S3.SS3.SSS0.Px1)\)\. §[5\.2](https://arxiv.org/html/2605.14040#S5.SS2)reports Physics\-R1 results to validatePhysCorp\-Aas trainable\. The Physics\-R1 \(binary, seed 42\) row in Table[3](https://arxiv.org/html/2605.14040#S5.T3)is the headline single\-seed checkpoint; the 3\-seed mean row aggregates seed 42 with two additional seeds \(seed\-17/step\-63 and seed\-23/step\-60\) on the auditedPhysR1Corpcorpus\.

#### Scoring protocol\.

All open\-ended columns of Table[3](https://arxiv.org/html/2605.14040#S5.T3)use*problem\-level*liberal Sonnet\-as\-judge accuracy \(Appendix[D](https://arxiv.org/html/2605.14040#A4)\): for multi\-sub\-part problems on PhysReason and PhysUniBench\-OE,llm\_judge\_v2\_alignment\.pyandllm\_judge\_v3\_pubeo\.pyrespectively call Sonnet 4\.5 once per gold sub\-answer with YES/NO, and the problem is judged correct only if every sub\-part is correct \(AND across sub\-parts\); OlympiadBench\-Physics andPhysOlym\-Ausejudge\_olympiad\.py, which makes a single YES/NO call per problem against the full gold solution\. The unjudgeable rate onPhysOlym\-Ais13\.9%13\.9\\%\(gold solutions consisting of grading rubrics, administrative notes, or figure\-only references\)\. Three layers bound judge optimism: strict vs\. liberal gap on Sonnet \(4\.74\.7pp onPhysOlym\-A\); inter\-judge Cohen’sκ\\kappa\(Cohen,[1960](https://arxiv.org/html/2605.14040#bib.bib20)\)between two Sonnet seeds; and a100100\-problem human\-graded random subset \(Appendix[D](https://arxiv.org/html/2605.14040#A4)\)\. The cross\-vendor judge agreement on a5050\-problem Sonnet/GPT\-4o pair test shows GPT\-4o is*more*lenient than Sonnet \(16%16\\%vs\.8%8\\%positive rate\), bounding self\-grading concern in the opposite direction from naive worry\.

#### Judging concurrency and reproducibility\.

All Sonnet\-judge runs reported in Table[3](https://arxiv.org/html/2605.14040#S5.T3)are executed atworkers=22–44concurrency to stay below Anthropic API rate limits; sub\-judgments that exceed the per\-call timeout are retried at lower concurrency rather than counted as wrong\. Per\-cell judge\-error counts \(typically0–1010out of629629–12001200records, all≤1%\\leq 1\\%\) and per\-record verdicts are released asjudge\_audit\.jsonin the supplementary archive\. The Sonnet 4\.5 PhysReason responses \(Table[3](https://arxiv.org/html/2605.14040#S5.T3),†\) are regenerated withmax\_tokens=16384\\texttt\{max\\\_tokens\}\{=\}16384to match the response budget used by all open\-source baselines and Physics\-R1; intermediate\-length Sonnet responses \(mean<200<200chars under default API settings\) systematically fail to commit a\\boxed\{\}final answer on multi\-sub\-part problems, which v2\_alignment scores as wrong\.

### 5\.1PhysOlym\-Aas a Measurement Instrument

#### Same\-model evaluation reveals a 46\-point format\-and\-novelty gradient\.

On identical Sonnet 4\.5 weights evaluated in the same week, the score sweeps from79\.7%79\.7\\%on PhyX \(4\-way MCQ\) down to50\.4%50\.4\\%liberal on OlympiadBench\-Physics and33\.4%33\.4\\%liberal onPhysOlym\-A—a 46\-point gradient on fixed weights, the strongest evidence the paper has for the central claim that physics evaluation is format\- and novelty\-bound \(Finding 3\)\. Three forces drive the drop: format \(4\-way MCQ vs\. open\-ended\), genre \(PhyX is K\-12 to early\-undergraduate, the bottom two are competition\-grade\), and contamination\-removal \(onlyPhysOlym\-Ais three\-stage audited against the Physics\-R1 training pool\)\. The PhyX→\\toOlympiadBench\-Physics step accounts for∼29\{\\sim\}29pp of the gradient \(dominated by format \+ genre, since both are public and not contamination\-cleaned\), and the OlympiadBench\-Physics→\\toPhysOlym\-Astep adds∼17\{\\sim\}17pp on top \(dominated by audit and novelty since both are open\-ended and competition\-grade\); a controlled 2×\\times2 \(format×\\timesaudit on identical items\) is left to follow\-up work to attribute the residual cleanly\.PhysOlym\-Asits at the bottom of this gradient by construction\. The per\-physics\-category breakdown on OlympiadBench\-Physics \(electromagnetism hardest at38\.4%38\.4\\%, astrophysics easiest at72\.9%72\.9\\%\) and the saturation\-gradient table are in Appendix[H\.7](https://arxiv.org/html/2605.14040#A8.SS7)\.

#### Difficulty\-stratified accuracy from organizer\-issued labels\.

The Estonian Physics Olympiad is the only public physics olympiad whose problems carry organizer\-issued difficulty labels \(1–10\) by construction, eliminating self\-annotation circularity\. On the 131 Estonian problems carrying native annotation, Sonnet 4\.5 strict accuracy decays near\-monotonically from62\.5%62\.5\\%at difficulty 1 to a hard0%0\\%floor at difficulties 3, 6, 8, and 10 \(full table: Appendix[H\.7](https://arxiv.org/html/2605.14040#A8.SS7), Table[12](https://arxiv.org/html/2605.14040#A8.T12)\)\. The trivial end of the Estonian olympiad \(62\.5%62\.5\\%at difficulty 1\) already lies below Sonnet’s PhyX score \(79\.7%79\.7\\%\); the clean zeros at four difficulty bins are the empirical signature of a non\-saturating benchmark, the property required forPhysOlym\-Ato serve as a stopping signal during Physics\-R1 training\.

#### Cross\-lingual ablation: 17\-point translation delta on identical problems\.

The Estonian PhO bilingual subset enables a controlled cross\-lingual experiment on 59 paired problems graded against the same gold by the same Sonnet 4\.5:30\.5%\\mathbf\{30\.5\}\\%strict on Estonian originals vs\.13\.6%\\mathbf\{13\.6\}\\%on English translations \(sign testp=0\.011p\{=\}0\.011; McNemarp=0\.021p\{=\}0\.021; paired bootstrap95%95\\%CI\[\+5\.1,\+28\.9\]\[\+5\.1,\+28\.9\]pp\)\. The per\-problem agreement matrix is asymmetric: 13 correct on Estonian but wrong on English, 3 the reverse\. For weaker open\-source 8B\-class models with limited Estonian training the gap is expected to flip in sign \(pre\-registered as a follow\-up direction; Appendix[H\.4](https://arxiv.org/html/2605.14040#A8.SS4), item \(viii\)\)\.

### 5\.2Physics\-R1

#### Recipe\.

Physics\-R1 is the GSPO\(Zhenget al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib12)\)\+DAPO\(Yuet al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib11)\)recipe of §[4](https://arxiv.org/html/2605.14040#S4)cold\-started from Qwen3\-VL\-8B\-Thinking BASE\(Qwen Team,[2025](https://arxiv.org/html/2605.14040#bib.bib17)\)onPhysR1Corp\(§[3\.1](https://arxiv.org/html/2605.14040#S3.SS1)\) under MM\-Eureka difficulty filtering\(Menget al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib33)\)and a binary correctness reward\.PhyX\-mini\-MC\(1,0001\{,\}000\-problem audit\-clean MCQ subset\(Shenet al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib1)\)\) is held out as the in\-training early\-stop signal: 4\-way MCQ gives a cleaner per\-step trajectory than open\-ended judging, and it is certified disjoint fromPhysR1Corpunder the audit pipeline of §[3\.3](https://arxiv.org/html/2605.14040#S3.SS3)\.

Table 3:Capability across MCQ, numerical, open\-ended, and held\-out olympiad benchmarks\.MCQ random baseline: PhyX\-1k/3k25%25\\%\. All open\-ended columns \(PhysReason, PUB\-OE, OlymBench\-Phys,PhysOlym\-A\) use*problem\-level*liberal Sonnet\-as\-judge accuracy under our v2/v3 judges \(Appendix[D](https://arxiv.org/html/2605.14040#A4)\): every sub\-question of a multi\-part problem must be judged correct for the problem to count\. All Sonnet\-judge runs useworkers=22–44concurrency for rate\-limit safety, with errored sub\-judgments retried at lower concurrency\. Sonnet PhysReason cell \(†\) is generated withmax\_tokens==16384 to match the protocol used by Physics\-R1 and the open\-source baselines\. GPT\-4o PhyX\-1k/3k fromShenet al\.\([2025](https://arxiv.org/html/2605.14040#bib.bib1)\); Gemini PhyX\-1k/3k cells \(∗\) measured here\. All Physics\-R1 evals usemax\_tokens==16384\.MCQOpen\-ended \(problem\-AND aggregation, liberal Sonnet\-judge\)ModelPhyX\-1kPhyX\-3kPhysReasonPUB\-OEOlymBenchPhysOlym\-AMCQ\-exactMCQ\-exactsubpart\-AND \(v2\)subpart\-AND \(v3\)problem\-lvlproblem\-lvl*Closed\-source frontier*Claude Sonnet 4\.579\.780\.649\.1†\\mathbf\{49\.1\}^\{\\dagger\}25\.425\.450\.4\\mathbf\{50\.4\}33\.4\\mathbf\{33\.4\}Gemini 2\.5 Pro75\.1∗49\.8∗38\.833\.437\.412\.2GPT\-4o70\.453\.651\.1¯\\underline\{51\.1\}31\.019\.719\.5*Open\-source bases \(best / second\-best across this block \+ ours:bold/underline\)*Qwen3\-VL\-32B\-Thinking73\.884\.225\.132\.853\.913\.2Qwen3\-VL\-8B\-Thinking \(base\)73\.774\.423\.935\.339\.38\.0InternVL3\-8B46\.843\.113\.323\.510\.44\.0*This work \(subscripts:Δ\\Deltavs\. Qwen3\-VL\-8B\-Thinking base\)*Physics\-R1 \(dense\)78\.3\+4\.6\\mathbf\{78\.3\}\_\{\+4\.6\}77\.5¯\+3\.1\\underline\{77\.5\}\_\{\+3\.1\}23\.3−0\.623\.3\_\{\-0\.6\}37\.7\+2\.4\\mathbf\{37\.7\}\_\{\+2\.4\}40\.5\+1\.240\.5\_\{\+1\.2\}19\.2¯\+11\.2\\underline\{19\.2\}\_\{\+11\.2\}Physics\-R1 \(binary, seed 42\)78\.0¯\+4\.3\\underline\{78\.0\}\_\{\+4\.3\}76\.9\+2\.576\.9\_\{\+2\.5\}32\.2¯\+8\.3\\underline\{32\.2\}\_\{\+8\.3\}37\.0¯\+1\.7\\underline\{37\.0\}\_\{\+1\.7\}45\.4¯\+6\.1\\underline\{45\.4\}\_\{\+6\.1\}25\.6\+17\.625\.6\_\{\+17\.6\}Physics\-R1 \(binary, 3\-seed mean±σ\\pm\\sigma\)‡77\.8±0\.3\+4\.177\.8\_\{\\pm 0\.3\}^\{\+4\.1\}76\.9±0\.3\+2\.576\.9\_\{\\pm 0\.3\}^\{\+2\.5\}39\.6±6\.4\+15\.7\\mathbf\{39\.6\}\_\{\\pm 6\.4\}^\{\+15\.7\}34\.8±3\.3−0\.534\.8\_\{\\pm 3\.3\}^\{\-0\.5\}46\.2±1\.5\+6\.9\\mathbf\{46\.2\}\_\{\\pm 1\.5\}^\{\+6\.9\}26\.3±1\.7\+18\.3\\mathbf\{26\.3\}\_\{\\pm 1\.7\}^\{\+18\.3\}
†Sonnet PhysReason regenerated atmax\_tokens=16384\\texttt\{max\\\_tokens\}=16384,1,192/1,2001\{,\}192/1\{,\}200clean records\.‡3\-seed mean over seeds\{42,17,23\}\\\{42,17,23\\\}on the auditedPhysR1Corp\(2,2682\{,\}268records, binary reward; seed\-42 also the dense\-ablation seed\)\. Per\-seed \(42/17/23\): PR32\.2/43\.1/43\.432\.2/43\.1/43\.4; PUB\-OE37\.0/36\.4/30\.937\.0/36\.4/30\.9; OlymBench\-Phys45\.4/45\.3/48\.045\.4/45\.3/48\.0;PhysOlym\-A25\.6/25\.0/28\.225\.6/25\.0/28\.2; PhyX\-mini78\.0/77\.4/77\.978\.0/77\.4/77\.9; PhyX\-3k76\.9/77\.2/76\.676\.9/77\.2/76\.6\.PhysOlym\-Alift decomposes into∼3\.5\\sim 3\.5pp from\\boxed\{\}\-emission rate \(33\.8%→64\.4%33\.8\\%\{\\to\}64\.4\\%from base to Physics\-R1\) and∼14\.1\\sim 14\.1pp from conditional accuracy \(22\.5→36\.322\.5\{\\to\}36\.3\); details in Appendix[H\.7](https://arxiv.org/html/2605.14040#A8.SS7)\.

#### Capability across formats and the held\-out olympiad column\.

Table[3](https://arxiv.org/html/2605.14040#S5.T3)reports Physics\-R1 alongside closed\-frontier and open\-source bases on three answer formats and the held\-out olympiad split\. The Qwen3\-VL\-8B\-Thinking BASE checkpoint attains73\.7%73\.7\\%on PhyX\-mini\-1k; the 32B sibling is indistinguishable on PhyX \(73\.8%73\.8\\%\), so scale alone in the 8B–32B Thinking band does not move PhyX\. All Physics\-R1 evals usemax\_tokens==16384 to permit full thinking\-mode CoT; eval\-budget sensitivity, harness\-canonical re\-evaluation, and per\-judge gap analysis are in Appendix[H\.7](https://arxiv.org/html/2605.14040#A8.SS7)\.

#### Physics\-R1 \(binary, recommended\): step 60 closes the audited held\-out gap\.

The binary\-reward checkpoint at step 60 lifts the 8B base across all formats \(Table[3](https://arxiv.org/html/2605.14040#S5.T3)\); the largest lift lands onPhysOlym\-Aliberal \(\+18\.3\+18\.3pp at 3\-seed mean\), with lifts on saturated public open\-ended splits substantially smaller—the contamination signal Finding 1 predicts: where the 8B base is already close to ceiling \(PUB\-OE35\.335\.3, OlymBench\-Phys39\.339\.3\), there is little headroom; where the eval is novel\-source and audited \(PhysOlym\-Aliberal8\.08\.0\), the post\-training lift is large\. The numerical/open\-ended jumps from step 40 to step 60 are driven by the additional2020GRPO steps lifting the\\boxed\{\}\-emission rate from46%46\\%to8787–96%96\\%\. The 3\-seed mean \(seed 42 \+ seed\-17/step\-63 \+ seed\-23/step\-60, all on the auditedPhysR1Corp\) is tight across most open\-ended columns: per\-seed PR\{32\.2,43\.1,43\.4\}\\\{32\.2,43\.1,43\.4\\\}\(mean39\.6±6\.439\.6\\pm 6\.4; seed\-42 outlier on PR,∼11\\sim 11pp below seeds 17/23 with otherwise comparable performance on other columns\),PhysOlym\-Aliberal\{25\.6,25\.0,28\.2\}\\\{25\.6,25\.0,28\.2\\\}\(mean26\.3±1\.726\.3\\pm 1\.7\), OlymBench\-Phys\{45\.4,45\.3,48\.0\}\\\{45\.4,45\.3,48\.0\\\}\(mean46\.2±1\.546\.2\\pm 1\.5\), PUB\-OE\{37\.0,36\.4,30\.9\}\\\{37\.0,36\.4,30\.9\\\}\(mean34\.8±3\.334\.8\\pm 3\.3\)\. The audited corpus is trainable and the lift over the 8B base is reproducible across seeds\.

#### Where Physics\-R1 helps: failure modes of the base it mitigates\.

The\+18\.3\+18\.3\-pp 3\-seed\-mean lift onPhysOlym\-Aliberal corresponds to∼92\{\\sim\}92problems flipped from wrong\-on\-base to correct\-on\-Physics\-R1 \(per\-seed range8585–101101\)\. Hand\-inspecting 30 such flips reveals three recurring failure modes of the 8B base, each addressed by a specific recipe lever: \(i\)*reasoning\-without\-committing*\(long correct CoT, no\\boxed\{\}final\) — fixed byrbinr\_\{\\mathrm\{bin\}\}\(§[4](https://arxiv.org/html/2605.14040#S4)\); \(ii\)*unit/dimensional shortcuts*\(dimensionally\-consistent but answer\-wrong\) — fixed by MM\-Eureka curriculum filtering ofN/NN/Nsurface\-heuristic prompts; \(iii\)*multi\-image evidence integration*\(base attends only to the first panel\) — fixed by the cold\-start from Qwen3\-VL\-8B\-Thinking BASE under FSDP1, which preserves the visual encoder\. Physics\-R1 does*not*fix genuine physics\-content gaps \(graduate\-level Tripos\-style perturbation theory remains wrong on both\)\. Full transcripts in Appendix[H\.1](https://arxiv.org/html/2605.14040#A8.SS1)\.

#### PhysOlym\-Agrounds the central training\-utility claim\.

ThePhysOlym\-A\-liberal column of Table[3](https://arxiv.org/html/2605.14040#S5.T3)is the cleanest non\-saturating capability signal in our 7\-axis comparison \(Table[1](https://arxiv.org/html/2605.14040#S2.T1)\)\. Sonnet attains33\.4%33\.4\\%; Physics\-R1 binary at the 3\-seed mean reaches26\.3±1\.7%\\mathbf\{26\.3\\pm 1\.7\}\\%\(per\-seed\{25\.6,25\.0,28\.2\}\\\{25\.6,25\.0,28\.2\\\}across seeds 42/17/23\), exceeding every open\-source baseline \(Qwen3\-VL\-32B13\.2%13\.2\\%, 8B8\.0%8\.0\\%, InternVL34\.0%4\.0\\%\) and the non\-Sonnet closed APIs \(GPT\-4o19\.5%19\.5\\%, Gemini 2\.5 Pro12\.2%12\.2\\%\), trailing only Sonnet by7\.17\.1pp\.

#### Reward\-shape ablation: dense gives a small saturated\-MCQ benefit, binary wins on open\-ended\.

Under problem\-level liberal Sonnet\-judge scoring \(seed 42\), dense at step 60 slightly leads on saturated MCQ \(PhyX\-mini78\.378\.3vs\. binary78\.078\.0; PhyX\-3k77\.577\.5vs\.76\.976\.9\) but trails binary on every non\-MCQ split: PhysReason \(23\.323\.3vs\.32\.2\\mathbf\{32\.2\},\+8\.9\+8\.9pp\), OlympiadBench\-Physics \(40\.540\.5vs\.45\.4\\mathbf\{45\.4\},\+4\.9\+4\.9pp\), andPhysOlym\-Aliberal \(19\.219\.2vs\.25\.6\\mathbf\{25\.6\},\+6\.4\+6\.4pp\)\. On PUB\-OE, dense and binary are within0\.70\.7pp \(37\.737\.7vs\.37\.037\.0\)\. The binary advantage is concentrated on the multi\-sub\-part numerical\-answer column \(PR\) and the held\-out audited olympiad column \(PhysOlym\-A\)—the two columns where the recipe contribution should matter most\. Dense is a reward\-shape ablation; binary is the recommended default\. The five\-component reward drop\-out \(Table[11](https://arxiv.org/html/2605.14040#A8.T11)\), recipe\-flag, and SFT data\-scaling ablations are left to follow\-up work\.

## 6Discussion and Limitations

Physics\-R1 uses unmodified GSPO\+DAPO; the dense reward and recommended baseline \(Algorithm[3](https://arxiv.org/html/2605.14040#alg3)\) are reproducibility artifacts, not method claims\. The audit pipeline catches verbatim and lightly\-paraphrased duplicates and is empirically robust to both embedder choice \(Spearmanρ=0\.78\\rho\{=\}0\.78vs\.text\-embedding\-3\-large; OpenAI candidate set is a strict subset of mxbai’s at every threshold tested; Appendix[A](https://arxiv.org/html/2605.14040#A1)\) and judge choice \(Sonnet 4\.5 vs\. GPT\-4o cross\-judgeκ=0\.44\\kappa\{=\}0\.44on a 50\-problemPhysOlym\-Asubset, with GPT\-4o more lenient—self\-grading direction is opposite to the feared bias; Appendix[D](https://arxiv.org/html/2605.14040#A4)\); the Sonnet\-as\-judge13\.9%13\.9\\%unjudgeable rate is a disclosed noise floor\. The cross\-lingual finding is Sonnet\-4\.5\-specific onn=59n\{=\}59paired items \(65\.7%65\.7\\%MC power, all three tests rejectH0H\_\{0\}\); direction is pre\-registered to reverse for cross\-lingual\-weak models \(Appendix[H\.4](https://arxiv.org/html/2605.14040#A8.SS4), item viii\)\.

## 7Conclusion

Three findings—𝟏𝟑𝟒\\mathbf\{134\}near\-duplicates in SciInstruct surfaced only by three\-stage audit \(Jaccard→\\tocosine→\\toHaiku\-4\.5 judge\), a1717\-pp Estonian–English translation delta on identical olympiad problems, and a4646\-pp format\-and\-novelty gradient on fixed Sonnet 4\.5 weights—motivate four released artifacts:PhysCorp\-A\(6,4326\{,\}432\-record audited corpus, fully Stage\-3 clean against all six public physics evals; Table[2](https://arxiv.org/html/2605.14040#S3.T2)\),PhysR1Corp\(2,2682\{,\}268\-record closed\-form RL pool\),PhysOlym\-A\(500500\-problem held\-out olympiad eval,99\.8%99\.8\\%novel\-source\), and Physics\-R1, a binary\-reward GSPO\+DAPO recipe that liftsPhysOlym\-Aliberal\+18\.3\+18\.3pp over the 8B base at the 3\-seed mean \(8\.0→26\.3±1\.78\.0\{\\to\}\\mathbf\{26\.3\\pm 1\.7\}, still7\.17\.1pp below Sonnet 4\.5; per\-seed\{25\.6,25\.0,28\.2\}\\\{25\.6,25\.0,28\.2\\\}across seeds\{42,17,23\}\\\{42,17,23\\\}on the auditedPhysR1Corp\)\. Audit\-pipeline robustness to embedder and judge choice is established in §[6](https://arxiv.org/html/2605.14040#S6)\(Appendices[A](https://arxiv.org/html/2605.14040#A1),[D](https://arxiv.org/html/2605.14040#A4)\)\. We recommend binary correctness reward as the deployable default \(variance\-optimal under GSPO with group\-normalized advantages, Goodhart\-robust against unit/format proxies; §[4](https://arxiv.org/html/2605.14040#S4)\); reward\-component drop\-out \(Table[11](https://arxiv.org/html/2605.14040#A8.T11)\) is left to follow\-up work; the 3\-seed mean reported in Table[3](https://arxiv.org/html/2605.14040#S5.T3)\(σ≤3\.3\\sigma\\leq 3\.3pp on PUB\-OE, OlymBench\-Phys, andPhysOlym\-A;σ=6\.4\\sigma\{=\}6\.4pp on PhysReason driven by a seed\-42 outlier\) demonstrates that the audited corpus retains training signal across seeds\.

## Acknowledgments

The author thanks Kevin Zhou for granting permission to redistribute his olympiad handouts under CC BY\-NC 4\.0, the Estonian Physics Olympiad committee for making their archived problems and solutions publicly available at[https://fyysika\.ee/](https://fyysika.ee/), the international olympiad committees \(IPhO, NBPhO, EuPhO, APhO, USAPhO, INPhO\) for the public archives that enabled the novel\-source held\-out evaluation, and the maintainers of the public physics\-VL benchmarks \(PhyX, MMMU\-Pro, OlympiadBench, UGPhysics, PhysReason, PhysUniBench\) whose released training pools and evals enabled the contamination audit reported in this work\. Compute support for Physics\-R1 training and evaluation was provided by RunPod\.

## References

- S\. Ahuja, V\. Gumma, and S\. Sitaram \(2024\)Contamination report for multilingual benchmarks\.Note:Tests 7 LLMs on multilingual benchmarks; finds nearly all are contaminated\.External Links:2410\.16186Cited by:[§2](https://arxiv.org/html/2605.14040#S2.SS0.SSS0.Px3.p1.1)\.
- M\. Akhtar, O\. Benjelloun, C\. Conforti,et al\.\(2024\)Croissant: a metadata format for ML\-ready datasets\.InNeurIPS Workshop on Data\-centric Machine Learning Research,External Links:[Link](https://arxiv.org/abs/2403.19546)Cited by:[§G\.7](https://arxiv.org/html/2605.14040#A7.SS7.SSS0.Px1.p1.1)\.
- American Association of Physics Teachers \(2025\)U\.s\. physics olympiad: archived problems and solutions\.Note:[https://www\.aapt\.org/physicsteam/](https://www.aapt.org/physicsteam/)Cited by:[§3\.1](https://arxiv.org/html/2605.14040#S3.SS1.p1.4)\.
- Anthropic \(2025\)Claude sonnet 4\.5\.Note:Model used as judge and frontier\-baseline reference throughout this paper\.External Links:[Link](https://www.anthropic.com/news/claude-sonnet-4-5)Cited by:[§1](https://arxiv.org/html/2605.14040#S1.SS0.SSS0.Px2.p1.6)\.
- Asian Physics Olympiad Committee \(2025\)Asian physics olympiad: archived problems and solutions\.Note:[https://apho2025\.fkfi\.lt/](https://apho2025.fkfi.lt/)Cited by:[§3\.1](https://arxiv.org/html/2605.14040#S3.SS1.p1.4)\.
- W\. Chow, J\. Mao, B\. Li, D\. Seita, V\. Guizilini, and Y\. Wang \(2025\)PhysBench: benchmarking and enhancing vision\-language models for physical world understanding\.InICLR,Note:Oral;[https://openreview\.net/forum?id=Q6a9W6kzv5](https://openreview.net/forum?id=Q6a9W6kzv5)Cited by:[§2](https://arxiv.org/html/2605.14040#S2.SS0.SSS0.Px3.p1.1)\.
- J\. Cohen \(1960\)A coefficient of agreement for nominal scales\.Educational and Psychological Measurement20\(1\),pp\. 37–46\.Cited by:[§5](https://arxiv.org/html/2605.14040#S5.SS0.SSS0.Px1.p1.7)\.
- DeepSeek\-AI \(2025\)DeepSeek\-r1: incentivizing reasoning capability in llms via reinforcement learning\.arXiv preprint arXiv:2501\.12948\.Cited by:[§1](https://arxiv.org/html/2605.14040#S1.p2.1),[§2](https://arxiv.org/html/2605.14040#S2.SS0.SSS0.Px1.p1.1)\.
- J\. Dekoninck, M\. N\. Mueller, and M\. Vechev \(2024\)ConStat: performance\-based contamination detection in large language models\.InNeurIPS,Cited by:[§2](https://arxiv.org/html/2605.14040#S2.SS0.SSS0.Px3.p1.1)\.
- Estonian Physics Olympiad \(2018\)Estonian physics olympiad: problem collection 2004–2018\.Note:[https://www\.fyysika\.ee/](https://www.fyysika.ee/)Cited by:[§3\.1](https://arxiv.org/html/2605.14040#S3.SS1.p1.4)\.
- EuPhO Committee \(2025\)European physics olympiad: archived problems and solutions\.Note:[https://eupho\.ee/](https://eupho.ee/)Cited by:[§3\.1](https://arxiv.org/html/2605.14040#S3.SS1.p1.4)\.
- T\. Gebru, J\. Morgenstern, B\. Vecchione, J\. W\. Vaughan, H\. Wallach, H\. Daume III, and K\. Crawford \(2021\)Datasheets for datasets\.Communications of the ACM\.Cited by:[Appendix G](https://arxiv.org/html/2605.14040#A7.p1.1)\.
- E\. Glazer, E\. Erdil, T\. Besiroglu, D\. Chicharro,et al\.\(2024\)FrontierMath: a benchmark for evaluating advanced mathematical reasoning in AI\.arXiv preprint arXiv:2411\.04872\.Cited by:[§2](https://arxiv.org/html/2605.14040#S2.SS0.SSS0.Px3.p1.1),[Table 1](https://arxiv.org/html/2605.14040#S2.T1.22.14.3)\.
- C\. He, R\. Luo, Y\. Bai, S\. Hu,et al\.\(2024\)OlympiadBench: a challenging benchmark for promoting AGI with olympiad\-level bilingual multimodal scientific problems\.InACL,Cited by:[§1](https://arxiv.org/html/2605.14040#S1.SS0.SSS0.Px3.p1.4),[§2](https://arxiv.org/html/2605.14040#S2.SS0.SSS0.Px2.p1.1),[Table 1](https://arxiv.org/html/2605.14040#S2.T1.16.8.3)\.
- Homi Bhabha Centre for Science Education \(2025\)Indian national physics olympiad: archived problems and solutions\.Note:[https://olympiads\.hbcse\.tifr\.res\.in/](https://olympiads.hbcse.tifr.res.in/)Cited by:[§3\.1](https://arxiv.org/html/2605.14040#S3.SS1.p1.4)\.
- International Physics Olympiad \(2025\)International physics olympiad: archived problems and solutions\.Note:[https://ipho\-unofficial\.org/](https://ipho-unofficial.org/)Cited by:[§3\.1](https://arxiv.org/html/2605.14040#S3.SS1.p1.4)\.
- S\. Lee, A\. Shakir, D\. Koenig, and J\. Lipp \(2024\)Open source strikes bread \- new fluffy embeddings model\.Mixedbread\.Note:mxbai\-embed\-large\-v1: 1024\-dim sentence embedding model used in our Stage\-2 audit\.External Links:[Link](https://www.mixedbread.com/blog/mxbai-embed-large-v1)Cited by:[Appendix A](https://arxiv.org/html/2605.14040#A1.SS0.SSS0.Px2.p2.3)\.
- F\. Meng, L\. Du, Z\. Liu, Z\. Zhou, Q\. Lu, D\. Fu, B\. Han, B\. Shi, W\. Wang, J\. He, K\. Zhang, T\. Zhao, Y\. Qiao, and P\. Luo \(2025\)MM\-eureka: exploring visual aha\-moment with rule\-based large\-scale reinforcement learning\.arXiv preprint arXiv:2503\.07365\.Cited by:[Table 8](https://arxiv.org/html/2605.14040#A8.T8.9.5.3),[§1](https://arxiv.org/html/2605.14040#S1.p2.1),[§2](https://arxiv.org/html/2605.14040#S2.SS0.SSS0.Px1.p1.1),[§2](https://arxiv.org/html/2605.14040#S2.SS0.SSS0.Px2.p1.1),[§4](https://arxiv.org/html/2605.14040#S4.p1.12),[§5\.2](https://arxiv.org/html/2605.14040#S5.SS2.SSS0.Px1.p1.1)\.
- NBPhO Committee \(2025\)Nordic\-baltic physics olympiad: archived problems and solutions\.Note:[https://www\.nbpho\.eu/](https://www.nbpho.eu/)Cited by:[§3\.1](https://arxiv.org/html/2605.14040#S3.SS1.p1.4)\.
- OpenStax \(2024\)College Physics 2e and University Physics Volumes 1\-3\.Note:[https://openstax\.org/subjects/science](https://openstax.org/subjects/science)Cited by:[§3\.1](https://arxiv.org/html/2605.14040#S3.SS1.p1.4)\.
- L\. Phan, A\. Gatti, Z\. Han, N\. Li, J\. Hu, H\. Zhang, C\. B\. C\. Zhang, M\. Shaaban, J\. Ling, S\. Shi,et al\.\(2025\)Humanity’s last exam\.External Links:2501\.14249Cited by:[§2](https://arxiv.org/html/2605.14040#S2.SS0.SSS0.Px3.p1.1),[Table 1](https://arxiv.org/html/2605.14040#S2.T1.24.16.3)\.
- S\. Qiu, S\. Guo, Z\. Song, Y\. Sun, Z\. Cai, J\. Wei, T\. Luo, Y\. Yin, H\. Zhang, Y\. Hu,et al\.\(2025\)PHYBench: holistic evaluation of physical perception and reasoning in LLMs\.InNeurIPS Datasets and Benchmarks Track,Note:[https://openreview\.net/forum?id=brG8FPq1cf](https://openreview.net/forum?id=brG8FPq1cf)Cited by:[§2](https://arxiv.org/html/2605.14040#S2.SS0.SSS0.Px2.p1.1),[Table 1](https://arxiv.org/html/2605.14040#S2.T1.9.1.2)\.
- Qwen Team \(2025\)Qwen3\-VL\.Note:[https://huggingface\.co/Qwen/Qwen3\-VL\-8B\-Thinking](https://huggingface.co/Qwen/Qwen3-VL-8B-Thinking)Vision\-language model releaseCited by:[§1](https://arxiv.org/html/2605.14040#S1.p2.1),[§4](https://arxiv.org/html/2605.14040#S4.p1.12),[§5\.2](https://arxiv.org/html/2605.14040#S5.SS2.SSS0.Px1.p1.1)\.
- R\. Rafailov, A\. Sharma, E\. Mitchell, C\. D\. Manning, S\. Ermon, and C\. Finn \(2023\)Direct preference optimization: your language model is secretly a reward model\.InNeurIPS,Cited by:[Table 8](https://arxiv.org/html/2605.14040#A8.T8.16.14.1)\.
- M\. Ravaut, B\. Ding, F\. Jiao,et al\.\(2024\)A comprehensive survey of contamination detection methods in large language models\.External Links:2404\.00699Cited by:[§2](https://arxiv.org/html/2605.14040#S2.SS0.SSS0.Px3.p1.1)\.
- O\. Sainz, J\. A\. Campos, I\. García\-Ferrero,et al\.\(2023\)NLP evaluation in trouble: on the need to measure LLM data contamination for each benchmark\.InEMNLP Findings,Cited by:[§2](https://arxiv.org/html/2605.14040#S2.SS0.SSS0.Px3.p1.1)\.
- Z\. Shao, P\. Wang, Q\. Zhu,et al\.\(2024\)DeepSeekMath: pushing the limits of mathematical reasoning in open language models\.arXiv preprint arXiv:2402\.03300\.Cited by:[Table 8](https://arxiv.org/html/2605.14040#A8.T8.5.1.2),[§1](https://arxiv.org/html/2605.14040#S1.p2.1),[§2](https://arxiv.org/html/2605.14040#S2.SS0.SSS0.Px1.p1.1)\.
- H\. Shen, T\. Wu, Q\. Han,et al\.\(2025\)PhyX: does your model have the “wits” for physical reasoning?\.arXiv preprint arXiv:2505\.15929\.Cited by:[§H\.3](https://arxiv.org/html/2605.14040#A8.SS3.p1.3),[§1](https://arxiv.org/html/2605.14040#S1.SS0.SSS0.Px3.p1.4),[§2](https://arxiv.org/html/2605.14040#S2.SS0.SSS0.Px2.p1.1),[§5\.2](https://arxiv.org/html/2605.14040#S5.SS2.SSS0.Px1.p1.1),[Table 3](https://arxiv.org/html/2605.14040#S5.T3)\.
- G\. Sheng, C\. Zhang, Z\. Ye, X\. Wu,et al\.\(2024\)HybridFlow: a flexible and efficient RLHF framework \(verl\)\.arXiv preprint arXiv:2409\.19256\.Cited by:[§4](https://arxiv.org/html/2605.14040#S4.p1.12)\.
- A\. K\. Singh, M\. Y\. Kocyigit, A\. Poulton,et al\.\(2024\)Evaluation data contamination in LLMs: how do we measure it and \(when\) does it matter?\.arXiv preprint arXiv:2411\.03923\.Cited by:[§2](https://arxiv.org/html/2605.14040#S2.SS0.SSS0.Px3.p1.1)\.
- Stack Exchange Inc\. \(2024\)Physics stack exchange: question and answer archive\.Note:[https://physics\.stackexchange\.com/](https://physics.stackexchange.com/)Cited by:[§3\.1](https://arxiv.org/html/2605.14040#S3.SS1.p1.4)\.
- G\. Tsoukalas, J\. Lee, J\. Jennings, Y\. Xin,et al\.\(2024\)PutnamBench: evaluating neural theorem\-provers on the putnam mathematical competition\.NeurIPS Datasets & Benchmarks\.Cited by:[§2](https://arxiv.org/html/2605.14040#S2.SS0.SSS0.Px3.p1.1),[Table 1](https://arxiv.org/html/2605.14040#S2.T1.18.10.3)\.
- C\. J\. Wang, D\. Lee, C\. Menghini,et al\.\(2025a\)EnigmaEval: a benchmark of long multimodal reasoning challenges\.Note:Frontier\-model pass\-rate so low that contamination is empirically dismissed\.External Links:2502\.08859Cited by:[§2](https://arxiv.org/html/2605.14040#S2.SS0.SSS0.Px3.p1.1)\.
- L\. Wang, E\. Su, J\. Liu,et al\.\(2025b\)PhysUniBench: a multi\-modal physics reasoning benchmark at undergraduate level\.External Links:2506\.17667Cited by:[§2](https://arxiv.org/html/2605.14040#S2.SS0.SSS0.Px2.p1.1),[Table 1](https://arxiv.org/html/2605.14040#S2.T1.12.4.4)\.
- Y\. Wang, X\. Ma, G\. Zhang, Y\. Ni, A\. Chandra, S\. Guo, W\. Ren, A\. Arulraj, X\. He, Z\. Jiang,et al\.\(2024\)MMLU\-Pro: a more robust and challenging multi\-task language understanding benchmark\.InNeurIPS Datasets and Benchmarks Track,Note:Spotlight;[https://openreview\.net/forum?id=y10DM6R2r3](https://openreview.net/forum?id=y10DM6R2r3)Cited by:[Table 1](https://arxiv.org/html/2605.14040#S2.T1.25.17.2)\.
- M\. Wu, W\. Wang, S\. Liu,et al\.\(2025\)The bitter lesson learned from 2,000\+ multilingual benchmarks\.External Links:2504\.15521Cited by:[§2](https://arxiv.org/html/2605.14040#S2.SS0.SSS0.Px3.p1.1)\.
- X\. Xu, Q\. Xu, T\. Xiao,et al\.\(2025\)UGPhysics: a comprehensive benchmark for undergraduate physics reasoning with large language models\.arXiv preprint arXiv:2502\.00334\.Cited by:[Appendix E](https://arxiv.org/html/2605.14040#A5.SS0.SSS0.Px1.p1.1),[§2](https://arxiv.org/html/2605.14040#S2.SS0.SSS0.Px2.p1.1),[Table 1](https://arxiv.org/html/2605.14040#S2.T1.41.36.1),[§3\.1](https://arxiv.org/html/2605.14040#S3.SS1.p1.4)\.
- W\. Xuan, R\. Yang, H\. Qi, Q\. Zeng, Y\. Xiao, Y\. Feng, J\. Liu, J\. Hou, J\. Zhao, W\. Yu,et al\.\(2025\)MMLU\-ProX: a multilingual benchmark for advanced reasoning across languages\.Note:29 languages, 11,829 identical questions per language\.External Links:2503\.10497Cited by:[§2](https://arxiv.org/html/2605.14040#S2.SS0.SSS0.Px3.p1.1)\.
- S\. Yang, W\. Chiang, L\. Zheng, J\. E\. Gonzalez, and I\. Stoica \(2023\)Rethinking benchmark and contamination for language models with rephrased samples\.Note:Demonstrates that n\-gram contamination audits miss rephrased duplicates; motivates Stage\-2 of our pipeline\.External Links:2311\.04850Cited by:[§2](https://arxiv.org/html/2605.14040#S2.SS0.SSS0.Px3.p1.1)\.
- Q\. Yu, Z\. Zhang, R\. Zhu,et al\.\(2025\)DAPO: an open\-source LLM reinforcement learning system at scale\.arXiv preprint arXiv:2503\.14476\.Cited by:[Table 8](https://arxiv.org/html/2605.14040#A8.T8.7.3.3),[§1](https://arxiv.org/html/2605.14040#S1.p2.1),[§2](https://arxiv.org/html/2605.14040#S2.SS0.SSS0.Px1.p1.1),[§4](https://arxiv.org/html/2605.14040#S4.p1.4),[§5\.2](https://arxiv.org/html/2605.14040#S5.SS2.SSS0.Px1.p1.1)\.
- X\. Yue, Y\. Ni, K\. Zhang, T\. Zheng,et al\.\(2024a\)MMMU: a massive multi\-discipline multimodal understanding and reasoning benchmark for expert AGI\.InCVPR,Cited by:[Appendix E](https://arxiv.org/html/2605.14040#A5.SS0.SSS0.Px4.p1.1),[§2](https://arxiv.org/html/2605.14040#S2.SS0.SSS0.Px2.p1.1),[§3\.1](https://arxiv.org/html/2605.14040#S3.SS1.p1.4)\.
- X\. Yue, T\. Zheng, Y\. Ni, Y\. Wang, K\. Zhang, S\. Tong, Y\. Sun, B\. Yu, G\. Zhang, H\. Sun,et al\.\(2024b\)MMMU\-Pro: a more robust multi\-discipline multimodal understanding benchmark\.InarXiv preprint,External Links:2409\.02813Cited by:[§2](https://arxiv.org/html/2605.14040#S2.SS0.SSS0.Px2.p1.1),[Table 1](https://arxiv.org/html/2605.14040#S2.T1.27.19.3)\.
- D\. Zhang, Z\. Hu, S\. Zhoubian, Z\. Du, K\. Yang, Z\. Wang, Y\. Yue, Y\. Dong, and J\. Tang \(2024\)SciInstruct: a self\-reflective instruction annotated dataset for training scientific language models\.InNeurIPS Datasets and Benchmarks Track,Note:[https://openreview\.net/forum?id=LC1QAqhePv](https://openreview.net/forum?id=LC1QAqhePv)Cited by:[Table 1](https://arxiv.org/html/2605.14040#S2.T1.41.39.1)\.
- X\. Zhang, Y\. Dong, Y\. Wu,et al\.\(2025\)PhysReason: a comprehensive benchmark towards physics\-based reasoning\.arXiv preprint arXiv:2502\.12054\.Cited by:[Appendix E](https://arxiv.org/html/2605.14040#A5.SS0.SSS0.Px5.p1.1),[§2](https://arxiv.org/html/2605.14040#S2.SS0.SSS0.Px2.p1.1),[Table 1](https://arxiv.org/html/2605.14040#S2.T1.14.6.3),[§3\.1](https://arxiv.org/html/2605.14040#S3.SS1.p1.4)\.
- C\. Zheng, S\. Liu, M\. Li,et al\.\(2025\)Group sequence policy optimization\.arXiv preprint arXiv:2507\.18071\.Cited by:[Table 8](https://arxiv.org/html/2605.14040#A8.T8.7.3.3),[§1](https://arxiv.org/html/2605.14040#S1.p2.1),[§2](https://arxiv.org/html/2605.14040#S2.SS0.SSS0.Px1.p1.1),[§4](https://arxiv.org/html/2605.14040#S4.p1.4),[§5\.2](https://arxiv.org/html/2605.14040#S5.SS2.SSS0.Px1.p1.1)\.
- K\. Zhou \(2018\)Olympiad physics handouts\.Note:[https://knzhou\.github\.io/](https://knzhou.github.io/)Cited by:[§3\.1](https://arxiv.org/html/2605.14040#S3.SS1.p1.4)\.
- Y\. Zhu, J\. Wang, Y\. Li,et al\.\(2025\)OIBench: benchmarking strong reasoning models with olympiad in informatics\.InNeurIPS Datasets and Benchmarks Track,Note:arXiv:2506\.10481Cited by:[Table 1](https://arxiv.org/html/2605.14040#S2.T1.20.12.3)\.

## Appendix AAudit pipeline details and worked examples

#### Stage\-3 LLM\-judge: SciInstruct cosine\-bucket near\-duplicate rate\.

For each of the4,8464\{,\}846SciInstruct↔\\leftrightarroweval Stage\-2 candidate pairs \(cos≥0\.85\\geq 0\.85\), the Stage\-3 Haiku\-4\.5 judge receives both problem statements and returns*close duplicate*\(paraphrase / numeric variation of the same problem\) or*same\-topic neighbor*\(related physics, distinct setup\)\. The close\-duplicate share is sharply threshold\-driven: atcos≥0\.95\\cos\{\\geq\}0\.95every flagged pair is a close duplicate; at the threshold edge\[0\.85,0\.87\)\[0\.85,0\.87\)only1\.5%1\.5\\%are\. Stage\-2 still surfaces genuinely close physics content at the low\-cos end—the topic overlap is real, just not strict duplication\.

Cosine bucketN pairsclose duplicatesame\-topic neighbor% close\-dup\[0\.95,0\.99\)\[0\.95,0\.99\)1717𝟏𝟕\\mathbf\{17\}0100\.0%\\mathbf\{100\.0\\%\}\[0\.90,0\.95\)\[0\.90,0\.95\)12712710101171177\.9%7\.9\\%\[0\.87,0\.90\)\[0\.87,0\.90\)1,1591\{,\}15954541,1051\{,\}1054\.7%4\.7\\%\[0\.85,0\.87\)\[0\.85,0\.87\)3,5433\{,\}54353533,4903\{,\}4901\.5%1\.5\\%Total𝟒,𝟖𝟒𝟔\\mathbf\{4\{,\}846\}𝟏𝟑𝟒\\mathbf\{134\}𝟒,𝟕𝟏𝟐\\mathbf\{4\{,\}712\}2\.8%\\mathbf\{2\.8\\%\}The pattern matches the threshold\-sensitivity table \(Table[4](https://arxiv.org/html/2605.14040#A1.T4)\): cosine≥0\.95\\geq 0\.95is precision\-dominant,≥0\.85\\geq 0\.85is recall\-dominant; Stage\-3 is the precision filter that converts a recall\-dominant Stage\-2 candidate set into a high\-precision near\-duplicate set\. Per\-eval Stage\-3 near\-duplicate counts \(out of Stage\-2 raw, matching Table[2](https://arxiv.org/html/2605.14040#S3.T2)\): PhysReason\-full36/2,68736/2\{,\}687\(1\.3%1\.3\\%\), PhysUniBench\-en22/1,02722/1\{,\}027\(2\.1%2\.1\\%\), PhyX\-mini46/70346/703\(6\.5%6\.5\\%\), OlymBench\-Phys15/13015/130\(11\.5%11\.5\\%\),PhysOlym\-A8/1638/163\(4\.9%4\.9\\%\), MMMU\-Pro7/1417/141\(5\.0%5\.0\\%\)\.

#### External\-corpora↔\\leftrightarrowheld\-out pairwise audit \(Stage 2, OlympiadBench shared\-source\)\.

The Stage\-1 / Stage\-2 cells of the main contamination matrix \(Table[2](https://arxiv.org/html/2605.14040#S3.T2)\) cover competitor training pools against the held\-out evals\. The complementary cross\-channel audit below pairs the four public physics\-olympiad benchmarks againstPhysOlym\-Aand PhyX 1000q, exposing shared\-source paraphrase overlap between Olympiad\-style problems composed independently\. The single Stage\-1 hit*OlympiadBench\-Physics*→\\toPhysOlym\-Ais the EuPhO 2020 “Mechanical accelerator” problem that grounds the99\.8%99\.8\\%\(rather than100%100\\%\) novel\-source claim\.

PhysOlym\-APhyX 1000qOlympiadBench\-Physics \(692\)1 / 1360 / 2PhysReason\-mini \(200\)0 / 2—PhysReason\-full \(1,200\)0 / 35—PhysUniBench\-en \(1,022\)0 / 27—
Stage 1 \(5\-gram Jaccard\)\. Tokenize each problem statement with a unicode word tokenizer; build the 5\-gram shingle set; compute Jaccard similarity over shingle sets; flag pairs with similarity≥0\.4\\geq 0\.4\. The threshold is calibrated against worked examples: an OpenStax College Physics↔\\leftrightarrowUniversity Physics duplicate scores Jaccard 1\.0; a UGPhysics paraphrase missed by Stage 1 \(Jaccard 0\.31\) is flagged at Stage 2 \(cosine 0\.91\); plausible misses include numerical\-value substitution \(same setup, different constants\) and translation across languages\. Stage 2 \(embedding cosine\)\. Encode each problem statement withmxbai\-embed\-large\-v1\[Leeet al\.,[2024](https://arxiv.org/html/2605.14040#bib.bib52)\]\(1024\-dim normalized embeddings\); compute pairwise cosine; flag pairs with cosine≥0\.85\\geq 0\.85\.

### A\.1Threshold\-Sensitive Leakage Finding on a Researcher\-Curated Baseline

Table 4:Threshold\-sensitivity of the leakage finding on a researcher\-curated baseline\.A 1,679\-record sample drawn fromPhysCorp\-pre\-audit\(the released 14,294\-record pre\-audit pool\) under conventional 5\-gram\-Jaccard \+ within\-pool embedding dedup is paired against a 500\-record internal analysis eval; each cell reports records flagged atJ≥jthrJ\\geq j\_\{\\text\{thr\}\}*or*cos≥cthr\\cos\\geq c\_\{\\text\{thr\}\}\. The boxed cell \(J≥0\.4J\\geq 0\.4,cos≥0\.85\\cos\\geq 0\.85\) is the operating point referenced throughout the paper\. Stage\-1 alone \(Jaccard\) is bimodal: all 56 leaks are exact atJ=1\.0J\{=\}1\.0\. Stage\-2 \(cosine\) catches an additional∼92\{\\sim\}92paraphrases the n\-gram audit misses at the published threshold; loosening cosine to0\.800\.80pushes the leak rate above27%27\\%\.cosine thresholdJ≥0\.3J\\geq 0\.3J≥0\.4J\\geq 0\.4J≥0\.5J\\geq 0\.5cos≥0\.80\\cos\\geq 0\.80\(lax\)455 \(27\.1%\)455 \(27\.1%\)455 \(27\.1%\)cos≥0\.85\\cos\\geq 0\.85\(paper op\.\)148 \(8\.8%\)148 \(8\.8%\)148 \(8\.8%\)cos≥0\.90\\cos\\geq 0\.90\(strict\)79 \(4\.7%\)79 \(4\.7%\)79 \(4\.7%\)*Single\-stage rates \(no union\):*Stage\-1 only \(J≥jthrJ\\geq j\_\{\\text\{thr\}\}, no cos\)56 \(3\.3%\)56 \(3\.3%\)56 \(3\.3%\)Stage\-2 only \(cos≥cthr\\cos\\geq c\_\{\\text\{thr\}\}, noJJ\)455 \(27\.1%\)148 \(8\.8%\)79 \(4\.7%\)To ground Finding 1 we audit a researcher\-curated baseline—a 1,679\-record sample drawn fromPhysCorp\-pre\-audit\(the released 14,294\-record pre\-audit pool\) under conventional 5\-gram\-Jaccard \+ within\-pool embedding dedup, paired against a 500\-record internal analysis eval\. The 500\-record eval is held internal to ground this finding and is distinct fromPhysOlym\-A\(constructed post\-audit\); the 1,679\-record sample is reproducible from the releasedPhysCorp\-pre\-auditso the audit can be re\-run by downstream users\. Stage\-1 catches 56 records \(3\.3%\) all atJ=1\.0J\{=\}1\.0\(bimodal distribution—verbatim duplication under our normalization\)\. Stage\-2 cosine≥0\.85\\geq 0\.85flags an additional 92 records \(5\.5%\), bringing the joint leak rate to 148 \(8\.8%\); the cosine dimension sweeps4\.74\.7–27\.1%27\.1\\%acrosscos∈\{0\.90,0\.85,0\.80\}\\cos\\in\\\{0\.90,0\.85,0\.80\\\}while Jaccard is flat at 3\.3% \(Table[4](https://arxiv.org/html/2605.14040#A1.T4)\)\. The 148 flagged decompose into 56 exact \(J=1\.0J\{=\}1\.0\), 23 strong paraphrases \(cos≥0\.90\\cos\\geq 0\.90\), 69 weak paraphrases \(0\.85≤cos<0\.900\.85\\leq\\cos<0\.90\)\. The 5\.5\-pp gap is the rephrasing dark\-matter that single\-stage audits miss because the*same*problem appears across upstream aggregations under wording that evadesJ≥0\.4J\\geq 0\.4but trips cosine≥0\.85\\geq 0\.85\.

#### Pipeline reproducibility note\.

Threshold\-sensitivity numbers were produced byaudit\_pipeline/threshold\_sensitivity\.py\. Normalization: lower\-case, strip LaTeX commands \(\\frac,\\sqrt, …\), remove$\{\}\[ \]\(\)delimiters, collapse whitespace, drop<5<5\-word shingles\. Embedding:sentence\-transformers 5\.4\.1, batch 32,L2L\_\{2\}\-normalized\. Best\-overlap via inverted\-index pruning \(Stage 1\) and fullN×MN\{\\times\}Mmatmul \(Stage 2\)\. Saved scores inthreshold\_sensitivity\_scores\.npz\.

#### Bimodal Jaccard distribution\.

Every Stage\-1 leak in our researcher\-curated baseline sits atJ=1\.0J\{=\}1\.0: aggressive normalization collapses surface variance, so shared problem statements yield identical shingle sets while distinct records land belowJ=0\.3J\{=\}0\.3\. Cosine≥0\.85\\geq 0\.85is therefore the sole paraphrase\-class detector here; bimodality is corpus\-specific\.

#### Re\-audit and the cleaned PhysR1Corp\.

Starting from a2,4332\{,\}433\-record candidate closed\-form pool, three sequential cleanup passes against the six paper\-canonical comparison corpora \(PhyX, MMMU\-Pro Physics, OlympiadBench\-Physics, PhysReason\-full, PhysUniBench\-en,PhysOlym\-A\) produced the releasedPhysR1Corp: \(i\) MMMU\-Pro Physics jointJ≥0\.4∨cos≥0\.85J\\geq 0\.4\\vee\\cos\\geq 0\.85re\-audit dropped 87 records \(Stage\-1: 16 records,16/60=26\.7%16/60\{=\}26\.7\\%MMMU\-Pro coverage; Stage\-2 union: 87 records,52/60=86\.7%52/60\{=\}86\.7\\%\); \(ii\) PhyX\-mini cosine≥0\.85\\geq 0\.85Stage\-2 audit dropped 69 records, all real\-near\-duplicate physics\-problem variants confirmed by manual inspection \(e\.g\. MgF2anti\-reflection\-coating problem with slightly tweakednglassn\_\{\\text\{glass\}\}and option labels, top cos0\.930\.93\); \(iii\) PhysUniBench\-en cosine≥0\.85\\geq 0\.85Stage\-2 audit dropped 9 template duplicates \(top cos0\.930\.93, shared problem\-template stems\)\. Total dropped:87\+69\+9=16587\+69\+9=165records \(with 3 records flagged in multiple channels netting to87\+78=16587\+78=165unique\)\. The released𝟐,𝟐𝟔𝟖\\mathbf\{2\{,\}268\}\-recordPhysR1Corpretains 19 Stage\-2 hits across PhysOlym\-A \(3\), MMMU\-Pro \(1\), OlymBench\-Phys \(4\), PhysReason\-full \(11\) classified as same\-topic neighbors on manual inspection \(top cos≤0\.87\\leq 0\.87, different problem setups, dimensionalities, or geometries; Table[2](https://arxiv.org/html/2605.14040#S3.T2)\)\. MMMU\-Pro Physics remains excluded from the headline Table[3](https://arxiv.org/html/2605.14040#S5.T3), the saturation\-gradient narrative in §[5\.1](https://arxiv.org/html/2605.14040#S5.SS1.SSS0.Px1), and the dense\-vs\-binary ablation in §[5\.2](https://arxiv.org/html/2605.14040#S5.SS2.SSS0.Px6), because the eval is small \(60 records\) and prior work has flagged it as contaminated against multiple frontier training corpora; Sonnet’s MMMU\-Pro Physics number is reported only as a frontier\-model reference \(Sonnet was not trained onPhysR1Corp\)\.

#### Embedding\-model and threshold rationale\.

We chosemxbai\-embed\-large\-v1for Stage 2 overbge\-large\-en\-v1\.5,e5\-large\-v2, OpenAItext\-embedding\-3\-large, and Voyagevoyage\-3because it \(i\) is permissively licensed \(Apache 2\.0, no API dependency\); \(ii\) ships 1024\-dim normalized vectors with cosine tuning; \(iii\) scored highest on MTEB physics\-adjacent retrieval at audit time\. Thecos≥0\.85\\cos\\geq 0\.85threshold was calibrated against worked examples \(UGPhysics paraphrase:J=0\.31J\{=\}0\.31Stage\-1\-miss,cos=0\.91\\cos\{=\}0\.91Stage\-2\-catch\); thecos∈\{0\.80,0\.85,0\.90\}\\cos\\in\\\{0\.80,0\.85,0\.90\\\}sensitivity grid brackets the operating point\.

#### Embedder\-sensitivity ablation \(mxbai vs\.text\-embedding\-3\-large\)\.

We re\-encode the releasedPhysCorp\-pre\-auditpool \(14,29414\{,\}294records\) andPhysOlym\-A\(500500records\) under OpenAItext\-embedding\-3\-largeand compute pairwise cosines, then compare per\-train\-record max\-cosine rankings against the mxbai baseline\.Spearmanρ=0\.78\\rho\{=\}0\.78\(p≈0p\{\\approx\}0\); the candidate\-set relationship at the operating threshold is summarized below\.

Thresholdcos≥\\cos\\geqmxbai cand\.text\-embedding\-3\-largecand\.bothonly mxbaionly OpenAI0\.850\.85\(paper op\.\)763763514514514514249249𝟎\\mathbf\{0\}0\.870\.87631631512512512512119119𝟎\\mathbf\{0\}0\.900\.905515515115115115114040𝟎\\mathbf\{0\}0\.950\.955205205065065065061414𝟎\\mathbf\{0\}
text\-embedding\-3\-largeflags a*strict subset*of mxbai’s candidates at every threshold \(only\-OpenAI count is0at all four levels\): every recordtext\-embedding\-3\-largewould catch is also caught by mxbai\. mxbai is therefore the more conservative \(higher\-recall\) Stage\-2 embedder, and the audit cannot have missed any contamination thattext\-embedding\-3\-largewould have surfaced\. The candidate\-set Jaccard at the operating threshold is0\.670\.67\.*Caveat\.*This ablation is on the user\-facing audit case \(released pool↔\\leftrightarrowreleased eval\), not the SciInstruct competitor\-pool case from Table[2](https://arxiv.org/html/2605.14040#S3.T2); the strict\-subset direction does not formally transfer, but theρ=0\.78\\rho\{=\}0\.78rank correlation suggests the SciInstruct134134\-near\-duplicate count is robust in direction under embedder change\. Sensitivity ablation againstvoyage\-3is left to follow\-up work\.

#### Stage\-3 judge model dependence and reproducibility\.

Stage\-3 introduces a dependence on Anthropic’s Haiku\-4\.5 \(the close\-duplicate vs\. same\-topic\-neighbor classifier\)\. To ensure long\-term reproducibility we \(i\) pin the exact model identifier \(claude\-haiku\-4\-5as of 2026\-05\) in the releasedaudit\_three\_stage\.py; \(ii\) release the full per\-pair judge prompts and per\-pair verdict labels \(judge\_labelarrays\) alongside the cosine scores inthreshold\_sensitivity\_scores\.npz, so the contamination flag set is reproducible without re\-querying the API; \(iii\) document a fallback protocol \(Sonnet 4\.5 or GPT\-4o on the same prompt template\) for cases where Haiku\-4\.5 access is no longer available\. Because Stage\-3 is a precision filter applied only to Stage\-2 candidates \(≤0\.5%\\leq 0\.5\\%of the train pool\), its labels are also the most amenable to manual re\-verification by a downstream auditor; the cosine\-bucketed precision profile \(Table[A](https://arxiv.org/html/2605.14040#A1.SS0.SSS0.Px1)\) gives a cheap calibration signal against any future judge\.

## Appendix BHeld\-out splits and corpus annotation schema

The released splits arePhysCorp\-A\(the 6,432\-record audited corpus\) with its closed\-form RL carve\-outPhysR1Corp\(2,268 records\) andPhysOlym\-A\(the 500\-problem held\-out olympiad eval\)\. The annotation schema below applies uniformly across all three\.

#### Eight\-field annotation schema\.

Each record carries the following annotations, generated by Sonnet 4\.5 batch annotation \(3,900 records with full annotation; the remainder carry source\-native labels merged into the same schema for∼\\sim31,000 total label values\):

- •difficulty∈\\in\{1,2,3,4,5\} \(Sonnet\-aggregated\), plus optional source\-native: Estonian organizer\-issued 1–10, Zhou pedagogical 1–5, Zhou advanced \[A\] flag\.
- •concept∈\\in\{Mechanics, Electromagnetism, Quantum, Thermodynamics, Waves, Optics, Modern, Relativity, Particle\}\.
- •problem\_type∈\\in\{Conceptual, Computational, Proof\-based, Experimental\}\.
- •expected\_solution\_length∈\\in\{S, M, L\} \(short / medium / long\)\.
- •math\_level∈\\in\{Algebra, Calculus, Vector, LinearAlgebra, DiffEq\}\.
- •modality∈\\in\{text, multimodal\}\.
- •language\(BCP\-47\):en,et,en\-et\(bilingual paired\)\.
- •license\(SPDX\-style\):CC\-BY\-4\.0,CC\-BY\-SA\-4\.0,CC\-BY\-NC\-4\.0,Public\-Domain\.

#### Worked example record \(JSONL\)\.

A typical record fromPhysOlym\-A:

> \{ "index": "estonian\_2017\_lahtine\_2", "source": "estonian\_olympiad", "license": "CC\-BY\-NC\-4\.0", "language": "en\-et", "messages": \[\{"role": "user", "content": "An ideal gas…"\}\], "solution": "By the first law…\\b​o​x​e​d​\{1\.5\}\\backslash boxed\\\{1\.5\\\}", "images": \[\], "concept": "Thermodynamics", "difficulty": 4, "native\_difficulty": \{"scale": "1\-\-10", "value": 7\}, "problem\_type": "Computational", "expected\_solution\_length": "M", "math\_level": "Calculus", "modality": "text", "audit\_passed": true \}

#### Inter\-annotator agreement\.

For the 3,900\-record fully\-annotated subset, Sonnet 4\.5 is run with two seeds on the same 100 records \(random sample, seed 42\); per\-field Cohen’sκ\\kappabetween the two annotation runs is reported in the released dataset card\. Preliminary inspection showsκ≥0\.85\\kappa\\geq 0\.85onconceptandproblem\_type\(categorical with sharp boundaries\),κ∼0\.7\\kappa\\sim 0\.7ondifficulty\(ordinal with neighboring\-class confusion\), andκ∼0\.6\\kappa\\sim 0\.6onexpected\_solution\_length\. The lowerκ\\kappaonexpected\_solution\_lengthreflects genuine ambiguity in the medium\-vs\-long boundary; downstream users requiring stable solution\-length labels should treat the field as a noisy proxy\.

#### Native difficulty preserved\.

For the 27% ofPhysOlym\-Arecords with Estonian native difficulty 1–10 and the 38% with Zhou pedagogical 1–5 point values, the*Sonnet\-aggregated*difficulty field is reported for cross\-source consistency, but the*native*difficulty is preserved in a separatenative\_difficultyfield with explicitscaleandvaluekeys\. The Sonnet difficulty curve in Table[12](https://arxiv.org/html/2605.14040#A8.T12)uses native Estonian 1–10, not the aggregated 1–5, because native labels avoid the self\-annotation circularity in which the difficulty estimate depends on the same model whose accuracy is being measured\.

## Appendix CReward function and full hyperparameter table

This appendix specifies both reward shapes referenced from §[4](https://arxiv.org/html/2605.14040#S4): the recommended binary correctness rewardrbinr\_\{\\mathrm\{bin\}\}defined inline in §[4](https://arxiv.org/html/2605.14040#S4)\(§[C\.1](https://arxiv.org/html/2605.14040#A3.SS1)\) and the dense five\-component reward of Equation[2](https://arxiv.org/html/2605.14040#S4.E2)reported as an ablation \(§[C\.2](https://arxiv.org/html/2605.14040#A3.SS2)\)\. Both share the matching/extraction primitives below; binary usesransr\_\{\\mathrm\{ans\}\}alone, dense composes all five components and clips\.

### C\.1Binary correctness reward \(recommended\)

The recommended reward is

rbin​\(y,x\)=1​\[Match​\(ExtractBoxed​\(y\),g​\(x\)\)\]∈\{0,1\},r\_\{\\mathrm\{bin\}\}\(y,x\)\\;=\\;\\mathbb\{1\}\\\!\\big\[\\,\\textsc\{Match\}\\\!\\big\(\\textsc\{ExtractBoxed\}\(y\),\\,g\(x\)\\big\)\\,\\big\]\\in\\\{0,1\\\},whereExtractBoxedparses the last\\boxed\{\}via brace\-counting \(handles unlimited nesting, e\.g\.\\sqrt\{\\frac\{T\_0\}\{\\eta\}\}\) andMatchaccepts: \(i\) MCQ\-letter equality on\{\\\{A,B,C,D,…\}\\\}gold; \(ii\) multi\-part numeric agreement within±1%\\pm 1\\%relative tolerance, afterlatex\_to\_plainnormalization \(\\text\{\}/\\mathrm\{\}stripped,\\frac\{a\}\{b\}→\\to\(a\)/\(b\),\\times/\\cdot→\\to\*,\\pi→\\topi\),float\(\), theneval\(\)on expression\-like strings, then prefix\-numeric extraction; \(iii\) symbolic equivalence viasympy\.simplify\(expr\_pred \- expr\_gold\) == 0for symbolic gold\. Released asreward\_physics\.py; selected by env varDENSE\_REWARD=0\(the default\), which returnsransr\_\{\\mathrm\{ans\}\}as a clean 0/1 binary\.

#### Why binary is the deployable default \(full theoretical analysis\)\.

The \(P1\)/\(P2\) intuitions are stated inline in §[4](https://arxiv.org/html/2605.14040#S4); we add the \(P3\) derivation here\. \(P1\) Group\-normalization rendersAkA\_\{k\}invariant to affine rescaling ofrrwithin a group, so dense shaping only matters when it*reorders*rollouts; we measure14\.3%14\.3\\%of within\-group pairs flipped by dense,87%87\\%inside the all\-wrong subgroup\. \(P2\) Those wrong\-group flips reward LaTeX\-format proxies satisfiable without solving the physics—a Goodhart channel that hurts prose\-and\-equation open\-ended evaluation; the matched\-step\-60 binary\-vs\-dense gap is largest on the auditedPhysOlym\-A\-liberal split\.

#### \(P3\) Binary reward maximizes per\-prompt advantage variance after the difficulty curriculum\.

The MM\-Eureka curriculum drops prompts where allKKrollouts are correct or all are wrong, so on every surviving prompt the rollout\-correct rate isp∈\(0,1\)p\\in\(0,1\)\. Binary reward is then a Bernoulli over rollouts:r¯=p\\bar\{r\}=p,σr2=p​\(1−p\)\\sigma\_\{r\}^\{2\}=p\(1\-p\), and the group\-normalized advantages take exactly two values

Akbin=\{\(1−p\)/prk=1−p/\(1−p\)rk=0,Var​\(Abin\)=1\.A^\{\\mathrm\{bin\}\}\_\{k\}\\;=\\;\\begin\{cases\}\\;\\sqrt\{\(1\-p\)/p\}&r\_\{k\}=1\\\\\[2\.0pt\] \-\\sqrt\{p/\(1\-p\)\}&r\_\{k\}=0\\end\{cases\},\\qquad\\mathrm\{Var\}\(A^\{\\mathrm\{bin\}\}\)\\;=\\;1\.\(3\)The advantage variance is exactly11, the maximum aKK\-sample group\-normalized estimator can carry on a Bernoulli reward\. Adding a bounded shaping termδk∈\[0,Δ\]\\delta\_\{k\}\\in\[0,\\Delta\]withΔ=0\.45\\Delta\{=\}0\.45\(therfmt\+rdim\+rsymr\_\{\\mathrm\{fmt\}\}\{\+\}r\_\{\\mathrm\{dim\}\}\{\+\}r\_\{\\mathrm\{sym\}\}budget of the dense reward\) inflates the within\-group standard deviationσr\\sigma\_\{r\}by an O\(Δ2\)\(\\Delta^\{2\}\)term while leaving the between\-correctness mean separation roughly unchanged, so the magnitude of the average correct\-vs\-wrong advantage*shrinks*:

\|Acorrectdense\|≈1−pp​\(1−p\)\+Δ2/12<1−pp​\(1−p\)=\|Acorrectbin\|\.\\big\|A^\{\\mathrm\{dense\}\}\_\{\\mathrm\{correct\}\}\\big\|\\;\\approx\\;\\frac\{1\-p\}\{\\sqrt\{p\(1\-p\)\+\\Delta^\{2\}/12\}\}\\;<\\;\\frac\{1\-p\}\{\\sqrt\{p\(1\-p\)\}\}\\;=\\;\\big\|A^\{\\mathrm\{bin\}\}\_\{\\mathrm\{correct\}\}\\big\|\.\(4\)Dense reward thus trades*between*\-correctness gradient signal\-to\-noise for*within*\-correctness rank information, but the within\-correctness flips are exactly the Goodhart channel of \(P2\)\.

### C\.2Dense five\-component physics\-native reward \(ablation\)

The dense five\-component Physics\-R1 reward \(Equation[2](https://arxiv.org/html/2605.14040#S4.E2), Section[4](https://arxiv.org/html/2605.14040#S4)\) is implemented asr=rans\+rfmt\+rdim\+rsym\+rconsr=r\_\{\\mathrm\{ans\}\}\+r\_\{\\mathrm\{fmt\}\}\+r\_\{\\mathrm\{dim\}\}\+r\_\{\\mathrm\{sym\}\}\+r\_\{\\mathrm\{cons\}\}, clipped to\[−1,1\]\[\-1,1\]\.

#### Per\-component implementation\.

Full code inreward\_physics\.py\.*rans∈\{0,\+1\}r\_\{\\mathrm\{ans\}\}\\in\\\{0,\+1\\\}:*brace\-counting\\boxedparser; MCQ\-letter equality; numeric/symbolic vialatex\_to\_plainnormalization \(\\frac,\\text,\\pi,\\times/\\cdot\) thenfloat/eval/prefix\-numeric; multi\-part requires all parts±1%\\pm 1\\%\.*rfmt∈\{0,\+0\.1\}r\_\{\\mathrm\{fmt\}\}\\in\\\{0,\+0\.1\\\}:*non\-empty\\boxed\{\}\.*rdim∈\{0,\+0\.15\}r\_\{\\mathrm\{dim\}\}\\in\\\{0,\+0\.15\\\}:*number\-prefix\-guarded regex extracts unit tokens, mapped tosympy\.physics\.units\(32 tokens\); fired only when every detected unit resolves\.*rsym∈\{0,\+0\.20\}r\_\{\\mathrm\{sym\}\}\\in\\\{0,\+0\.20\\\}:*first\-successsympy\.sympifyon LaTeX\-cleaned\\frac\{NUM\}\{DEN\}intermediates\.*rcons∈\{−0\.25,0\}r\_\{\\mathrm\{cons\}\}\\in\\\{\-0\.25,0\\\}:*negative\-only penalty when energy/momentum balance from corpus annotation is violated by\>5%\>5\\%relative\.

#### Mode and version pins\.

Released asreward\_physics\.py\. The mode is selected by env varDENSE\_REWARD:0\(default, recommended\) returnsransr\_\{\\mathrm\{ans\}\}only as 0/1 binary, recoveringrbinr\_\{\\mathrm\{bin\}\}of §[C\.1](https://arxiv.org/html/2605.14040#A3.SS1);1returns the full clipped dense sum \(ablation\)\. The audit factor \(envAUDIT\_LAMBDA\) zeros out reward on contaminated training items when set\. Version pins:sympy==1\.13\.3,transformers==4\.57\.0,vllm==0\.11\.0,verl==0\.6\.1\.

#### Worked example \(rollout group; binary vs\. dense advantages\)\.

A concrete realization of the \(P1\)/\(P2\) Goodhart channel of §[4](https://arxiv.org/html/2605.14040#S4)on one prompt withK=8K\{=\}8rollouts\. The prompt is a two\-step kinematics problem with gold19\.6\\boxed\{19\.6\}N \(problem idefo\-2014\-3a\); rolloutsy1,…,y4y\_\{1\},\\dots,y\_\{4\}commit to the correct answer with varying CoT quality,y5,…,y8y\_\{5\},\\dots,y\_\{8\}commit to wrong answers ranging from arithmetic\-slip \(19\.8\\boxed\{19\.8\},\>1%\{\>\}1\\%\) to off\-by\-physics \(42\.0\\boxed\{42\.0\}\)\.

kkFinal⋅\\boxed\{\\cdot\}CoT shaperansr\_\{\\mathrm\{ans\}\}rfmtr\_\{\\mathrm\{fmt\}\}rdimr\_\{\\mathrm\{dim\}\}rsymr\_\{\\mathrm\{sym\}\}rconsr\_\{\\mathrm\{cons\}\}rbinr\_\{\\mathrm\{bin\}\}rdenser\_\{\\mathrm\{dense\}\}119\.619\.6full units \+\\frac110\.100\.100\.150\.150\.200\.2001\.00\\mathbf\{1\.00\}1\.00\\mathbf\{1\.00\}219\.619\.6units only, no\\frac110\.100\.100\.150\.15001\.00\\mathbf\{1\.00\}1\.00\\mathbf\{1\.00\}319\.619\.6\\fraconly, no units110\.100\.1000\.200\.2001\.00\\mathbf\{1\.00\}1\.00\\mathbf\{1\.00\}419\.619\.6sparse CoT \(“answer is19\.619\.6”\)110\.100\.100001\.00\\mathbf\{1\.00\}1\.00\\mathbf\{1\.00\}542\.042\.0full units \+\\frac00\.100\.100\.150\.150\.200\.2000\.00\\mathbf\{0\.00\}0\.450\.45619\.819\.8units only, no\\frac00\.100\.100\.150\.15000\.00\\mathbf\{0\.00\}0\.250\.25742\.042\.0\\fraconly, no units00\.100\.1000\.200\.2000\.00\\mathbf\{0\.00\}0\.300\.308no⋅\\boxed\{\\cdot\}rambling, no commit000000\.00\\mathbf\{0\.00\}0\.000\.00
After clipping \(r≤1r\{\\leq\}1\) the four correct rollouts collapse to identical reward under both shapes\. After group\-normalization \(Eq\.[1](https://arxiv.org/html/2605.14040#S4.E1)\) the binary advantage vector isAbin=\(\+1,\+1,\+1,\+1,−1,−1,−1,−1\)A^\{\\mathrm\{bin\}\}\{=\}\(\+1,\+1,\+1,\+1,\-1,\-1,\-1,\-1\), every correct rollout receives equal positive gradient and every wrong rollout equal negative gradient\. The dense advantage vector isAdense≈\(\+1\.04,\+1\.04,\+1\.04,\+1\.04,−0\.55,−1\.00,−0\.93,−1\.69\)A^\{\\mathrm\{dense\}\}\\\!\\approx\\\!\(\+1\.04,\+1\.04,\+1\.04,\+1\.04,\-0\.55,\-1\.00,\-0\.93,\-1\.69\)\(computed exactly:r¯=0\.625\\bar\{r\}\{=\}0\.625,σr≈0\.428\\sigma\_\{r\}\{\\approx\}0\.428\)\.*Three observations land the \(P1\)–\(P3\) theory at the sample level:*

- •\(P1\) realized\. Among the four correct rollouts \(k=1,2,3,4k\{=\}1,2,3,4\), dense and binary assign*the same*advantage to every rollout—rank\-equivalent in the correct subgroup, even though rollouts 1–3 “deserve” more credit by the*a priori*physics\-native intuition\. Clipping at11erases the dense\-side variation among correct rollouts\.
- •\(P2\) realized\. Among the four wrong rollouts \(k=5,6,7,8k\{=\}5,6,7,8\), dense reorders them by LaTeX surface form:k=5k\{=\}5\(well\-formatted, units,\\frac,42\.0\\boxed\{42\.0\}way wrong\) gets the*smallest*negative advantage \(−0\.55\-0\.55\),k=8k\{=\}8\(no boxed commit\) gets the most negative \(−1\.69\-1\.69\)\. The policy gradient is therefore pushed toward producing well\-formatted wrong reasoning over poorly\-formatted wrong reasoning—the canonical Goodhart channel\. Notek=5k\{=\}5has dense advantage−0\.55\-0\.55vs\. binary−1\.00\-1\.00: the wrong\-answer gradient is*weakened*for the format\-compliant rollout, exactly the bias toward proxy satisfaction\.
- •\(P3\) realized\. The magnitude of the average correct\-rollout advantage is\|Acorrectbin\|=1\.0\|A^\{\\mathrm\{bin\}\}\_\{\\mathrm\{correct\}\}\|\{=\}1\.0vs\.\|Acorrectdense\|≈1\.04\|A^\{\\mathrm\{dense\}\}\_\{\\mathrm\{correct\}\}\|\{\\approx\}1\.04\(very close because clipping caps both atr=1r\{=\}1\)\. The magnitude of the average wrong\-rollout advantage is\|Awrongbin\|=1\.0\|A^\{\\mathrm\{bin\}\}\_\{\\mathrm\{wrong\}\}\|\{=\}1\.0vs\.\|Awrongdense\|≈1\.04\|A^\{\\mathrm\{dense\}\}\_\{\\mathrm\{wrong\}\}\|\{\\approx\}1\.04on the dense side as well, but the within\-wrong spread ofσ=0\.42\\sigma\{=\}0\.42across\{−0\.55,−0\.93,−1\.00,−1\.69\}\\\{\-0\.55,\-0\.93,\-1\.00,\-1\.69\\\}is the within\-group rank\-flipping variance that absorbs gradient capacity into the Goodhart direction\. Binary spends zero capacity on within\-correctness ranking and all of it on the correct\-vs\-wrong axis\.

The aggregate effect of running this calculus across∼\\sim1,024 prompts and 60 GRPO steps is the matched\-step\-60 binary\-vs\-dense gap of Table[3](https://arxiv.org/html/2605.14040#S5.T3)\(problem\-level liberal Sonnet\-judge accuracy across all open\-ended columns; bug\-corrected per §[5](https://arxiv.org/html/2605.14040#S5.SS0.SSS0.Px2)\): PhysReason32\.232\.2vs\.23\.323\.3\(\+8\.9\+8\.9pp\), PUB\-OE37\.037\.0vs\.37\.737\.7\(−0\.7\-0\.7pp, tied\), OlymBench\-Phys liberal45\.445\.4vs\.40\.540\.5\(\+4\.9\+4\.9pp\),PhysOlym\-Aliberal25\.625\.6vs\.19\.219\.2\(\+6\.4\+6\.4pp\)\.

#### Hyperparameter and framework details\.

The full GSPO\+DAPO configuration with all flags, lambdas, clip ranges, dynamic\-sampling settings, and the difficulty\-curriculum thresholds is in Table[10](https://arxiv.org/html/2605.14040#A8.T10)\(main text\) and the released YAML\. Theverl0\.6\.1 FSDP1 reproducibility note \(Section[6](https://arxiv.org/html/2605.14040#S6)\) and the upstream GitHub issue link are tracked in the released README\.

Algorithm 2Dense five\-component physics\-native reward for one rollout\.1:Solution string

ss, gold answer

gg, optional conservation flag

cons\\mathrm\{cons\}fromextra\_info

2:

r∈\[−1,1\]r\\in\[\-1,1\]
3:

ans←ExtractBoxed​\(s\)\\mathrm\{ans\}\\leftarrow\\textsc\{ExtractBoxed\}\(s\)⊳\\trianglerightbrace\-counting parser; nested\\boxedOK

4:

rans←\+1r\_\{\\mathrm\{ans\}\}\\leftarrow\+1if

Match​\(ans,g\)\\textsc\{Match\}\(\\mathrm\{ans\},g\)under MCQ\-letter / multi\-part /

±1%\\pm 1\\%numeric tolerance, else

0
5:

rfmt←\+0\.1r\_\{\\mathrm\{fmt\}\}\\leftarrow\+0\.1if

ans\\mathrm\{ans\}is non\-empty \(well\-formed\\boxed\{…\}\), else

0
6:

U←ExtractUnits​\(s\)U\\leftarrow\\textsc\{ExtractUnits\}\(s\)⊳\\trianglerightregex\(number\)\(unit\)\(exponent\)with number\-prefix guard

7:

rdim←\+0\.15r\_\{\\mathrm\{dim\}\}\\leftarrow\+0\.15if

\|U\|≥1\|U\|\\geq 1and

∀u∈U:ResolvesInSymPy​\(u\)\\forall u\\in U:\\textsc\{ResolvesInSymPy\}\(u\), else

0
8:

F←ExtractFracs​\(s\)F\\leftarrow\\textsc\{ExtractFracs\}\(s\)⊳\\trianglerightfind all\\frac\{NUM\}\{DEN\}

9:

rsym←\+0\.20r\_\{\\mathrm\{sym\}\}\\leftarrow\+0\.20if

∃\(NUM,DEN\)∈F:Sympifies​\(NUM\)∧Sympifies​\(DEN\)\\exists\(\\mathrm\{NUM\},\\mathrm\{DEN\}\)\\in F:\\textsc\{Sympifies\}\(\\mathrm\{NUM\}\)\\land\\textsc\{Sympifies\}\(\\mathrm\{DEN\}\), else

0
10:if

cons\\mathrm\{cons\}is provided \(

∼\\sim3,100\-record subset\)then

11:

p^←TryFloat​\(ans\)\\hat\{p\}\\leftarrow\\textsc\{TryFloat\}\(\\mathrm\{ans\}\);

p∗←∑cons\.in−∑cons\.out∖ansp^\{\*\}\\leftarrow\\sum\\\!\\mathrm\{cons\.in\}\-\\sum\\\!\\mathrm\{cons\.out\}\\setminus\\mathrm\{ans\}
12:

rcons←−0\.25r\_\{\\mathrm\{cons\}\}\\leftarrow\-0\.25if

\|p^−p∗\|/max⁡\(\|p∗\|,\|p^\|,ε\)\>0\.05\|\\hat\{p\}\-p^\{\*\}\|/\\max\(\|p^\{\*\}\|,\|\\hat\{p\}\|,\\varepsilon\)\>0\.05, else

0
13:else

14:

rcons←0r\_\{\\mathrm\{cons\}\}\\leftarrow 0⊳\\trianglerightnegative\-only; no positive reward for satisfaction

15:endif

16:return

clip​\(rans\+rfmt\+rdim\+rsym\+rcons,−1,1\)\\mathrm\{clip\}\(r\_\{\\mathrm\{ans\}\}\+r\_\{\\mathrm\{fmt\}\}\+r\_\{\\mathrm\{dim\}\}\+r\_\{\\mathrm\{sym\}\}\+r\_\{\\mathrm\{cons\}\},\-1,1\)

## Appendix DLLM\-judge details: three judges, scoring conventions, and reproducibility

This appendix documents the three Sonnet\-4\.5\-as\-judge variants used in Table[3](https://arxiv.org/html/2605.14040#S5.T3)and the scoring conventions adopted across columns\. Source code for all three judges is released in thejudge/directory of the code repository \([github\.com/shanyang\-me/physics\-r1\-neurips2026](https://arxiv.org/html/2605.14040v1/github.com/shanyang-me/physics-r1-neurips2026)\)\.

#### Why three judges\.

Open\-ended physics olympiad problems differ structurally: PhysReason and PhysUniBench\-OE problems are explicitly multi\-sub\-part \(often22–55sub\-questions per record, each with its own gold answer\), whilePhysOlym\-Aand OlympiadBench\-Physics problems are graded at the problem level \(a single gold solution document, with the model’s final answer compared against it\)\. A single judge prompt cannot serve both\. We therefore use three judges, each tuned to its eval’s structure:

- •judge\_olympiad\.py\(problem\-level, used forPhysOlym\-Aand OlympiadBench\-Physics\)\. One Sonnet 4\.5 call per problem; the prompt provides the full gold solution paragraph and a\\boxed\{\}\-extracted candidate answer \(with a600600\-char response\-tail fallback if no boxed is emitted\), and asks for a single YES/NO verdict on whether the candidate’s final answer is mathematically/physically equivalent to the gold’s\. Tolerance is2%2\\%relative\.
- •llm\_judge\_v2\_alignment\.py\(per\-subpart, AND across sub\-parts, used for PhysReason\)\. For each gold sub\-answergig\_\{i\}, a separate Sonnet 4\.5 call asks:*does ANY of the candidate’s predictions equalgig\_\{i\}?*Per\-subpart verdict is YES/NO\.judge\_problem\_correctis the AND across all sub\-parts\. Tolerance:1%1\\%\.
- •llm\_judge\_v3\_pubeo\.py\(per\-subpart, AND across sub\-parts, with cached clean gold \+ tail fallback, used for PhysUniBench\-OE\)\. Same per\-subpart structure as v2, but with a pre\-extracted*clean*per\-subpart gold list \(e\.g\.,\["2\.68nC", "7853\.1W"\]\) that bypasses PUB\-OE’s verbose paragraph\-form gold\. If the candidate’s\\boxed\{\}list is empty, a regex fallback scans the last300300chars of the response for likely numeric/symbolic answers\. Per\-subpart verdict is YES/NO;judge\_problem\_correct\_v3is the AND across all sub\-parts\. Tolerance:2%2\\%\.

#### Scoring conventions in Table[3](https://arxiv.org/html/2605.14040#S5.T3)\.

All open\-ended cells use*problem\-level*accuracy: for multi\-sub\-part problems \(PhysReason, PhysUniBench\-OE\) every sub\-part must be judged correct for the problem to count, via thejudge\_problem\_correctfield of the v2/v3 judges \(AND across sub\-parts\); for problem\-level evals \(PhysOlym\-A, OlympiadBench\-Physics\)judge\_olympiad\.pyreturns one YES/NO per problem\. We also computed a softer*per\-subpart*variant \(partial credit,∑i𝟙​\[subi​correct\]/∑i1\\sum\_\{i\}\\mathbb\{1\}\[\\text\{sub\}\_\{i\}\\text\{ correct\}\]/\\sum\_\{i\}1\) for the multi\-sub\-part columns and verified it is uniformly44–1717pp higher across rows; per\-subpart values are released alongside the dataset for users who want to study the partial\-credit lens but are not reported as headline in Table[3](https://arxiv.org/html/2605.14040#S5.T3)\.

#### Reproducibility note\.

All Sonnet\-judge runs in Table[3](https://arxiv.org/html/2605.14040#S5.T3)useworkers=22–44concurrency; errored sub\-judgments are filtered and re\-judged at lower concurrency rather than counted as wrong\. Per\-cell judge\-error counts \(typically≤1%\\leq 1\\%of records\) and per\-record verdicts are released in the supplementary archive atjudge\_audit\.json, so downstream auditors can verify each cell independently\.

#### Verbatim judge prompt \(problem\-level,PhysOlym\-A\+ OlymBench\)\.

The Sonnet 4\.5 judge is invoked with the following template:

> You are grading a physics olympiad answer\. GOLD \(full reference solution; the final numeric/symbolic answer is what matters\): \{gold\} CANDIDATE answers \(extracted from the model’s \\boxed\{\} markers\): \{preds\} \(If the candidate emitted no \\boxed\{\}, the candidate is the last 600 chars of its full response:\) \{tail\} Task: decide whether the candidate’s final answer is mathematically/physically equivalent to the gold’s final answer\. Allow: different but equivalent algebraic forms; trivial unit/format differences; rounding within 2% relative tolerance; trailing prose\. Reject: different magnitude, different sign, different functional form, missing or wrong physical content, no answer\. Respond with EXACTLY one word: YES or NO\.

#### Verbatim judge prompt \(per\-subpart, PhysReason \+ PhysUniBench\-OE\)\.

For each gold sub\-answergig\_\{i\}, the judge is invoked with:

> You are grading a physics olympiad answer\. GOLD answer: \{g\_i\} CANDIDATE predictions \(one or more, separated by ===\): \{preds\} Task: decide whether ANY of the candidate predictions is mathematically/physically equivalent to the gold answer\. Allow: different but equivalent algebraic forms; trivial unit/format differences \("450 N" == "450 \\text\{ N\}"\); rounding within 1\-\-2% relative tolerance; trailing prose \("approximately", "to the right"\); different variable names mapping cleanly\. Reject: different magnitude, different sign, different functional form, missing or wrong physical content\. Respond with EXACTLY one word: YES \(if any candidate matches\) or NO\. No other text\.

#### Per\-chunk verdict breakdown forPhysOlym\-A\(Sonnet 4\.5\)\.

The 500 problems are partitioned into 5 chunks of 100 \(random seed 42, source\-stratified\)\. Per\-chunk strict\-correct counts: 38, 20, 35, 30, 40 \(judgeable\-only denominators 98, 100, 99, 98, 96 after removing13\.9%13\.9\\%unjudgeables\)\. The chunk\-1 outlier at20%20\\%accuracy is concentrated in reference\-pointer Zhou problems whose gold solutions cite external olympiad\-handbook material rather than providing self\-contained answers; these are the unjudgeable category we surface as a known noise floor\.

#### Inter\-judge agreement\.

The Sonnet judge is run with two seeds on the same 500 problems forPhysOlym\-A\. Cohen’sκ\\kappaon the YES/NO label between the two passes is reported alongside the released dataset; preliminary inspection showsκ≥0\.8\\kappa\\geq 0\.8on thePhysOlym\-Acorpus\.

#### Human\-graded subset\.

A 100\-problem random subsample ofPhysOlym\-ASonnet predictions is graded by a single physics\-trained annotator using the same YES/NO rubric\. Per\-record human\-vs\-LLM agreement and Cohen’sκ\\kappaare released as a JSON alongside the dataset\. We report thisκ\\kappaas a calibration check on the LLM judge, not as a calibrated human\-baseline ground\-truth \(cf\. Section[6](https://arxiv.org/html/2605.14040#S6)\)\.

#### Released artifacts\.

The verbatim judge prompts \(problem\-level \+ per\-subpart\), the YES/NO rubric, the per\-chunk verdict breakdown, the 100\-problem human\-graded subset, and the per\-record agreement matrices are released asjudge\_prompts\.txt,rubric\.json,per\_chunk\_verdicts\.json, andhuman\_graded\_subset\.json\. Thepub\_oe\_gold\_cache\.json\(per\-id clean per\-subpart gold list\) is released alongsidellm\_judge\_v3\_pubeo\.pyand is reproducible from the original PhysUniBench\-OE source via the cache\-builder script in the same directory\.

#### Self\-grading concern\.

Sonnet 4\.5 is both the highest\-scoring frontier baseline onPhysOlym\-Aand the judge used for liberal accuracy\. Three checks bound any self\-favoring bias: \(i\) the strict \(numeric/symbolic match, judge\-independent\) Sonnet score is reported alongside liberal—the4\.74\.7\-pp Sonnet strict\-vs\-liberal gap \(28\.7%28\.7\\%vs\.33\.4%33\.4\\%\) bounds maximum leniency; \(ii\) on the failure\-mode taxonomy \(Appendix[H\.1](https://arxiv.org/html/2605.14040#A8.SS1)\) the dominant Sonnet error category iswrong\_subpart\(structural mismatch any judge marks wrong\), notvalid\_partial; \(iii\) we report a cross\-vendor judge agreement against GPT\-4o on a5050\-problem random subsample ofPhysOlym\-A\(Qwen3\-VL\-8B\-Thinking responses, seed4242\)\. Sonnet 4\.5 and GPT\-4o under an identical prompt template show88%88\\%raw agreement\(44/5044/50\) andCohen’sκ=0\.44\\kappa\{=\}0\.44\(moderate, per Landis & Koch\)\. The disagreement is asymmetric: GPT\-4o flips55Sonnet\-NO records to YES, while only11Sonnet\-YES record is flipped by GPT\-4o to NO \(McNemar exact on66discordant pairs\{5:Sonnet\-NO→GPT\-YES,1:Sonnet\-YES→GPT\-NO\}\\\{5\{:\}\\text\{Sonnet\-NO\}\\\!\\to\\\!\\text\{GPT\-YES\},\\,1\{:\}\\text\{Sonnet\-YES\}\\\!\\to\\\!\\text\{GPT\-NO\}\\\}, two\-sidedp=0\.219p\{=\}0\.219; the asymmetry is not significant atn=6n\{=\}6, but the direction is preserved on a larger200200\-problem ablation left to follow\-up work\)\. GPT\-4o’s positive rate \(16%16\\%,8/508/50\) is roughly*twice*Sonnet’s \(8%8\\%,4/504/50\), which means the cross\-vendor judge would assign Physics\-R1*higher*numbers than the Sonnet\-judge headlines, not lower—the self\-grading direction is the opposite of what the self\-favoring concern would predict\. The full per\-pair verdicts are released ascross\_judge\_50\.jsonlalongside the dataset\.

## Appendix EPer\-source license and provenance log

Per\-source provenance for each of the nine source families: full name, scrape URL, scrape date, original license string, redistribution tier, and outreach log\.

#### UGPhysics\.

Source:Xuet al\.\[[2025](https://arxiv.org/html/2605.14040#bib.bib3)\]\(ICML 2025\)\. 5,520 EN/ZH undergraduate physics problems\. Scrape URL:[https://huggingface\.co/datasets/UGPhysics/ugphysics\-bench](https://huggingface.co/datasets/UGPhysics/ugphysics-bench)\. Scrape date: 2026\-03\-15\. Original license: CC BY\-NC\-SA 4\.0\. Redistribution: CC BY\-NC\-SA 4\.0 carried through\. Outreach: not contacted \(license is permissive for academic redistribution under same\-license sharing\)\.

#### OpenStax College \+ University Physics\.

#### Physics Stack Exchange\.

Source: Physics Stack Exchange Q&A archive\. 2,291 problem\-and\-accepted\-answer records filtered for olympiad\-style physics\. Scrape URL:[https://physics\.stackexchange\.com/](https://physics.stackexchange.com/)via Stack Exchange data dump\. Scrape date: 2026\-01\-08\. Original license: CC BY\-SA 4\.0 \(Stack Exchange contributor agreement\)\. Redistribution: CC BY\-SA 4\.0 carried through\.

#### MMMU \+ o1\-CoT seed \(RL\-SFT seed\)\.

Source: 1,293\-record pool of MMMU physics problems augmented with o1\-style CoT solutions generated by Sonnet 4\.5\. MMMU base license: MIT\[Yueet al\.,[2024a](https://arxiv.org/html/2605.14040#bib.bib5)\]; generated CoT is our contribution\. Scrape date: 2026\-02\-10\. Redistribution: MIT \(MMMU base\) with generated CoT released under CC BY 4\.0\.

#### PhysReason\.

Source:Zhanget al\.\[[2025](https://arxiv.org/html/2605.14040#bib.bib4)\]\(ACL 2025\)\. 1,200 step\-graded physics reasoning problems\. Scrape URL:[https://huggingface\.co/datasets/PhysReason](https://huggingface.co/datasets/PhysReason)\. Scrape date: 2026\-02\-22\. Original license: CC BY 4\.0\. Redistribution: CC BY 4\.0 carried through\.

#### Estonian Physics Olympiad collection\.

Source: Estonian Physics Olympiad \(EFO\), 2004–2018\. 418 problems with organizer\-issued 1–10 difficulty labels and a 201\-problem bilingual EN\+ET subset\. Scrape URL:[https://fyysika\.ee/](https://fyysika.ee/)\(rounds: lahtine, koolivoor, vabariikilik\)\. Scrape date: 2026\-03\-30\. Provenance: publicly archived problems and solutions on the official EFO portal, released by the Estonian Physics Olympiad committee for educational use under competition policy \(consistent with international physics\-olympiad practice for IPhO, NBPhO, EuPhO, APhO, USAPhO, INPhO\)\. Redistribution: public\-domain by competition policy, used for non\-commercial research evaluation; downstream users redistributing for commercial purposes or in materially modified form should consult[https://fyysika\.ee/](https://fyysika.ee/)directly\.

#### Kevin Zhou’s olympiad handouts\.

Source: Kevin Zhou’s olympiad training documents \(USAPhO/IPhO/APhO/EuPhO content with Cambridge Tripos and graduate\-qualifier extensions\),[https://knzhou\.github\.io/](https://knzhou.github.io/)\. 692 problems with native point values 1–5 and a 3\.2% advanced \[A\] flag\. Scrape date: 2026\-02\-01\. Original license: CC BY\-NC 4\.0 \(written confirmation from Kevin Zhou \(kzhou7@gmail\.com\), Date headerSun, 3 May 2026 17:53:30 \+0800, archived in supplementary aszhou\_license\_2026\-05\-02\.eml, SHA\-2567f25c859f5ae1d790e45dbdfd23ab6be27aa1814a76de10f9bdbffb67088aba4\)\. Redistribution: the692692redistributed problems plus their reference solutions \(∼\\sim1,600 problem\-solution items in total counting solutions as separate documents\) are released under CC BY\-NC 4\.0 with attribution link to[https://knzhou\.github\.io/](https://knzhou.github.io/)preserved per record\. Third\-party content disclosure \(per Zhou’s reply\): some problems in the handouts are drawn from books or other olympiad archives, with the original source attributed inline by Zhou\. We preserve all such per\-problem internal attributions verbatim in the released records; downstream users should treat any in\-record secondary\-source attribution as binding under the original source’s terms, which may be more restrictive than CC BY\-NC 4\.0\. Outreach log: initial inquiry 2026\-04\-10 \(proposal:∼\\sim1,600 problems \+ solutions, CC BY\-NC 4\.0, attribution to source site\); written agreement to the proposed terms 2026\-05\-03 with the third\-party\-content caveat noted by Zhou\.

#### IPhO \+ NBPhO \+ EuPhO scrape\.

Sources: International Physics Olympiad \([ipho\-new\.org](https://arxiv.org/html/2605.14040v1/ipho-new.org)\), Nordic Baltic Physics Olympiad, European Physics Olympiad official archives\. 258 problems \(IPhO 66, NBPhO 165, EuPhO 27\)\. Scrape date: 2026\-03\-05\. Original license: public\-domain by competition policy \(problems published for educational use without restriction; we preserve attribution per record\)\. Redistribution: public\-domain carried through with per\-record attribution\.

#### APhO \+ USAPhO \+ INPhO scrape\.

Sources: Asian Physics Olympiad, USA Physics Olympiad \(Physics Olympiad Foundation\), Indian National Physics Olympiad\. 241 problems recovered after a multi\-pattern splitter fix that recognizesA1\.\-style markers \(USAPhO 1997–2006\) and1\. \(a\)solution markers \(INPhO\)\. The INPhO recovery alone added \+33 records relative to a single\-regex baseline\. Scrape date: 2026\-03\-12\. Original license: public\-domain by competition policy\. Redistribution: public\-domain carried through with per\-record attribution\.

Table 5:Per\-source license and provenance\. Each released record carries its source license through; non\-commercial sources \(UGPhysics, Zhou\) restrict downstream use to academic research, and the public\-domain\-by\-competition\-policy olympiad scrapes \(EFO, IPhO, NBPhO, EuPhO, APhO, USAPhO, INPhO\) preserve per\-record source attribution\.Source familyRecordsOriginal licenseOpenStax College \+ University Physics2,381CC BY 4\.0Physics Stack Exchange2,291CC BY\-SA 4\.0PhysReason1,200CC BY 4\.0MMMU \+ o1\-CoT seed \(RL\-SFT seed\)1,293MIT \(MMMU\); generated CoTIPhO \+ NBPhO \+ EuPhO scrape258Public\-domain \(competition policy\)APhO \+ USAPhO \+ INPhO scrape241Public\-domain \(competition policy\)UGPhysics5,520CC BY\-NC\-SA 4\.0Estonian Physics Olympiad418Public\-domain \(competition policy\)Kevin Zhou’s olympiad handouts692CC BY\-NC 4\.0 \(Zhou confirmed 2026\-05\-03\)Total pre\-audit14,294

## Appendix FReproducibility checklist

#### Compute budget per experiment\.

ExperimentWall\-clockHardwareCostNotesSonnet baselines \(PhyX / OlymBench /PhysOlym\-A\)∼5\{\\sim\}5hAPI∼$​80\\sim\\mathdollar 80Two\-judgeOpen\-source baseline sweep on PhyX\-mini\-MC 1000q∼90\{\\sim\}90min1×\{\\times\}H200∼$​5\{\\sim\}\\mathdollar 5vLLM 0\.11\.0 bf16Two\-stage contamination audit∼3\{\\sim\}3minMPS/CPUfree5 k records pairwiseThreshold\-sensitivity analysis∼3\{\\sim\}3minMPSfreesentence\-transformersPhysics\-R1 RL training to step 60\+∼30\{\\sim\}30h4×\{\\times\}H200∼$​120\{\\sim\}\\mathdollar 120verl 0\.6\.1, FSDP1Reward\-component drop\-out \(follow\-up work\)∼40\{\\sim\}40h4×\{\\times\}H200∼$​700\{\\sim\}\\mathdollar 700Table[11](https://arxiv.org/html/2605.14040#A8.T11)3\-seed Physics\-R1 sensitivity \(seeds 17 \+ 23 retrainings\)∼60\{\\sim\}60h4×\{\\times\}H200∼$​200\{\\sim\}\\mathdollar 200added to seed\-42 in Table[3](https://arxiv.org/html/2605.14040#S5.T3)Total \(this paper\)∼$​700\\sim\\mathdollar 700seeds 42/17/23 \+ Sonnet baselines \+ Sonnet\-judge runsFollow\-up budget \(reward drop\-out\)∼$​1,000\\sim\\mathdollar 1\{,\}000Table[11](https://arxiv.org/html/2605.14040#A8.T11)
#### Random seeds\.

The headline single\-seed Physics\-R1 binary checkpoint uses seed4242\. The 3\-seed mean reported in Table[3](https://arxiv.org/html/2605.14040#S5.T3)aggregates seeds\{42,17,23\}\\\{42,17,23\\\}, all on the auditedPhysR1Corpcorpus under the binary correctness reward; checkpoint selection per seed uses MM\-Eureka difficulty\-curriculum saturation on the held\-out PhyX\-mini\-MC \(1,0001\{,\}000\-problem\) early\-stop signal\. The data\-build pipeline \(PhysOlym\-Asampling, train/val splits, audit pass\) usesnumpy\.random\.default\_rng\(42\)\.

#### Version pins\.

transformers==4\.57\.0\(re\-evaluation with4\.57\.6drifts results by∼1\.5\{\\sim\}1\.5points; we pin to4\.57\.0for reproducibility\),vllm==0\.11\.0,verl==0\.6\.1,sympy==1\.13\.3,sentence\-transformers==5\.4\.1,torch==2\.8\.0\+cu128\. Embedding model:mxbai\-embed\-large\-v1from MixedBread AI\.

#### Code and data hosting\.

#### Quick\-start audit\.

```
python audit/audit_two_stage.py \
    --train_jsonl your_pool.jsonl --eval_jsonl data/physolym_a.jsonl \
    --jaccard_thr 0.4 --cosine_thr 0.85 --emit report.json
```

Writes per\-record audit \+ aggregate3×33\{\\times\}3threshold\-sensitivity table \(Table[4](https://arxiv.org/html/2605.14040#A1.T4)\);∼3\{\\sim\}3min on MPS or a CUDA GPU for a5,0005\{,\}000\-record pool against44held\-out splits, with shingle/embedding caches inaudit\_cache/\.

#### Supplementary materials index\.

\(i\) Croissant 1\.0 \+ RAI JSON\-LD per artifact \(croissant\_rai\_<artifact\>\.json\); \(ii\) the four released datasets \(physcorp\_a\.jsonl,physr1corp\.jsonl,physolym\_a\.jsonl, plusPhysCorp\-pre\-audit\); \(iii\) audit pipeline \(audit\_two\_stage\.py\+threshold\_sensitivity\_scores\.npz\); \(iv\) reward implementation \(reward\_physics\.py\); \(v\) LLM\-judge artifacts \(judge\_prompt\.txt,rubric\.json,per\_chunk\_verdicts\.json,human\_graded\_subset\.json\); \(vi\) archived license confirmation \(zhou\_license\_2026\-05\-02\.eml, the Kevin Zhou CC BY\-NC 4\.0 grant\); \(vii\) training config \(configs/physics\-r1\.yaml\); \(viii\) checkpoints at steps\{20,40,60,80\}\\\{20,40,60,80\\\}\.

#### Reproducibility checklist\.

Code released withrequirements\.txtand deterministic build script; data released with per\-source license provenance \(Table[5](https://arxiv.org/html/2605.14040#A5.T5)\) and Croissant 1\.0\+RAI metadata; datasheet in Appendix[G](https://arxiv.org/html/2605.14040#A7); compute, seeds, and hyperparameters above \+ Table[10](https://arxiv.org/html/2605.14040#A8.T10)\. Hosting: HuggingFace \(≥\\geq5 yr\) \+ GitHub \+ Zenodo\. Maintenance: quarterly contamination audit\. Follow\-up commitments: reward\-component drop\-out ablation \(Table[11](https://arxiv.org/html/2605.14040#A8.T11)\), embedder\-sensitivity audit againstvoyage\-3andtext\-embedding\-3\-large, paraphrase/translation\-aware audit pass\.

#### Recommended baseline configuration\.

Algorithm[3](https://arxiv.org/html/2605.14040#alg3)captures the joint setting; each individual choice is small in magnitude\. Binary correctness is the recommended default; dense \(Algorithm[2](https://arxiv.org/html/2605.14040#alg2)\) is reported as an ablation\.

Algorithm 3Recommended baseline configuration for physics\-VL RL post\-training\.1:Base thinking\-mode VLM

πbase\\pi\_\{\\mathrm\{base\}\}, audited training pool

T′T^\{\\prime\}\(Algorithm[1](https://arxiv.org/html/2605.14040#alg1)\), held\-out MCQ early\-stop signal

HMCH\_\{\\mathrm\{MC\}\}, dense reward

rr\(Algorithm[2](https://arxiv.org/html/2605.14040#alg2)\)

2:Init:

πθ←πbase\\pi\_\{\\theta\}\\leftarrow\\pi\_\{\\mathrm\{base\}\}⊳\\trianglerightcold\-start from base, no SFT pass \(MM\-Eureka thesis\)

3:Optimizer: GSPO\+DAPO \(sequence\-level importance, decoupled clip\), unmodified

4:KL anchor: add

βKL​DKL​\(πθ∥πbase\)\\beta\_\{\\mathrm\{KL\}\}\\,D\_\{\\mathrm\{KL\}\}\(\\pi\_\{\\theta\}\\\|\\pi\_\{\\mathrm\{base\}\}\)with

βKL=10−3\\beta\_\{\\mathrm\{KL\}\}=10^\{\-3\}⊳\\trianglerightbounds drift

5:Entropy bonus: add

βH​ℋ​\(πθ\)\\beta\_\{H\}\\,\\mathcal\{H\}\(\\pi\_\{\\theta\}\)with

βH=10−3\\beta\_\{H\}=10^\{\-3\}⊳\\trianglerightprevents entropy collapse

6:Difficulty curriculum: drop train items where the base model gets

0/N0/Nor

N/NN/Nrollouts⊳\\trianglerightMM\-Eureka\-style; preserves learning signal

7:LR schedule:

1×10−61\{\\times\}10^\{\-6\}initial, step\-cosine decay; halve LR after the first plateau on

HMCH\_\{\\mathrm\{MC\}\}⊳\\trianglerightcounters drift and length collapse

8:Response budget: fixed long\-CoT budget \(12,288 tokens for thinking\-mode\),*not adaptive*⊳\\trianglerightstable per\-step cost

9:Early stopping: stop at the saturation peak on held\-out MCQ

HMCH\_\{\\mathrm\{MC\}\},*not*on the train\-distribution validation set⊳\\trianglerightaddresses train\-up/eval\-down divergence

10:Reward:*recommended*binary correctness reward

rans∈\{0,\+1\}r\_\{\\mathrm\{ans\}\}\\in\\\{0,\+1\\\}\(simpler, fully reproducible onvllm== 0\.11\.0 multi\-image eval\); dense five\-component physics\-native reward

rr\(Algorithm[2](https://arxiv.org/html/2605.14040#alg2)\) is reported as an ablation

11:Audit gate: train pool must be Algorithm[1](https://arxiv.org/html/2605.14040#alg1)\-audited against held\-out splits*and*external benchmarks before training begins

12:Reproducibility pins:transformers== 4\.57\.0,vllm== 0\.11\.0,verl== 0\.6\.1, FSDP1 sharding for Qwen3\-VL \(§[6](https://arxiv.org/html/2605.14040#S6)\)

## Appendix GDatasheet for Physics\-R1

We followGebruet al\.\[[2021](https://arxiv.org/html/2605.14040#bib.bib21)\]with seven sections: Motivation, Composition, Collection process, Preprocessing/cleaning/labeling, Uses, Distribution, and Maintenance\. The full per\-source provenance log is in Appendix[E](https://arxiv.org/html/2605.14040#A5); the audit pipeline in Appendix[A](https://arxiv.org/html/2605.14040#A1); the LLM\-judge protocol in Appendix[D](https://arxiv.org/html/2605.14040#A4)\.

### G\.1Motivation

For what purpose was the dataset created? To support contamination\-audited evaluation and post\-training of multimodal vision\-language models on visual physics reasoning, with three specific gaps the field had not closed: \(i\) no public physics\-VL training pool was audited under a three\-stage \(n\-gram, embedding, LLM\-judge\) protocol that catches paraphrase\-class duplicates and recovers threshold\-edge topic\-similarity false positives; \(ii\) no public physics\-olympiad eval was both novel\-source and contamination\-clean against the major training\-side aggregations \(PhyX, MMMU\-Pro Physics, OlympiadBench\-Physics, UGPhysics\); \(iii\) no public physics\-VL benchmark exposed the format\-and\-novelty saturation gradient at frontier\-model scale\. Who created the dataset and on whose behalf? Shan Yang\. Who funded the creation? Self\-funded; no third\-party sponsor\. Other comments\. The corpus aggregates and audits material that already existed in scattered formats; the Estonian Physics Olympiad, Kevin Zhou’s olympiad handouts, and seven international olympiads are first\-time ML\-format releases\.

### G\.2Composition

*Instances:*one physics problem per record \(statement, optional images PNG/JPEG, gold MCQ\-letter / numeric / symbolic / multi\-part answer, optional reference solution, 14\-field annotation; schema in §[B](https://arxiv.org/html/2605.14040#A2)\)\.*Counts:*PhysCorp\-A6,432 audited;PhysR1Corp2,268 closed\-form RL pool;PhysOlym\-A500 held\-out \(stratified sample, seed 42, source\-family\-stratified\);PhysCorp\-pre\-audit14,294 raw\.*Labels:*gold answer \+ schema labels;∼\\sim3,900 records carry full Sonnet\-4\.5 annotation, rest from source\-native labels;native\_difficultypresent where organizers publish \(Estonian 27%, Zhou 38%\)\.*Splits:*see §[B](https://arxiv.org/html/2605.14040#A2)\.*Noise:*LLM\-judge unjudgeable rate13\.9%13\.9\\%onPhysOlym\-A; Stage\-1 audit misses paraphrase/translation, Stage\-2 misses numerical substitution—both reported as floors\.*Self\-contained:*yes for problems and solutions; some Zhou records carry inline secondary\-source attribution \(Appendix[E](https://arxiv.org/html/2605.14040#A5)\)\.*No PII or sensitive content\.*

### G\.3Collection process

*Acquisition:*5 repackaged benchmark releases \(UGPhysics, OpenStax, Physics Stack Exchange, MMMU\+o1\-CoT, PhysReason\) \+ 4 first\-ML scrapes by authors \(Estonian PhO, Kevin Zhou’s handouts, 7 international olympiads\); per\-source URLs and dates in Appendix[E](https://arxiv.org/html/2605.14040#A5)\.*Sampling:*per\-source complete enumeration over public archive date ranges\.*Authorship:*the author \(Shan Yang\) handled scraping/parsing/audit; Sonnet 4\.5 batch produced annotation labels\.*Time frame:*scrapes 2025\-12 to 2026\-04, audit/curation 2026\-02 to 2026\-04\.*Ethics:*no human subjects, no PII, no IRB\.*Consent:*Kevin Zhou confirmed CC BY\-NC 4\.0 redistribution of his olympiad handouts in writing 2026\-05\-03; the Estonian Physics Olympiad and other olympiad sources \(IPhO, NBPhO, EuPhO, APhO, USAPhO, INPhO\) are released under public\-domain competition policy with per\-record source attribution; repackaged sources are redistributed under their original CC/MIT licenses with attribution preserved\.

### G\.4Preprocessing, cleaning, labeling

Pre\-tokenization normalization for Stage\-1 audit \(Appendix[A](https://arxiv.org/html/2605.14040#A1)\); three\-stage audit \(Stage\-1J≥0\.4J\\geq 0\.4, Stage\-2cos≥0\.85\\cos\\geq 0\.85, Stage\-3 Haiku\-4\.5 LLM\-judge close\-duplicate vs\. same\-topic\-neighbor classification\) pairwise across PhyX, MMMU\-Pro Physics, OlympiadBench\-Physics, UGPhysics\-Train,PhysOlym\-A, with Stage\-3 close\-duplicate records removed; 14\-field schema annotation by Sonnet 4\.5 batch \(3,900 records\) \+ source\-native labels\. Raw pre\-audit pool \(PhysCorp\-pre\-audit, 14,294 records\) released alongside so users can re\-run audits\. Audit pipeline released asaudit\_three\_stage\.pywith savedbest\_jaccard/best\_cosine/judge\_labelarrays\.

### G\.5Uses

*Used for:*Physics\-R1 RL recipe \(§[4](https://arxiv.org/html/2605.14040#S4)\) trains on the audited pool and evaluates onPhysOlym\-A\+ auxiliary splits\. Supplementary archive ships paper, datasets, Croissant\+RAI metadata, \.eml license confirmations, checkpoints\.*Other use cases:*visual physics reasoning eval, contamination\-audit methodology, cross\-lingual studies \(EN/ET subset\), native\-difficulty calibration, VLM RL post\-training\.*Caveats:*cross\-lingual finding is Sonnet\-specific; dense reward is an ablation not a tuned standard\.*Out of scope:*general physics ability beyond visual reasoning, human olympiad grading substitution, experimental physics, research capability; held\-out splits must not enter pretraining/fine\-tuning\.

### G\.6Distribution

*Distribution:*HuggingFace \(≥5\\geq 5yr\) \+ GitHub \+ Zenodo DOI\. URLs:[huggingface\.co/datasets/shanyangmie/physolym\-a](https://arxiv.org/html/2605.14040v1/huggingface.co/datasets/shanyangmie/physolym-a)and[huggingface\.co/datasets/shanyangmie/physics\-r1\-corpus](https://arxiv.org/html/2605.14040v1/huggingface.co/datasets/shanyangmie/physics-r1-corpus); code at[github\.com/shanyang\-me/physics\-r1\-neurips2026](https://arxiv.org/html/2605.14040v1/github.com/shanyang-me/physics-r1-neurips2026)\. All four artifacts released alongside this paper with per\-source licenses documented in Appendix[E](https://arxiv.org/html/2605.14040#A5); the Kevin Zhou CC BY\-NC 4\.0 grant is preserved as a written \.eml in the supplementary archive, and the remaining olympiad sources are released under public\-domain competition policy\.*Licenses:*CC BY 4\.0, CC BY\-SA 4\.0, public\-domain by competition policy \(EFO, IPhO, NBPhO, EuPhO, APhO, USAPhO, INPhO\), MIT, CC BY\-NC 4\.0 \(Zhou\), CC BY\-NC\-SA 4\.0 \(UGPhysics\)—each record carries its source license; Table[5](https://arxiv.org/html/2605.14040#A5.T5)\. CC BY\-NC sources are non\-commercial; Zhou records honor inline secondary\-source attribution\. No export controls\.

### G\.7Maintenance

Maintained by the author \(Shan Yang,alexyangshan@gmail\.com\); contact via email or GitHub Issues at[github\.com/shanyang\-me/physics\-r1\-neurips2026](https://arxiv.org/html/2605.14040v1/github.com/shanyang-me/physics-r1-neurips2026)\. Versioned CHANGELOG \(v1\.0\.0initial release,v1\.1\.xadditive,v2\.0\.0schema\-breaking\); per\-record diffs per release\. Quarterly contamination audit against new physics\-VL benchmarks;≥1%\\geq 1\\%leakage triggers documented\-diff removal\. Planned follow\-ups: paraphrase/translation\-aware audit, embedder\-sensitivity ablation againstvoyage\-3/text\-embedding\-3\-large\. Hosting≥5\\geq 5yr on HF \+ Zenodo DOI per release; all versions tagged and accessible\. Contributions via GitHub Issues / PR \+ per\-recorderratum\.json, reviewed within 60 days\.

#### Machine\-readable metadata: Croissant \+ RAI JSON\-LD\.

The release ships a Croissant 1\.0\[Akhtaret al\.,[2024](https://arxiv.org/html/2605.14040#bib.bib35)\]JSON\-LD descriptor \(croissant\.json\) declaring distribution objects forPhysR1Corp,PhysOlym\-A, and the audit\-pipeline source archive, plus a 14\-fieldproblem\_recordschema and the full RAI extension \(rai:dataCollection,rai:dataAnnotationProtocol,rai:dataPreprocessingProtocol,rai:personalSensitiveInformation,rai:dataLimitations,rai:dataReleaseMaintenancePlan,rai:dataUseCases,rai:dataBiases,rai:dataSocialImpact\) covering the audit methodology, the Kevin Zhou CC BY\-NC 4\.0 written grant \(Appendix[E](https://arxiv.org/html/2605.14040#A5)\), the public\-domain\-by\-competition\-policy basis for the olympiad scrapes, three documented distributional biases \(EN/ET asymmetry, per\-physics\-category variance, difficulty\-stratified decay\), and the Stage\-1/Stage\-2 thresholds\. Passesmlcroissant validateand the MLCommons RAI checker\.

## Appendix HExtended discussion: failure modes, ethics, methodological notes, and future work

This appendix collects extended\-discussion content that was trimmed from the main body to fit the page limit\. Each subsection below corresponds to a one\-line pointer in the analysis section \(§[6](https://arxiv.org/html/2605.14040#S6)\)\.

### H\.1Failure\-mode taxonomy \(extended\)

A manual taxonomy of 100 randomly\-sampled wrong or partial Sonnet predictions on OlympiadBench\-Physics was completed post\-evaluation\. Methodology\. The 100 cases were drawn \(random seed 42\) from the 354 total wrong predictions on OlympiadBench\-Physics; each case was categorized by a single physics\-trained annotator \(graduate\-level physics background,∼\\sim3 hours of total annotation time,∼\\sim1\.8 min/case median\) using a nine\-category mutually\-exclusive rubric \(wrong\_subpart,missing\_physics,calc\_error,different\_question,valid\_partial,symbolic\_vs\_numeric,magnitude\_error,sign\_error,diagram\_misread\)\. The annotator saw the problem statement, the gold solution, the gold final answer, and the Sonnet response; they did not see Sonnet’s confidence or the verdict from the LLM judge\. We report this single\-annotator taxonomy as a methodological diagnostic,*not*as a calibrated human grade; a second\-annotator pass on a 50\-problem random subsample with inter\-annotatorκ\\kappaper category is left to follow\-up work \(distinct from the audit\-flag inter\-annotatorκ\\kappapre\-registered in Appendix[A](https://arxiv.org/html/2605.14040#A1), which targets contamination\-flag agreement, not failure\-category agreement\)\.

Table 6:Failure\-mode taxonomy of 100 randomly\-sampled Sonnet wrong/partial predictions on OlympiadBench\-Physics\. Categories are mutually exclusive; each prediction received one label\.Categoryn%Descriptionwrong\_subpart3030%Answered a neighboring sub\-question instead of the one askedmissing\_physics2222%Wrong physical model, law, or geometry; crucial effect absentcalc\_error1616%Correct approach; small numerical slip \(factor 2–4 off\)different\_question1010%Reasoning chain for an unrelated problem in the same promptvalid\_partial88%Right approach with arithmetic slip; partial\-credit territorysymbolic\_vs\_numeric77%Formula returned where a number was required, or vice versamagnitude\_error55%Correct expression, wrong order of magnitude \(≥10×\\geq 10\\times\)sign\_error00%None observeddiagram\_misread00%None observed \(text\-only evaluation; figures unavailable\)Total100100%#### Reading the taxonomy\.

wrong\_subpartdominates \(30%\): multi\-part olympiad prompts containing \(a\)/\(b\)/\(c\) sub\-questions where the model fluently answers a neighboring sub\-question instead of the asked one—a 5×\\timesunderestimate by regex\-only flagging \(∼\\sim6% in §[5](https://arxiv.org/html/2605.14040#S5.SS0.SSS0.Px1)\)\.missing\_physics\(22%\) anddifferent\_question\(10%\) together account for nearly a third of failures and represent a real reasoning floor unlikely to be format\-induced\.calc\_error\(16%\) andvalid\_partial\(8%\) are correct\-approach/failed\-execution, matching the 4\.7\-pp strict\-vs\-liberal gap \(§[5\.1](https://arxiv.org/html/2605.14040#S5.SS1.SSS0.Px1)\)\.sign\_erroranddiagram\_misreadare zero \(the latter because OlympiadBench\-Physics is text\-only in the public release\)\.

#### Per\-physics\-category breakdown\.

Table 7:Per\-physics\-category accuracy on OlympiadBench\-Physics \(Sonnet 4\.5, strict\)\. The\+34\.5\+34\.5\-pp EM \(38\.4%38\.4\\%\) vs\. astrophysics \(72\.9%72\.9\\%\) gap on identical weights motivates per\-category reporting\.CategorynAcc\.CategorynAcc\.CategorynAcc\.Electromagnetism23738\.4%Relativity8951\.7%Other1060\.0%Quantum12841\.4%Classical mech\.1957\.9%Astrophysics7072\.9%Waves3946\.2%Thermodynamics8759\.8%All692—Fluid mechanics1346\.2%

### H\.2Ethical considerations and intended use

#### Intended and out\-of\-scope use\.

The released artifacts support research on visual physics reasoning, contamination\-audit methodology, cross\-lingual LLM evaluation, and RL post\-training of VLMs\.PhysOlym\-Ais held\-out: please treat it as test\-only, not for pretraining or fine\-tuning\. Out of scope: general physics ability beyond visual reasoning, human olympiad grading substitution, experimental\-physics evaluation, paper\-writing, derivation novelty, or open\-ended hypothesis generation\.

#### Source provenance, consent, and misuse risk\.

Kevin Zhou confirmed CC BY\-NC 4\.0 redistribution of his olympiad handouts in writing on 2026\-05\-03 \(∼\\sim1,600 problem\-solution items, Appendix[E](https://arxiv.org/html/2605.14040#A5)\); the Estonian Physics Olympiad collection and the six international olympiad scrapes \(IPhO, NBPhO, EuPhO, APhO, USAPhO, INPhO\) are redistributed under public\-domain competition policy with per\-record source attribution\. No PII; problems are olympiad/textbook content\. Public\-domain olympiad scrapes are released by competition policy\. The dense reward and Physics\-R1 recipe ship for reproducibility, not as a tuned gold standard; the audit pipeline is a measurement tool, not a certification authority, and we disclose threshold sensitivity \(Table[4](https://arxiv.org/html/2605.14040#A1.T4)\) so users do not over\-claim “contamination\-free\.”

#### Caveats and compute\.

The 17\-pp EN/ET cross\-lingual delta is Sonnet\-4\.5\-specific onn=59n\{=\}59paired items; direction may flip on low\-resource\-language\-weak open\-source models—treat as model\-specific\. Total compute∼\\sim360 H200\-GPU\-hours across the 3 seeds in Table[3](https://arxiv.org/html/2605.14040#S5.T3)\(seed\-42 at3030h×4\\times 4H200; seed\-17 and seed\-23 retrainings∼60\\sim 60h×4\\times 4H200 combined; Appendix[F](https://arxiv.org/html/2605.14040#A6)\), plus frontier\-API inference for baselines and Sonnet\-judge runs\. Per\-experiment carbon estimates are omitted because cloud\-provider electricity mix is not consistently disclosed\.

### H\.3Construction\-process disclosures

\(i\) The initialPhysOlym\-Adescription claimed 100% novel\-source; the Stage\-1 audit surfaced one EuPhO 2020 overlap with OlympiadBench\-Physics \(J=0\.91J\{=\}0\.91\), so the honest claim is99\.8%99\.8\\%\(499/500499/500\)—disclosed, not dropped\. \(ii\) Our initial header describedPhyX\-mini\-MCas 500 problems; the canonical MC subset\[Shenet al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib1)\]is 1,000\.

### H\.4Future work and named follow\-ups

Ten follow\-ups consolidated\.*Eval refinements:*\(i\) post\-hoc MCQ\-ification ofPhysOlym\-Ato isolate the format axis \(§[5\.1](https://arxiv.org/html/2605.14040#S5.SS1.SSS0.Px1)\); \(ii\) paraphrase\- and translation\-aware audit pass\.*Cross\-corpus:*\(iii\) audit OlympiadBench\-Physics against our pool, then add the Physics\-R1 row; \(iv\) frontier cross\-evaluation against PHYBench, PhysUniBench, HLE\-physics,PhysOlym\-A\.*Recipe and scale:*\(v\) transfer to InternVL3\-8B and LLaVA\-OneVision\-7B; \(vi\) 32B \+ full\-RL comparator; \(vii\) SFT\-only data\-scaling curve at500/1,293/5,000/9,575500/1\{,\}293/5\{,\}000/9\{,\}575audited prompts\.*Cross\-lingual:*\(viii\) confirm/refute EN/ET sign\-flip on low\-resource open\-source models on the same 59\-pair Estonian Physics Olympiad subset;*pre\-registered*—if 8B\-class open\-source models with documented Estonian\-weak training \(e\.g\. Qwen2\.5\-VL\-7B, LLaVA\-OV\-7B\) show ET<<EN by≥5\\geq 5pp under paired sign test, F2’s Sonnet\-4\.5 ET\>\>EN direction is model\-specific \(confirmed\); a≥5\\geq 5pp result in the same direction \(ET\>\>EN\) on those same models would replicate F2 across model families; results within±5\\pm 5pp are inconclusive atn=59n\{=\}59\.*Methodology:*\(ix\) 50\-problem inter\-annotatorκ\\kappaon the failure\-mode taxonomy; \(x\) embedder\-sensitivity ablation againstvoyage\-3andtext\-embedding\-3\-large\.

### H\.5Versioning and maintenance commitment

Versioned releases with semantic\-version tags \(v1\.0\.0initial,v1\.1\.xadditive,v2\.0\.0schema\-breaking\)\. Each release ships Croissant 1\.0 \+ RAI JSON\-LD, per\-source license matrix \(Table[5](https://arxiv.org/html/2605.14040#A5.T5)\), threshold\-sensitivity grid \(Table[4](https://arxiv.org/html/2605.14040#A1.T4)\) recomputed against newly\-released physics\-VL benchmarks, and the audit\-pipeline source archive with saved best\-overlap scores\. Quarterly contamination audit against new benchmarks; benchmarks introducing≥1%\\geq 1\\%leakage to any held\-out split trigger a documented\-diff removal in the next minor release\. Hosting: HuggingFace \(≥5\\geq 5yr\), GitHub, Zenodo DOI per release\.

### H\.6Method comparison and benchmark comparison tables \(extended\)

Table 8:Physics\-R1 vs\. rule\-based RL recipes for thinking\-mode VLMs\.Init: base or SFT cold\-start\. Reward signals:*Ans*\(answer\),*Fmt*\(format\),*Dim*\(units\),*Sym*\(symbolic\),*Cons*\(conservation\)\. Audit and Filter as defined in §[3\.3](https://arxiv.org/html/2605.14040#S3.SS3)\.✓\\checkmarkpresent; — absent;∘\\circpartial\.RecipeInitAnsFmtDimSymConsAuditFilterDPO\[Rafailovet al\.,[2023](https://arxiv.org/html/2605.14040#bib.bib32)\]SFT———————GRPO\[Shaoet al\.,[2024](https://arxiv.org/html/2605.14040#bib.bib10)\]SFT✓\\checkmark——————GSPO\+DAPO\[Zhenget al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib12), Yuet al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib11)\]SFT✓\\checkmark∘\\circ—————MM\-Eureka\[Menget al\.,[2025](https://arxiv.org/html/2605.14040#bib.bib33)\]base✓\\checkmark✓\\checkmark—————Physics\-R1base✓\\checkmark✓\\checkmark✓\\checkmark✓\\checkmark✓\\checkmark✓\\checkmark✓\\checkmark
### H\.7Corpus composition, hyperparameters, and reward ablation \(extended\)

Table 9:PhysCorp\-pre\-auditcomposition by source family \(14,29414\{,\}294records total before two\-stage audit\)\.The audited releasePhysCorp\-A\(6,4326\{,\}432records\) is the subset surviving Algorithm[1](https://arxiv.org/html/2605.14040#alg1)after a re\-audit pass against PhysReason\-full and PhysUniBench\-en dropped an additional 804 records from a7,2367\{,\}236\-record candidate \(the construction audit excluded those two corpora\)\. The 804\-record re\-audit drop concentrates on PhysReason\-cousin records \(540\) and PhysUniBench\-en\-cousin records \(186\), drawn predominantly from the*repackaged\-benchmark*source families \(UGPhysics / OpenStax / Physics SE / MMMU\+o1\-CoT seed / PhysReason\); the four*first\-ML\-format*source families \(Estonian PhO, Zhou’s handouts, the 7\-international\-olympiad scrapes\) are minimally affected and retain their full1,6091\{,\}609\-record contribution toPhysCorp\-A\. All 9 source families remain represented in the released pool\. Per\-source licenses are carried through to each released record; full provenance and outreach log in Appendix[E](https://arxiv.org/html/2605.14040#A5)\.Source familyCountLicenseAnswer typeUGPhysics5,520CC BY\-NC\-SAOpen\-ended numericOpenStax Physics2,381CC BY 4\.0NumericPhysics Stack Exchange2,291CC BY\-SA 4\.0EquationsRL\-SFT seed \(MMMU \+ o1 CoT\)1,293MIT \(MMMU\); generated CoTVariedPhysReason1,200CC BY 4\.0CoT \+ open\-endedZhou Olympiad Handouts692CC BY\-NC 4\.0†MixedEstonian Physics Olympiad418Public\-domain \(competition policy\)Open\-endedIPhO \+ NBPhO \+ EuPhO scrape258Public domainOpen\-endedAPhO \+ USAPhO \+ INPhO scrape241Public domainOpen\-endedTotal novel1,609——Total corpus14,294——
†\\daggerAuthor granted explicit redistribution permission with attribution \(email confirmation, May 2026\); we redistribute under CC BY\-NC 4\.0\.

Table 10:Physics\-R1 training configuration\.GSPO\+DAPO viaverl0\.6\.1 with vLLM 0\.11\.0 \(TP=4\) for rollouts\. FSDP1 is required for Qwen3\-VL underverl0\.6\.1 \(FSDP2 fails on the multimodal projector path; reproducibility note in §[6](https://arxiv.org/html/2605.14040#S6)\)\. Cells marked*\(recipe\)*differ from a default unconstrained GSPO\+DAPO configuration; together with the dense physics\-native reward \(Section[4](https://arxiv.org/html/2605.14040#S4)\), the audited training pool, and held\-out early stopping, they constitute the Physics\-R1 reference recipe\.ParameterValueRole in the recipeBase / init*\(recipe\)*Qwen3\-VL\-8B\-Thinking BASE \(*no SFT*\)Cold\-start RL, mirroring MM\-Eureka thesisAlgorithmGSPO \+ DAPOStable group\-policy backboneImportance sampling*\(recipe\)*truncated, sequence\-levelRollout\-vs\-train policy correctionLearning rate1×10−61\\\!\\times\\\!10^\{\-6\}Standard for 8B\-class RLLR decay*\(recipe\)*step\-cosine;0\.5×0\.5\\\!\\timesafter step 60Counters drift and length collapseBatch size96Effective batch via gradient accumulationRollouts per prompt16DAPO group sizeMax response length12,288 tokensLong\-CoT thinking budgetKL anchor*\(recipe\)*1×10−31\\\!\\times\\\!10^\{\-3\}to baseAnchors policy to base; bounds driftEntropy bonus*\(recipe\)*1×10−31\\\!\\times\\\!10^\{\-3\}Prevents entropy collapseClip range \(decoupled\)0\.2 / 0\.28DAPO decoupled clipDifficulty curriculum*\(recipe\)*drop 0/16 and 16/16 rollout promptsMM\-Eureka\-style; preserves learning signalReward*\(recipe\)*binary correctness \(§[4](https://arxiv.org/html/2605.14040#S4)\)Recommended default; dense \(Algo\.[2](https://arxiv.org/html/2605.14040#alg2)\) is the shape ablationTrain pool*\(recipe\)*2,268PhysR1CorppromptsClosed\-form carve\-out ofPhysCorp\-A\(§[3](https://arxiv.org/html/2605.14040#S3)\)Early stop*\(recipe\)*held\-out PhyX\-mini\-MCCatches the saturation peakSharding*\(recipe\)*FSDP1 \(not FSDP2\)Required for Qwen3\-VL underverl0\.6\.1Hardware4×\\timesH200 \(single node\)—Step time∼40\{\\sim\}40minutes/stepRollout\-bound \(gen∼65%\{\\sim\}65\\%, update∼23%\{\\sim\}23\\%\)Wall\-clock to saturation peak∼30\{\\sim\}30–4545hours \(steps 60–80\)Training cost∼$​120\{\\sim\}\\mathdollar 120–$​180\\mathdollar 180transformers4\.57\.0 \(pinned\)4\.57\.6 yields different forward pathvllm0\.11\.0TP=4 rollout backendverl0\.6\.1Training frameworkRandom seed \(headline\)423\-seed sweep\{17,23,42\}\\\{17,23,42\\\}reported as the headline binary row in Table[3](https://arxiv.org/html/2605.14040#S5.T3)Table 11:Physics\-R1 reward\-shape ablation on PhyX\-mini\-MC\.Each row toggles one of the five reward components on or off, training Qwen3\-VL\-8B\-Thinking\-base under the same GSPO\+DAPO\+KL\-anchor recipe \(Equation[1](https://arxiv.org/html/2605.14040#S4.E1)\) for the same step budget\. Ans = answer\-correctness binary \(\+1\+1,≡rbin\\equiv r\_\{\\mathrm\{bin\}\}of §[4](https://arxiv.org/html/2605.14040#S4)\); Fmt =\\boxed\{\}format \(\+0\.1\+0\.1\); Dim = dimensional consistency from regex\-detected units \+ sympy unit\-system \(\+0\.15\+0\.15\); Sym = symbolic equation verification of intermediate\\fracexpressions via sympy \(\+0\.20\+0\.20\); Cons = conservation\-law penalty \(energy/momentum\) when applicable \(−0\.25\-0\.25\)\. Composed reward is clipped to\[−1,1\]\[\-1,1\]\. Init in every cell is the Qwen3\-VL\-8B\-Thinking BASE checkpoint \(*no SFT*; cold\-start, mirroring the MM\-Eureka thesis\)\. The Ans\-only row is the recommended Physics\-R1 recipe \(binary correctness reward, §[4](https://arxiv.org/html/2605.14040#S4)\); the all\-on row is the dense ablation of §[4](https://arxiv.org/html/2605.14040#S4); the intermediate rows isolate the marginal effect of each physics\-native shaping component\. Drop\-out cells \(—\) are intermediate\-component runs left to follow\-up work \(compute estimate in Appendix[F](https://arxiv.org/html/2605.14040#A6)\)\.ConfigurationAnsFmtDimSymConsPhyX\-mini\-MCBase \(no RL\)—————73\.7%Physics\-R1 \(binary, recommended\)≡\\equivAns\-only✓————78\.0\+ Format✓✓————\+ Dim✓✓✓———\+ Sym✓✓✓✓——Physics\-R1 \(dense, ablation, all\-on\)✓✓✓✓✓78\.3*Single\-component drop\-outs \(dense ablation minus one\)*−\-Dim✓✓—✓✓—−\-Sym✓✓✓—✓—−\-Cons✓✓✓✓——Table 12:Sonnet 4\.5 strict accuracy on Estonian Physics Olympiad problems by organizer\-issued native difficulty \(n=131\)\. The curve is near\-monotonically decreasing in difficulty and hits a hard floor of 0% at difficulties 3, 6, 8, and 10\. Non\-monotone bumps at 4 and 5 are within sampling noise on small per\-bin counts\.DifficultynCorrectAcc\.1171062\.5%215320\.0%3800\.0%412325\.0%5271037\.0%6900\.0%714321\.4%8900\.0%915321\.4%10500\.0%All1313224\.4%Table 13:Cross\-lingual Sonnet performance on the Estonian Physics Olympiad bilingual subset \(n=59, identical problems, same judge protocol\)\. Strict% = numeric/symbolic match; Liberal% = LLM\-judge score≥0\.5\\geq 0\.5\.LanguagenCorrectPartialIncorrectUnjudgeableStrict%Liberal%English \(translated\)598843013\.6%20\.3%Estonian \(original\)5918927530\.5%38\.1%Table 14:Per\-problem agreement matrix for the EN/ET cross\-lingual ablation \(n=59\)\. The 4\.3:1 asymmetry in the off\-diagonal cells \(ET\-correct/EN\-wrong vs\. EN\-correct/ET\-wrong\) rules out a noise explanation\.ET correctET wrongEN correct5 \(both correct\)3 \(EN only\)EN wrong13 \(ET only\)38 \(both wrong\)

Similar Articles

Physics-IQ Verified

Hugging Face Daily Papers

This paper presents a systematic audit of the Physics-IQ benchmark for evaluating physical understanding in video generative models, proposing improvements to prompts and scoring to enhance reliability.