OmniPhys: A Unified Multimodal Benchmark for Physics Understanding and Generation from Chinese Educational Corpora
Summary
OmniPhys is a large-scale multimodal benchmark for physics understanding and generation, covering middle school to university-level problems from Chinese educational corpora, aimed at evaluating and advancing multimodal large language models in scientific domains.
View Cached Full Text
Cached at: 08/27/26, 09:19 AM
# OmniPhys: A Unified Multimodal Benchmark for Physics Understanding and Generation from Chinese Educational Corpora
Source: [https://arxiv.org/html/2608.25398](https://arxiv.org/html/2608.25398)
###### Abstract
Multimodal Large Language Models \(MLLMs\) have demonstrated strong abilities in solving diverse visual and textual reasoning tasks\. However, their development in the physics domain is significantly hindered by the lack of a comprehensive benchmark\. To fill this gap, we introduceOmniPhys, a large\-scale benchmark for multimodal physics understanding and reasoning, covering middle school through university\-level problems from Chinese Educational Corpora\. OmniPhys consists of15,246questions and19,850images, accompanied by detailed annotations that support fine\-grained analysis of reasoning processes and knowledge usage\. Beyond conventional evaluation, OmniPhys is a benchmark that systematically evaluates multimodal outputs in physics domain, including models’ ability to generate structured physics diagrams, which constitute a fundamental component of authentic physics problem solving\. Extensive evaluations reveal critical gaps in the capabilities of current MLLMs, especially in complex reasoning and visual generation\. To address this, we releaseOmniPhysto serve as a foundational resource for advancing multimodal intelligence in physics and scientific domains\. Codes and data are available at[https://github\.com/ECNU\-RAIL/OmniPhys\-EMNLP2026](https://github.com/ECNU-RAIL/OmniPhys-EMNLP2026)\.
## 1Introduction
The rapid evolution of Large Language Models \(LLMs\)[Jaech et al\. \(2024\)](https://arxiv.org/html/2608.25398#bib.bib1);[OpenAI \(2023\)](https://arxiv.org/html/2608.25398#bib.bib2);[Gheorghe Comanici and others \(2025\)](https://arxiv.org/html/2608.25398#bib.bib3);[Grattafiori et al\. \(2024\)](https://arxiv.org/html/2608.25398#bib.bib4);[Liu et al\. \(2024\)](https://arxiv.org/html/2608.25398#bib.bib5)and Multimodal Large Language Models \(MLLMs\)[Radford et al\. \(2021\)](https://arxiv.org/html/2608.25398#bib.bib6);[Alayrac et al\. \(2022\)](https://arxiv.org/html/2608.25398#bib.bib35);[Li et al\. \(2023a\)](https://arxiv.org/html/2608.25398#bib.bib7)has revolutionized language understanding and visual reasoning across general\-purpose tasks\. However, applying these models to specialized scientific domains, such as mathematical problem\-solving[Lu et al\. \(2023\)](https://arxiv.org/html/2608.25398#bib.bib51);[Yang et al\. \(2024\)](https://arxiv.org/html/2608.25398#bib.bib52), physics problem\-solving[Ding et al\. \(2023\)](https://arxiv.org/html/2608.25398#bib.bib12);[Jaiswal et al\. \(2024\)](https://arxiv.org/html/2608.25398#bib.bib19)and diagram interpretation[Masry et al\. \(2022\)](https://arxiv.org/html/2608.25398#bib.bib9);[Methani et al\. \(2020\)](https://arxiv.org/html/2608.25398#bib.bib8), remains challenging\.Physics, in particular, stands out as a prototypical multimodal arena, demanding the rigorous integration of textual descriptions, visual diagrams, and symbolic logic to achieve accurate reasoning\.
Physics underpins all branches of the natural sciences[Feynman \(1967\)](https://arxiv.org/html/2608.25398#bib.bib11);[Smith \(2007\)](https://arxiv.org/html/2608.25398#bib.bib10)\. While specific datasets have emerged to benchmark physical reasoning[Ding et al\. \(2023\)](https://arxiv.org/html/2608.25398#bib.bib12);[Anand et al\. \(2024\)](https://arxiv.org/html/2608.25398#bib.bib26);[Xu et al\. \(2025\)](https://arxiv.org/html/2608.25398#bib.bib13);[Luo et al\. \(2025\)](https://arxiv.org/html/2608.25398#bib.bib27), existing works rarely satisfy three critical criteria simultaneously: \(1\)Cross\-stage knowledge fusion, spanning the full spectrum from middle school to university levels; \(2\)Multimodal input comprehension, requiring the interpretation of complex textual and visual cues; and \(3\)Multimodal output generation, assessing the model’s ability to actively synthesize diagrams rather than merely selecting options\. The absence of such a unified benchmark hinders a systematic evaluation of current models’ capabilities\.
Figure 1:Overview of the OmniPhys Dataset\.The benchmark encompasses five major physics disciplines, illustrated by representative samples:Mechanics,Electromagnetism,Optics,Thermodynamics, andAcoustics\.To address these challenges, we presentOmniPhys, a unified Chinese benchmark designed to assess physics mastery from secondary education to university levels\. As shown in Figure[1](https://arxiv.org/html/2608.25398#S1.F1), the dataset features a rigorous combination of textual, visual, and symbolic inputs, covering five major physical disciplines including mechanics, electromagnetism, and optics\. To guarantee difficulty and pedagogical validity, all questions are meticulously curated from contemporary examination papers and authoritative textbooks, undergoing a strict multi\-stage filtering process\. Distinctively, OmniPhys introduces a pioneering subset for multimodal output tasks, specifically designed to assess the capabilities of MLLMs in physics diagram understanding and editing\. To establish a benchmark baseline, we conducted comprehensive evaluations on OmniPhys\. Our empirical results reveal that despite recent advancements, significant challenges persist for both proprietary and open\-source MLLMs\. Models across both categories demonstrate substantial room for improvement in handling multimodal physics reasoning across diverse educational stages\. Our main contributions are summarized as follows:
Holistic Benchmark\.We present OmniPhys, an authentic dataset covering middle school to university physics to evaluate cross\-stage reasoning transfer\.
Generative Subset\.We design a novel subset for multimodal output tasks, assessing the models’ capability to synthesize and edit physics diagrams\.
Comprehensive Evaluation\.Extensive experiments with SOTA MLLMs reveal persistent challenges in conceptual mastery, positioning OmniPhys as a rigorous baseline for future research\.
## 2Related Works
Table 1:Comparison with Existing Physics Datasets\.OmniPhysdistinguishes itself by supportingmultimodal outputsand covering the full spectrum of educational stages\. \(Lang\.: Language, PS: Primary School, MS: Middle School, HS: High School, Uni: University\)Overlap\(rightmost column\) indicates the percentage of samples in OmniPhys that overlap with existing benchmarks, computed via string similarity and perceptual hashing\. Our dataset maintains near\-zero overlap while uniquely supporting multimodal outputs across all stages\.### 2\.1Multimodal Large Language Models
Recent MLLMs, such as LLaVA[Liu et al\. \(2023\)](https://arxiv.org/html/2608.25398#bib.bib36)and BLIP\-2[Li et al\. \(2023b\)](https://arxiv.org/html/2608.25398#bib.bib34), leverage instruction tuning to adapt LLM reasoning to multimodal contexts\. While frontier models like GPT\-4V[Wu et al\. \(2024\)](https://arxiv.org/html/2608.25398#bib.bib37)and Gemini[Gheorghe Comanici and others \(2025\)](https://arxiv.org/html/2608.25398#bib.bib3)now support complex interleaved inputs and outputs, they remains susceptible to hallucinations[Bai et al\. \(2024\)](https://arxiv.org/html/2608.25398#bib.bib38)and often falter in rigorous logical deduction or numerical calculation[Yan et al\. \(2025\)](https://arxiv.org/html/2608.25398#bib.bib39);[Jia et al\. \(2026\)](https://arxiv.org/html/2608.25398#bib.bib65)\. These persistent vulnerabilities underscore the urgent necessity for challenging benchmarks to probe the upper bounds of deep multimodal reasoning\.
### 2\.2Benchmark for Physics Reasoning
Table[1](https://arxiv.org/html/2608.25398#S2.T1)summarizes representative physical reasoning datasets\. Early research primarily focused on text\-only modalities[Ding et al\. \(2023\)](https://arxiv.org/html/2608.25398#bib.bib12);[Jaiswal et al\. \(2024\)](https://arxiv.org/html/2608.25398#bib.bib19);[Xu et al\. \(2025\)](https://arxiv.org/html/2608.25398#bib.bib13);[Zheng et al\. \(2025\)](https://arxiv.org/html/2608.25398#bib.bib25);[Qu et al\. \(2026\)](https://arxiv.org/html/2608.25398#bib.bib64)or general K\-12 benchmarks with limited physics subsets[Hendrycks et al\. \(2020\)](https://arxiv.org/html/2608.25398#bib.bib42);[Li et al\. \(2024\)](https://arxiv.org/html/2608.25398#bib.bib17);[Zhong et al\. \(2024\)](https://arxiv.org/html/2608.25398#bib.bib16);[Huang et al\. \(2023\)](https://arxiv.org/html/2608.25398#bib.bib14);[Zhang et al\. \(2023\)](https://arxiv.org/html/2608.25398#bib.bib15)\. While recent works have introduced image inputs[He et al\. \(2024\)](https://arxiv.org/html/2608.25398#bib.bib20);[Huang et al\. \(2024\)](https://arxiv.org/html/2608.25398#bib.bib22);[Li et al\. \(2025a\)](https://arxiv.org/html/2608.25398#bib.bib24);[Zhou et al\. \(2025\)](https://arxiv.org/html/2608.25398#bib.bib41);[Zhang et al\. \(2025a\)](https://arxiv.org/html/2608.25398#bib.bib66), they still suffer from incomplete educational coverage and a lack of multimodal outputs[Guo et al\. \(2025\)](https://arxiv.org/html/2608.25398#bib.bib40)\. To bridge these gaps, we introduceOmniPhys, a unified benchmark meticulously curated from authentic Chinese educational corpora\. Unlike existing resources that frequently overlap in web\-crawled content, OmniPhys maintains a near\-zero overlap \(<0\.05%\) with current benchmarks, ensuring a non\-contaminated and independent evaluation\. By spanning the full spectrum from K\-12 to university levels and uniquely supporting multimodal outputs, OmniPhys provides a rigorous testbed for both physics understanding and diagrammatic generation\.
## 3The OmniPhys Benchmark
We introduceOmniPhys, a comprehensive multimodal physics benchmark spanning junior high to university curricula\. It challenges MLLMs to synergize diagrams, formulas, and text for rigorous, visually grounded reasoning\.
Table 2:Statistics of OmniPhys\.The benchmark maintains a high\-quality distribution across three educational stages\. In response to the growing need for generative AI, we specifically curated a subset forMultimodal Generation, featuring expert\-annotated diagrams\.### 3\.1Data Sources and Distribution
OmniPhys is curated from authoritative Chinese pedagogical resources and authentic examination papers, spanning middle school to university curricula\. The source materials are collected from publicly accessible educational platforms \(e\.g\., Zxxk\.com\) and open examination archives, and used under fair\-use principles for non\-commercial academic research\. To respect intellectual property, the released benchmark includes only structured annotations \(text, reasoning chains, and re\-rendered diagrams\), rather than verbatim source PDFs\. We utilize a multi\-stage pipeline, integrating MinerU[Niu et al\. \(2025\)](https://arxiv.org/html/2608.25398#bib.bib28)and DeepSeek\-OCR[Wei et al\. \(2025\)](https://arxiv.org/html/2608.25398#bib.bib29)to transform raw PDFs into a structuredJSONLformat\. Each entry features comprehensive annotations, including problem texts, associated diagrams, and step\-by\-step reasoning chains\. Representative samples of the source materials are provided in Appendix[D](https://arxiv.org/html/2608.25398#A4)\.
Table[2](https://arxiv.org/html/2608.25398#S3.T2)summarizes the composition of OmniPhys, which is strategically architected to mirror real\-world physics curricula\. Centered on a "difficulty pyramid," the benchmark scales from foundational K\-12 concepts to advanced university\-level challenges\. This hierarchical structure is designed to stress\-test the upper bounds of model reasoning in expert\-level scenarios, moving beyond broad but shallow evaluations\.
### 3\.2Data Preprocessing and Quality Filtering
To ensure the correctness, clarity, and reliability ofOmniPhys, we implement a multi\-stage data preprocessing and filtering pipeline\.
Completeness Screening\.Each sample is strictly required to adhere to a validJSONschema and contain all mandatory fields, including the problem statement, reference answer, and detailed solution\. Furthermore, we prune samples exhibiting insufficient textual length, effectively filtering out incomplete or low\-quality content\.
Semantic Deduplication\.We employ an embedding\-based deduplication strategy using thebge\-small\-zh\-v1\.5encoder[Xiao et al\. \(2023\)](https://arxiv.org/html/2608.25398#bib.bib30)\. We compute dense vector representations for all problem texts and identify duplicates based on a cosine similarity threshold of 0\.95\. This rigorous process eliminated 12\.1% of redundant instances to yield the final valid dataset\.
Visual Dependency Filtering\.To ensure OmniPhys targets genuine multimodal reasoning, we categorize instances into three levels\.Level 1 \(Text\-Solvable\)problems contain only decorative images where all necessary information is fully specified in the text\.Level 2 \(Text\-Descriptive\)problems include images that convey information, but the textual description completely duplicates the visual content\.Level 3 \(Image\-Essential\)problems require direct interpretation of the image, as critical quantities or relationships are only available in the diagram\. To ensure robust classification, we employ a diverse judge panel consisting of DeepSeek\-V3[Liu et al\. \(2024\)](https://arxiv.org/html/2608.25398#bib.bib5), Qwen2\.5[Team \(2025c\)](https://arxiv.org/html/2608.25398#bib.bib32), and GPT\-3\.5[Ye et al\. \(2023\)](https://arxiv.org/html/2608.25398#bib.bib33)\. Only samples classified asLevel 3by all three judges are admitted into the final dataset\. The detailed prompting strategy is provided in Appendix[A\.1](https://arxiv.org/html/2608.25398#A1.SS1)\. To validate this text\-based inference strategy, we conducted human\-machine alignment on a random sample of 300 instances in Appendix[D](https://arxiv.org/html/2608.25398#A4), achieving 89\.3% agreement and Cohen’sκ\\kappa= 0\.82 with expert judgment\. This confirms that visual dependency can be reliably inferred from textual cues alone\.
Difficulty Screening\.We implement a model\-based difficulty screening\. We employ two representative lightweight MLLMs: Qwen2\.5\-VL\-3B[Team \(2025b\)](https://arxiv.org/html/2608.25398#bib.bib55)and MiniCPM\-V\-2\.6[Yao et al\. \(2024\)](https://arxiv.org/html/2608.25398#bib.bib31)\. We adopt an adversarial filtering criterion: any problem correctly solved by both models is deemed insufficiently challenging and is subsequently discarded\. The process is parallelized using thevLLMframework on 2 NVIDIA RTX 4090 GPUs\. This step refines the dataset distribution, removing easy problems to better reflect meaningful cross\-stage physics mastery\.
Data Leakage Prevention\.We specifically selecting materials from frontline educators and physically digitized examinations that are effectively insulated from public indexing\. To empirically verify this isolation, we adopt a search\-based filtering protocol\. We eliminate samples whereGPT\-4[OpenAI \(2023\)](https://arxiv.org/html/2608.25398#bib.bib2)retrieves the exact solution via web browsing or exhibits inconsistent responses when the search function is toggled\. Finally, manual verification is conducted on the remaining subset to guarantee a valid zero\-shot evaluation\.
Human Validation\.Expert cross\-validation on randomized subsets \(Appendix[E](https://arxiv.org/html/2608.25398#A5)\) confirms high consistency between our automated pipeline and human judgment\. Meanwhile, all collected materials were manually screened to exclude personally identifying information and sensitive content\.
Table 3:Main Results on OmniPhys\.Metrics:S1S\_\{1\}\(Answer Accuracy\),S2S\_\{2\}/S3S\_\{3\}\(Process Quality for Objective/Open\-ended tasks\), andPobjP\_\{obj\},PopenP\_\{open\}\(Strict Mastery Rates\)\.Formatting: In each section \(Closed\-source MLLMs/Open\-source MLLMs\), thebestscore is bolded and thesecond bestis underlined\. Theglobal bestacross all models is highlighted inblue\.
## 4Experiments
### 4\.1Experimental Setup
We evaluate a diverse suite of state\-of\-the\-art MLLMs onOmniPhys, spanning both proprietary and open\-source categories\. The models included in the evaluation are detailed below\.
Proprietary Models\.We prioritize recently released models to benchmark the upper bounds of current capabilities\. Our evaluation covers a broad spectrum of high\-performance systems, including GPT\-5\.2[OpenAI \(2025\)](https://arxiv.org/html/2608.25398#bib.bib43), GPT\-5\.1, GPT\-4o, and o4\-mini, alongside competitive counterparts such as Gemini\-3\-pro[Deepmind \(2025\)](https://arxiv.org/html/2608.25398#bib.bib47), Gemini\-2\.5\-Flash[Gheorghe Comanici and others \(2025\)](https://arxiv.org/html/2608.25398#bib.bib3), Claude\-4\.5\-sonnet[Anthropic \(2025\)](https://arxiv.org/html/2608.25398#bib.bib46), Grok\-4[Xai \(2025\)](https://arxiv.org/html/2608.25398#bib.bib44), Doubao\-Seed\-1\.6[Seed \(2025\)](https://arxiv.org/html/2608.25398#bib.bib48), GLM\-4\.6v[AI \(2025\)](https://arxiv.org/html/2608.25398#bib.bib45), and Qwen3\-VL\-Plus[Bai et al\. \(2025\)](https://arxiv.org/html/2608.25398#bib.bib49)\. All proprietary models are accessed via their official APIs\.
Open\-source Models\.We assess the full spectrum of the Qwen2\.5\-VL[Bai et al\. \(2025\)](https://arxiv.org/html/2608.25398#bib.bib49)series \(3B, 7B, 32B, 72B\) and its successor Qwen3\-VL \(2B, 8B, 32B, 235B\), alongside InternVL\-3\.5[Wang et al\. \(2025\)](https://arxiv.org/html/2608.25398#bib.bib50)at 2B, 14B, and 38B scales\. To further ensure architectural diversity, we incorporate representative models such as Kimi\-VL\-A3B[Team et al\. \(2025\)](https://arxiv.org/html/2608.25398#bib.bib56), DeepSeek\-VL\-7B[Lu et al\. \(2024\)](https://arxiv.org/html/2608.25398#bib.bib57), LLaVA\-OneVision\-1\.5\-8B[An et al\. \(2025\)](https://arxiv.org/html/2608.25398#bib.bib58), and Phi\-4\-Multimodal[Abdin et al\. \(2024\)](https://arxiv.org/html/2608.25398#bib.bib59)\. The Qwen3\-VL\-235B\-a22b is accessed via its official API, while all other open\-source models are evaluated locally using 8 NVIDIA RTX 4090 and 2 NVIDIA RTX PRO 6000 GPUs, under identical inference settings\.
### 4\.2Dual\-Track Reasoning Evaluation
To achieve a granular assessment of multimodal reasoning capabilities, we propose theDual\-Track Reasoning Evaluation \(DTRE\)framework\. We categorize problems into two distinct streams: Objective Tasks with deterministic outputs \(e\.g\., multiple\-choice, fill\-in\-the\-blank\) and Open\-Ended Tasks requiring step\-by\-step derivation \(e\.g\., calculation, experimental design\)\.
To quantify model performance, we formalize two complementary metrics:Result Score\(SresS\_\{res\}\) andProcess Score\(SprocS\_\{proc\}\)\.
Figure 2:Performance Degradation Across Educational Stages\.We report the pass rates \(%\) of 12 representative MLLMs across Junior High, Senior High, and University levels\. Both closed\-source \(Left\) and open\-source \(Right\) models exhibit a consistent downward trend, validating the hierarchical difficulty design of OmniPhys\.#### Result Score \(SresS\_\{res\}\)
Designed forObjective Tasks, this metric quantifies the accuracy of the final answer\. We denote the ground truth set as𝒜gt\\mathcal\{A\}\_\{gt\}and the predicted answer set as𝒜pred\\mathcal\{A\}\_\{pred\}\. To ensure robustness against formatting variations, we apply a standard normalization functionϕ\(⋅\)\\phi\(\\cdot\)\.
For single\-choice questions and single\-slot fill\-in\-the\-blank tasks, the scoring is binary:
Sres=𝕀\(ϕ\(𝒜pred\)=ϕ\(𝒜gt\)\)S\_\{res\}=\\mathbb\{I\}\(\\phi\(\\mathcal\{A\}\_\{pred\}\)=\\phi\(\\mathcal\{A\}\_\{gt\}\)\)\(1\)where𝕀\(⋅\)\\mathbb\{I\}\(\\cdot\)is the indicator function\.
For multi\-select questions and multi\-slot tasks, we adopt astrict partial credit mechanismto penalize random guessing\. A score is awarded if and only if the predicted set is a subset of the ground truth \(i\.e\., no incorrect options are selected\)\. The score is calculated as:
Sres=\{\|ϕ\(𝒜pred\)\|\|ϕ\(𝒜gt\)\|ifϕ\(𝒜pred\)⊆ϕ\(𝒜gt\)0otherwiseS\_\{res\}=\\begin\{cases\}\\frac\{\|\\phi\(\\mathcal\{A\}\_\{pred\}\)\|\}\{\|\\phi\(\\mathcal\{A\}\_\{gt\}\)\|\}&\\text\{if \}\\phi\(\\mathcal\{A\}\_\{pred\}\)\\subseteq\\phi\(\\mathcal\{A\}\_\{gt\}\)\\\\ 0&\\text\{otherwise\}\\end\{cases\}\(2\)
#### Process Score \(SprocS\_\{proc\}\)
To quantitatively assess the quality of the Chain\-of\-Thought \(CoT\) in Open\-Ended Tasks, we define a metric based on the completeness of the logical derivation\. Let the ground truth reasoning path be decomposed into a set ofMMkey reasoning steps, denoted as𝒦=\{k1,k2,…,kM\}\\mathcal\{K\}=\\\{k\_\{1\},k\_\{2\},\.\.\.,k\_\{M\}\\\}\. We verify whether each key stepkik\_\{i\}is explicitly or implicitly present and correctly applied in the model’s reasoning pathℛpred\\mathcal\{R\}\_\{pred\}\. The process score is calculated as the recall rate of these key steps:
Sproc=1M∑i=1Mδ\(ki,ℛpred\)S\_\{proc\}=\\frac\{1\}\{M\}\\sum\_\{i=1\}^\{M\}\\delta\(k\_\{i\},\\mathcal\{R\}\_\{pred\}\)\(3\)whereδ\(ki,ℛpred\)∈\{0,1\}\\delta\(k\_\{i\},\\mathcal\{R\}\_\{pred\}\)\\in\\\{0,1\\\}indicates whether theii\-th key step is successfully recovered in the model’s generation\. we denote the Result Score on objective tasks asS1S\_\{1\}, the Process Score \(Sproc\) on objective tasks asS2S\_\{2\}, and the Process Score on open\-ended tasks asS3S\_\{3\}\.
We implement the LLM\-as\-a\-Judge framework[Li et al\. \(2025b\)](https://arxiv.org/html/2608.25398#bib.bib61)\. The final value is derived from the arithmetic mean of scores independently assigned byDeepSeek\-V3\.2[DeepSeek\-AI \(2025\)](https://arxiv.org/html/2608.25398#bib.bib60)andGPT\-4[OpenAI \(2023\)](https://arxiv.org/html/2608.25398#bib.bib2)\. We provide a human\-machine alignment study in Appendix[F](https://arxiv.org/html/2608.25398#A6)\. Although human experts are typically more stringent and assign marginally lower absolute scores, they exhibit high consistency with the LLM judges in terms of both model ranking and error diagnosis\. This strong correlation justifies the use of automated evaluation for large\-scale assessment\. Detailed prompts for both inference and evaluation are provided in Appendix[A\.2](https://arxiv.org/html/2608.25398#A1.SS2)and Appendix[A\.3](https://arxiv.org/html/2608.25398#A1.SS3)\.
To quantify strict mastery, we definePobjP\_\{obj\}as the rate of instances achieving perfect alignment in both result and reasoning \(i\.e\.,S1=S2=1\.0S\_\{1\}=S\_\{2\}=1\.0\), andPopenP\_\{open\}as the rate of flawless derivation in open\-ended tasks \(i\.e\.,S3=1\.0S\_\{3\}=1\.0\)\.
### 4\.3Main Results Analysis
The main evaluation results onOmniPhysare presented in Table[3](https://arxiv.org/html/2608.25398#S3.T3)and Figure[2](https://arxiv.org/html/2608.25398#S4.F2)\. We summarize the key observations as follows:
Dominance of Proprietary Models\.Proprietary systems maintain a clear advantage over open\-source counterparts\. Gemini\-3\-Pro establishes a new state\-of\-the\-art, closely trailed by Doubao\-Seed\-1\.6—both outperforming the flagship GPT\-5\.2 by over 15%, marking a divergence in the top\-tier landscape\. Yet even leading models fail to saturate, with strict mastery rates \(PobjP\_\{obj\}\) remaining below 70%, underscoring that OmniPhys evaluates genuine reasoning rather than rote memorization\.
Scaling Laws and Generational Gains in Open\-Source Models\.The Qwen3\-VL family scales monotonically, culminating in the 235B model that surpasses the proprietary flagship GPT\-5\.2\. Architectural superiority proves equally pivotal: Qwen3\-VL\-32B notably eclipses the substantially larger Qwen2\.5\-VL\-72B\. These findings confirm that algorithmic efficiency is as decisive as raw parameter scale\.
Validating Difficulty Hierarchy\.Figure[2](https://arxiv.org/html/2608.25398#S4.F2)reveals a universal inverse correlation between model performance and educational stages: accuracy peaks at Junior High and degrades monotonically toward University level\. This consistent drop validates the hierarchical design of OmniPhys, confirming that the benchmark captures the escalating cognitive complexity of advanced physics curricula\.
### 4\.4Research on Multimodal Outputs
Current benchmarks predominantly focus on text\-based tasks, overlooking the critical capacity tovisualizeandconstructphysical scenarios\. Since drawing diagrams is a fundamental demonstration of understanding, we extend our evaluation to multimodal generation, a domain rarely explored in existing physics benchmarks\.
Table 4:Performance comparison in multimodal output settings\.Figure 3:A case study of multimodal outputs inOmniPhysdataset\.Table 5:Comparison ofAverage Scores\(scaled to 0\-100\) on the full dataset versus the Test\-Mini set\. The significant score decline \(avg\. \-67\.8%\) confirms theelevated difficultyof the subset, validating itsdistinct research valueas a rigorous probe for robust physical reasoning\.To operationalize this, we define thePhysics Diagram Editingtask\. Unlike standard retrieval, this requires the model to transform a multimodal input \(Iin,TI\_\{in\},T\) into a target state \(IoutI\_\{out\}\) governed by strict physical laws\. Our pilot experiments reveal a critical divergence: while current models exhibit high visual fidelity in general editing, they frequently violate physical constraints—failing to maintain vector directionality during coordinate transformations or preserve the topological integrity of rigid bodies\. This gap underscores the necessity of our generative evaluation, exposing blind spots in physical reasoning that text\-only metrics fail to capture\.
Table 6:Ablation study results comparing MLLMs \(left\) and LLMs \(right\)\. The columns are arranged by decreasing modal information: from Text\+Image to Text\-only\. We report Objective \(PobjP\_\{obj\}\) and Open\-ended \(PopenP\_\{open\}\) Accuracy\.To rigorously assess generation quality, we employed a dual\-track strategy: \(1\)Human Evaluation, where three physics graduate students graded physical correctness, and \(2\)Automated Evaluation, utilizingGPT\-5\.1to score instruction adherence on a discrete\{0,0\.5,1\}\\\{0,0\.5,1\\\}scale\. Detailed annotation protocols and the specific judging prompts are provided in Appendix[A\.4](https://arxiv.org/html/2608.25398#A1.SS4)and Appendix[B](https://arxiv.org/html/2608.25398#A2)\.
As shown in Table[4](https://arxiv.org/html/2608.25398#S4.T4), the model evaluator exhibitsexcessive optimism, consistently inflating scores \(0\.510\.51\-0\.750\.75\) compared to human judgment \(0\.100\.10\-0\.290\.29\)\. This significant divergence indicates that the current MLLM\-as\-a\-judge paradigm lacks the robustness required for rigorous physical verification\. It suggests that relying solely on automated judges is currently insufficient; specifically, the accurate discrimination of visual quality in multimodal outputs remains heavily dependent on human evaluators\. Conversely, the low human ratings attest to the high difficulty and quality of our dataset, offering substantial headroom for future research and model development\.
Figure[3](https://arxiv.org/html/2608.25398#S4.F3)presents a case study revealing that despite high visual fidelity, model\-generated outputs frequently deviate from strict physical correctness\. Specifically, Nano banana and Doubao\-seedream\-4\.0 fail to correctly grasp the concept of a convex lens focus, while Gemini\-3\-pro\-Image and gpt\-Image\-1 introduce erroneous, redundant line segments\. These discrepancies highlight a significant capability gap in both multimodal understanding and precise generation, indicating substantial room for improvement toward human\-level rigor\.
## 5Ablation Study
To enable efficient yet rigorous experimentation, we construct aTest\-Minisubset \(10%10\\%of the data\) using an adversarial hardness\-based strategy\. Specifically, samples are selected according to the consensus of five SOTA models, emphasizing high empirical failure rates \(75%75\\%\) and reasoning complexity \(25%25\\%\)\. We adopt this subset for two reasons: \(1\) the ablation requires evaluating each model under three input settings \(Text\+Img/Caption/Only\), making full\-set evaluation computationally expensive; and \(2\) the harder subset exposes modality differences more clearly, whereas easier samples in the full set tend to saturate performance and obscure cross\-setting gaps\. As shown in Table[5](https://arxiv.org/html/2608.25398#S4.T5), this adversarial selection leads to a substantial performance drop, demonstrating strong discriminative power for ablation analysis\. More details are provided in Appendix[C](https://arxiv.org/html/2608.25398#A3)\.
We evaluate a diverse suite of models, including deep reasoning architectures such as DeepSeek\-R1\-Distill\-Qwen\-7B and P1\-30B\-A3B[Team \(2025a\)](https://arxiv.org/html/2608.25398#bib.bib62), as well as variants fine\-tuned on mathematics like InternLM\-Math\-20B[Cai et al\. \(2024\)](https://arxiv.org/html/2608.25398#bib.bib63)and Qwen2\.5\-Math\-7B or physics like P1\-30B\-A3B\. These models are assessed across three settings to quantify visual dependency:Text\+Img, which utilizes the original multimodal input;Text\+Caption, where diagrams are replaced by textual descriptions; andText\-only, which requires the model to perform blind inference based solely on the question text\.
As shown in Table[6](https://arxiv.org/html/2608.25398#S4.T6), the adversarialTest\-Miniset imposes a significantly higher difficulty than the main benchmark, serving as a rigorous probe for deep reasoning capabilities\. On average, theText\+Imgsetting yields the highest performance, with MLLMs consistently outperforming LLMs, validating the indispensable role of visual constraints in physics problems\. However, we also observe a counter\-intuitive phenomenon where certain models perform better in theText\-onlysetting\. This suggests that for specific architectures, visual inputs or captions may be misinterpreted as distractor noise rather than helpful context, highlighting persistent challenges in cross\-modal alignment that require further investigation\.
## 6Conclusion
In this paper, we presentOmniPhys, the first unified multimodal benchmark that evaluates physics reasoning across the full educational spectrum from junior high to university, covering both understanding and generation\. At its core, the DTRE protocol moves beyond surface\-level answer matching by jointly scoring final results and reasoning fidelity, exposing the reasoning shortcuts that conventional metrics overlook\. We further pioneer the Physics Diagram Editing task and reveal that even frontier image\-generation models systematically violate physical laws, a critical blind spot invisible to text\-only benchmarks\. Extensive experiments across both proprietary and open\-source MLLMs show that OmniPhys poses a substantial challenge, with leading models remaining below 70% strict mastery, while ablation studies confirm the indispensable role of visual grounding\. Collectively, these contributions establish OmniPhys not only as an evaluation suite, but also as a diagnostic instrument for advancing physically grounded multimodal intelligence\.
## 7Limitations
Our current work presents a foundational step with identified limitations that guide future research\. A primary limitation lies in the linguistic and curricular scope of our data sources, which are predominantly drawn from Chinese educational materials\. This monolingual setting introduces a confounding factor: weaker model performance may reflect either limited physics reasoning capability or insufficient Chinese multimodal alignment, particularly under our cross\-lingual prompting protocol where instructions are in English while problems remain in Chinese\. To enhance global representativeness and disentangle language effects from reasoning ability, future iterations will incorporate diverse international curricula, such as A\-Level and IPhO\. Methodologies for evaluating multimodal physical outputs remain underexplored\. A key challenge is the tendency of MLLM\-based judges to overestimate the quality of generated content, often overlooking subtle physical inconsistencies and necessitating costly human evaluation\. While we mitigate this through a structured rubric, a fully scalable solution is still lacking\. We aim to establish a robust evaluation framework tailored for multimodal physics, potentially combining CV\-based structural metrics with LLM semantic scoring to enable reproducible and large\-scale assessment\. Finally, our current failure analysis is limited in depth, offering only qualitative observations of representative error patterns\. We plan to conduct a rigorous error attribution study at scale, systematically distinguishing visual perception failures from logical reasoning errors, thereby providing more granular insights into MLLM mechanisms in physics reasoning\.
## Acknowledgments
This study was funded by Guangxi Science and Technology Program \(2025AB25069309\), the Open Research Fund of Key Laboratory of Advanced Theory and Application in Statistics and Data Science \(East China Normal University\)\.
## References
- Abdinet al\.\(2024\)M\. Abdin, J\. Aneja, H\. Behl, S\. Bubeck, R\. Eldan, S\. Gunasekar, M\. Harrison, R\. J\. Hewett, M\. Javaheripi, P\. Kauffmann, J\. R\. Lee, Y\. T\. Lee, Y\. Li, W\. Liu, C\. C\. T\. Mendes, A\. Nguyen, E\. Price, G\. de Rosa, O\. Saarikivi, A\. Salim, S\. Shah, X\. Wang, R\. Ward, Y\. Wu, D\. Yu, C\. Zhang, and Y\. ZhangPhi\-4 technical report\.External Links:2412\.08905,[Link](https://arxiv.org/abs/2412.08905)Cited by:[§4\.1](https://arxiv.org/html/2608.25398#S4.SS1.p3.1)\.
- AI \(2025\)Z\. AIGLM\-4\.6v system card\.Note:[https://docs\.bigmodel\.cn/cn/guide/models/vlm/glm\-4\.6v](https://docs.bigmodel.cn/cn/guide/models/vlm/glm-4.6v)Released on December 11, 2025Cited by:[§4\.1](https://arxiv.org/html/2608.25398#S4.SS1.p2.1)\.
- Alayracet al\.\(2022\)J\. Alayrac, J\. Donahue, P\. Luc, A\. Miech, I\. Barr, Y\. Hasson, K\. Lenc, A\. Mensch, K\. Millican, M\. Reynolds,et al\.Flamingo: a visual language model for few\-shot learning\.Advances in neural information processing systems35,pp\. 23716–23736\.Cited by:[§1](https://arxiv.org/html/2608.25398#S1.p1.1)\.
- Anet al\.\(2025\)X\. An, Y\. Xie, K\. Yang, W\. Zhang, X\. Zhao, Z\. Cheng, Y\. Wang, S\. Xu, C\. Chen, C\. Wu, H\. Tan, C\. Li, J\. Yang, J\. Yu, X\. Wang, B\. Qin, Y\. Wang, Z\. Yan, Z\. Feng, Z\. Liu, B\. Li, and J\. DengLLaVA\-onevision\-1\.5: fully open framework for democratized multimodal training\.External Links:2509\.23661,[Link](https://arxiv.org/abs/2509.23661)Cited by:[§4\.1](https://arxiv.org/html/2608.25398#S4.SS1.p3.1)\.
- Anandet al\.\(2024\)A\. Anand, J\. Kapuriya, A\. Singh, J\. Saraf, N\. Lal, A\. Verma, R\. Gupta, and R\. ShahMm\-phyqa: multimodal physics question\-answering with multi\-image cot prompting\.InPacific\-Asia Conference on Knowledge Discovery and Data Mining,pp\. 53–64\.Cited by:[§1](https://arxiv.org/html/2608.25398#S1.p2.1),[Table 1](https://arxiv.org/html/2608.25398#S2.T1.2.1.15.1)\.
- Anthropic \(2025\)AnthropicIntroducing claude sonnet 4\.5\.Note:[https://www\.anthropic\.com/news/claude\-sonnet\-4\-5](https://www.anthropic.com/news/claude-sonnet-4-5)Cited by:[§4\.1](https://arxiv.org/html/2608.25398#S4.SS1.p2.1)\.
- Baiet al\.\(2025\)S\. Bai, Y\. Cai, R\. Chen, K\. Chen, X\. Chen, Z\. Cheng, L\. Deng, W\. Ding, C\. Gao, C\. Ge, W\. Ge, Z\. Guo, Q\. Huang, J\. Huang, F\. Huang, B\. Hui, S\. Jiang, Z\. Li, M\. Li, M\. Li, K\. Li, Z\. Lin, J\. Lin, X\. Liu, J\. Liu, C\. Liu, Y\. Liu, D\. Liu, S\. Liu, D\. Lu, R\. Luo, C\. Lv, R\. Men, L\. Meng, X\. Ren, X\. Ren, S\. Song, Y\. Sun, J\. Tang, J\. Tu, J\. Wan, P\. Wang, P\. Wang, Q\. Wang, Y\. Wang, T\. Xie, Y\. Xu, H\. Xu, J\. Xu, Z\. Yang, M\. Yang, J\. Yang, A\. Yang, B\. Yu, F\. Zhang, H\. Zhang, X\. Zhang, B\. Zheng, H\. Zhong, J\. Zhou, F\. Zhou, J\. Zhou, Y\. Zhu, and K\. ZhuQwen3\-vl technical report\.External Links:2511\.21631,[Link](https://arxiv.org/abs/2511.21631)Cited by:[§4\.1](https://arxiv.org/html/2608.25398#S4.SS1.p2.1),[§4\.1](https://arxiv.org/html/2608.25398#S4.SS1.p3.1)\.
- Baiet al\.\(2024\)Z\. Bai, P\. Wang, T\. Xiao, T\. He, Z\. Han, Z\. Zhang, and M\. Z\. ShouHallucination of multimodal large language models: a survey\.arXiv preprint arXiv:2404\.18930\.Cited by:[§2\.1](https://arxiv.org/html/2608.25398#S2.SS1.p1.1)\.
- Caiet al\.\(2024\)Z\. Cai, M\. Cao, H\. Chen, K\. Chen, K\. Chen, X\. Chen, X\. Chen, Z\. Chen, Z\. Chen, P\. Chu, X\. Dong, H\. Duan, Q\. Fan, Z\. Fei, Y\. Gao, J\. Ge, C\. Gu, Y\. Gu, T\. Gui, A\. Guo, Q\. Guo, C\. He, Y\. Hu, T\. Huang, T\. Jiang, P\. Jiao, Z\. Jin, Z\. Lei, J\. Li, J\. Li, L\. Li, S\. Li, W\. Li, Y\. Li, H\. Liu, J\. Liu, J\. Hong, K\. Liu, K\. Liu, X\. Liu, C\. Lv, H\. Lv, K\. Lv, L\. Ma, R\. Ma, Z\. Ma, W\. Ning, L\. Ouyang, J\. Qiu, Y\. Qu, F\. Shang, Y\. Shao, D\. Song, Z\. Song, Z\. Sui, P\. Sun, Y\. Sun, H\. Tang, B\. Wang, G\. Wang, J\. Wang, J\. Wang, R\. Wang, Y\. Wang, Z\. Wang, X\. Wei, Q\. Weng, F\. Wu, Y\. Xiong, C\. Xu, R\. Xu, H\. Yan, Y\. Yan, X\. Yang, H\. Ye, H\. Ying, J\. Yu, J\. Yu, Y\. Zang, C\. Zhang, L\. Zhang, P\. Zhang, P\. Zhang, R\. Zhang, S\. Zhang, S\. Zhang, W\. Zhang, W\. Zhang, X\. Zhang, X\. Zhang, H\. Zhao, Q\. Zhao, X\. Zhao, F\. Zhou, Z\. Zhou, J\. Zhuo, Y\. Zou, X\. Qiu, Y\. Qiao, and D\. LinInternLM2 technical report\.External Links:2403\.17297Cited by:[§5](https://arxiv.org/html/2608.25398#S5.p2.1)\.
- Chenet al\.\(2023\)W\. Chen, M\. Yin, M\. Ku, P\. Lu, Y\. Wan, X\. Ma, J\. Xu, X\. Wang, and T\. XiaTheoremqa: a theorem\-driven question answering dataset\.arXiv preprint arXiv:2305\.12524\.Cited by:[Table 1](https://arxiv.org/html/2608.25398#S2.T1.2.1.12.1)\.
- Daiet al\.\(2025\)S\. Dai, Y\. Yan, J\. Su, D\. Zihao, Y\. Gao, Y\. Hei, J\. Li, J\. Zhang, S\. Tao, Z\. Gao,et al\.PhysicsArena: the first multimodal physics reasoning benchmark exploring variable, process, and solution dimensions\.arXiv preprint arXiv:2505\.15472\.Cited by:[Table 1](https://arxiv.org/html/2608.25398#S2.T1.2.1.16.1)\.
- Deepmind \(2025\)G\. DeepmindGemini3 – our most intelligent ai model that brings any idea to life\.Note:[https://deepmind\.google/models/gemini/](https://deepmind.google/models/gemini/)Cited by:[§4\.1](https://arxiv.org/html/2608.25398#S4.SS1.p2.1)\.
- DeepSeek\-AI \(2025\)DeepSeek\-AIDeepSeek\-v3\.2: pushing the frontier of open large language models\.Cited by:[§4\.2](https://arxiv.org/html/2608.25398#S4.SS2.SSS0.Px2.p2.1.1)\.
- Dinget al\.\(2023\)J\. Ding, Y\. Cen, and X\. WeiUsing large language model to solve and explain physics word problems approaching human level\.arXiv preprint arXiv:2309\.08182\.Cited by:[§1](https://arxiv.org/html/2608.25398#S1.p1.1),[§1](https://arxiv.org/html/2608.25398#S1.p2.1),[§2\.2](https://arxiv.org/html/2608.25398#S2.SS2.p1.1),[Table 1](https://arxiv.org/html/2608.25398#S2.T1.2.1.4.1)\.
- Feynman \(1967\)R\. FeynmanThe character of physical law \(1965\)\.Cox and Wyman Ltd\., London\.Cited by:[§1](https://arxiv.org/html/2608.25398#S1.p2.1)\.
- Gheorghe Comaniciet al\.\(2025\)E\. B\. Gheorghe Comaniciet al\.Gemini 2\.5: pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities\.Note:Technical report / arXivExternal Links:2507\.06261,[Link](https://arxiv.org/abs/2507.06261)Cited by:[§1](https://arxiv.org/html/2608.25398#S1.p1.1),[§2\.1](https://arxiv.org/html/2608.25398#S2.SS1.p1.1),[§4\.1](https://arxiv.org/html/2608.25398#S4.SS1.p2.1)\.
- Grattafioriet al\.\(2024\)A\. Grattafiori, A\. Dubey, A\. Jauhri, A\. Pandey, A\. Kadian, A\. Al\-Dahle, A\. Letman, A\. Mathur, A\. Schelten, A\. Vaughan,et al\.The llama 3 herd of models\.arXiv preprint arXiv:2407\.21783\.Cited by:[§1](https://arxiv.org/html/2608.25398#S1.p1.1)\.
- Guoet al\.\(2025\)M\. Guo, X\. Chu, Q\. Yang, Z\. Mo, Y\. Shen, P\. Li, X\. Lin, J\. Zhang, X\. Chen, Y\. Zhang,et al\.Rbench\-v: a primary assessment for visual reasoning models with multi\-modal outputs\.arXiv preprint arXiv:2505\.16770\.Cited by:[§2\.2](https://arxiv.org/html/2608.25398#S2.SS2.p1.1)\.
- Heet al\.\(2024\)C\. He, R\. Luo, Y\. Bai, S\. Hu, Z\. Thai, J\. Shen, J\. Hu, X\. Han, Y\. Huang, Y\. Zhang,et al\.Olympiadbench: a challenging benchmark for promoting agi with olympiad\-level bilingual multimodal scientific problems\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 3828–3850\.Cited by:[§2\.2](https://arxiv.org/html/2608.25398#S2.SS2.p1.1),[Table 1](https://arxiv.org/html/2608.25398#S2.T1.2.1.14.1)\.
- Hendryckset al\.\(2020\)D\. Hendrycks, C\. Burns, S\. Basart, A\. Zou, M\. Mazeika, D\. Song, and J\. SteinhardtMeasuring massive multitask language understanding\.arXiv preprint arXiv:2009\.03300\.Cited by:[§2\.2](https://arxiv.org/html/2608.25398#S2.SS2.p1.1),[Table 1](https://arxiv.org/html/2608.25398#S2.T1.2.1.6.1)\.
- Huanget al\.\(2023\)Y\. Huang, Y\. Bai, Z\. Zhu, J\. Zhang, J\. Zhang, T\. Su, J\. Liu, C\. Lv, Y\. Zhang, Y\. Fu,et al\.C\-eval: a multi\-level multi\-discipline chinese evaluation suite for foundation models\.Advances in Neural Information Processing Systems36,pp\. 62991–63010\.Cited by:[§2\.2](https://arxiv.org/html/2608.25398#S2.SS2.p1.1),[Table 1](https://arxiv.org/html/2608.25398#S2.T1.2.1.5.1)\.
- Huanget al\.\(2024\)Z\. Huang, Z\. Wang, S\. Xia, X\. Li, H\. Zou, R\. Xu, R\. Fan, L\. Ye, E\. Chern, Y\. Ye,et al\.Olympicarena: benchmarking multi\-discipline cognitive reasoning for superintelligent ai\.Advances in Neural Information Processing Systems37,pp\. 19209–19253\.Cited by:[§2\.2](https://arxiv.org/html/2608.25398#S2.SS2.p1.1)\.
- Jaechet al\.\(2024\)A\. Jaech, A\. Kalai, A\. Lerer, A\. Richardson, A\. El\-Kishky, A\. Low, A\. Helyar, A\. Madry, A\. Beutel, A\. Carney,et al\.Openai o1 system card\.arXiv preprint arXiv:2412\.16720\.Cited by:[§1](https://arxiv.org/html/2608.25398#S1.p1.1)\.
- Jaiswalet al\.\(2024\)R\. Jaiswal, D\. Jain, H\. P\. Popat, A\. Anand, A\. Dharmadhikari, A\. Marathe, and R\. R\. ShahImproving physics reasoning in large language models using mixture of refinement agents\.arXiv preprint arXiv:2412\.00821\.Cited by:[§1](https://arxiv.org/html/2608.25398#S1.p1.1),[§2\.2](https://arxiv.org/html/2608.25398#S2.SS2.p1.1)\.
- Jiaet al\.\(2026\)R\. Jia, Y\. Wei, R\. Li, Y\. Jiang, X\. Xie, Y\. Shen, M\. Zhang, and B\. JiangDiacdm: cognitive diagnosis in teacher\-student dialogues using the initiation\-response\-evaluation framework\.InICASSP 2026\-2026 IEEE International Conference on Acoustics, Speech and Signal Processing \(ICASSP\),pp\. 5186–5190\.Cited by:[§2\.1](https://arxiv.org/html/2608.25398#S2.SS1.p1.1)\.
- Liet al\.\(2025a\)C\. Li, C\. Zhu, T\. Zhang, M\. Lin, Z\. Zhou, and J\. XieK12Vista: exploring the boundaries of mllms in k\-12 education\.arXiv preprint arXiv:2506\.01676\.Cited by:[§2\.2](https://arxiv.org/html/2608.25398#S2.SS2.p1.1),[Table 1](https://arxiv.org/html/2608.25398#S2.T1.2.1.20.1)\.
- Liet al\.\(2025b\)D\. Li, B\. Jiang, L\. Huang, A\. Beigi, C\. Zhao, Z\. Tan, A\. Bhattacharjee, Y\. Jiang, C\. Chen, T\. Wu,et al\.From generation to judgment: opportunities and challenges of llm\-as\-a\-judge\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,pp\. 2757–2791\.Cited by:[§4\.2](https://arxiv.org/html/2608.25398#S4.SS2.SSS0.Px2.p2.1)\.
- Liet al\.\(2024\)H\. Li, Y\. Zhang, F\. Koto, Y\. Yang, H\. Zhao, Y\. Gong, N\. Duan, and T\. BaldwinCmmlu: measuring massive multitask language understanding in chinese\.InFindings of the Association for Computational Linguistics: ACL 2024,pp\. 11260–11285\.Cited by:[§2\.2](https://arxiv.org/html/2608.25398#S2.SS2.p1.1)\.
- Liet al\.\(2023a\)J\. Li, D\. Li, S\. Savarese, and S\. HoiBlip\-2: bootstrapping language\-image pre\-training with frozen image encoders and large language models\.InInternational conference on machine learning,pp\. 19730–19742\.Cited by:[§1](https://arxiv.org/html/2608.25398#S1.p1.1)\.
- Liet al\.\(2023b\)J\. Li, D\. Li, S\. Savarese, and S\. HoiBlip\-2: bootstrapping language\-image pre\-training with frozen image encoders and large language models\.InInternational conference on machine learning,pp\. 19730–19742\.Cited by:[§2\.1](https://arxiv.org/html/2608.25398#S2.SS1.p1.1)\.
- Liuet al\.\(2024\)A\. Liu, B\. Feng, B\. Xue, B\. Wang, B\. Wu, C\. Lu, C\. Zhao, C\. Deng, C\. Zhang, C\. Ruan,et al\.Deepseek\-v3 technical report\.arXiv preprint arXiv:2412\.19437\.Cited by:[§1](https://arxiv.org/html/2608.25398#S1.p1.1),[§3\.2](https://arxiv.org/html/2608.25398#S3.SS2.p4.1)\.
- Liuet al\.\(2023\)H\. Liu, C\. Li, Q\. Wu, and Y\. J\. LeeVisual instruction tuning\.Advances in neural information processing systems36,pp\. 34892–34916\.Cited by:[§2\.1](https://arxiv.org/html/2608.25398#S2.SS1.p1.1)\.
- Luet al\.\(2024\)H\. Lu, W\. Liu, B\. Zhang, B\. Wang, K\. Dong, B\. Liu, J\. Sun, T\. Ren, Z\. Li, Y\. Sun, C\. Deng, H\. Xu, Z\. Xie, and C\. RuanDeepSeek\-vl: towards real\-world vision\-language understanding\.External Links:2403\.05525Cited by:[§4\.1](https://arxiv.org/html/2608.25398#S4.SS1.p3.1)\.
- Luet al\.\(2023\)P\. Lu, H\. Bansal, T\. Xia, J\. Liu, C\. Li, H\. Hajishirzi, H\. Cheng, K\. Chang, M\. Galley, and J\. GaoMathvista: evaluating mathematical reasoning of foundation models in visual contexts\.arXiv preprint arXiv:2310\.02255\.Cited by:[§1](https://arxiv.org/html/2608.25398#S1.p1.1)\.
- Luoet al\.\(2025\)Z\. Luo, Z\. Yin, Y\. Guo, Z\. Wang, J\. Zhu, and X\. TangMulti\-physics: a comprehensive benchmark for multimodal llms reasoning on chinese multi\-subject physics problems\.arXiv preprint arXiv:2509\.15839\.Cited by:[§1](https://arxiv.org/html/2608.25398#S1.p2.1),[Table 1](https://arxiv.org/html/2608.25398#S2.T1.2.1.13.1)\.
- Masryet al\.\(2022\)A\. Masry, D\. X\. Long, J\. Q\. Tan, S\. Joty, and E\. HoqueChartqa: a benchmark for question answering about charts with visual and logical reasoning\.arXiv preprint arXiv:2203\.10244\.Cited by:[§1](https://arxiv.org/html/2608.25398#S1.p1.1)\.
- Menget al\.\(2025\)F\. Meng, L\. Du, Z\. Liu, Z\. Zhou, Q\. Lu, D\. Fu, T\. Han, B\. Shi, W\. Wang, J\. He,et al\.Mm\-eureka: exploring the frontiers of multimodal reasoning with rule\-based reinforcement learning\.arXiv preprint arXiv:2503\.07365\.Cited by:[Table 1](https://arxiv.org/html/2608.25398#S2.T1.2.1.17.1)\.
- Methaniet al\.\(2020\)N\. Methani, P\. Ganguly, M\. M\. Khapra, and P\. KumarPlotqa: reasoning over scientific plots\.InProceedings of the ieee/cvf winter conference on applications of computer vision,pp\. 1527–1536\.Cited by:[§1](https://arxiv.org/html/2608.25398#S1.p1.1)\.
- Niuet al\.\(2025\)J\. Niu, Z\. Liu, Z\. Gu, B\. Wang, L\. Ouyang, Z\. Zhao, T\. Chu, T\. He, F\. Wu, Q\. Zhang, Z\. Jin, G\. Liang, R\. Zhang, W\. Zhang, Y\. Qu, Z\. Ren, Y\. Sun, Y\. Zheng, D\. Ma, Z\. Tang, B\. Niu, Z\. Miao, H\. Dong, S\. Qian, J\. Zhang, J\. Chen, F\. Wang, X\. Zhao, L\. Wei, W\. Li, S\. Wang, R\. Xu, Y\. Cao, L\. Chen, Q\. Wu, H\. Gu, L\. Lu, K\. Wang, D\. Lin, G\. Shen, X\. Zhou, L\. Zhang, Y\. Zang, X\. Dong, J\. Wang, B\. Zhang, L\. Bai, P\. Chu, W\. Li, J\. Wu, L\. Wu, Z\. Li, G\. Wang, Z\. Tu, C\. Xu, K\. Chen, Y\. Qiao, B\. Zhou, D\. Lin, W\. Zhang, and C\. HeMinerU2\.5: a decoupled vision\-language model for efficient high\-resolution document parsing\.External Links:2509\.22186,[Link](https://arxiv.org/abs/2509.22186)Cited by:[§3\.1](https://arxiv.org/html/2608.25398#S3.SS1.p1.1)\.
- OpenAI \(2023\)OpenAIGPT\-4 technical report\.Note:Technical reportExternal Links:2303\.08774,[Link](https://arxiv.org/abs/2303.08774)Cited by:[§1](https://arxiv.org/html/2608.25398#S1.p1.1),[§3\.2](https://arxiv.org/html/2608.25398#S3.SS2.p6.1),[§4\.2](https://arxiv.org/html/2608.25398#S4.SS2.SSS0.Px2.p2.1.2)\.
- OpenAI \(2025\)OpenAIGPT\-5\.2 system card\.Note:[https://openai\.com/index/introducing\-gpt\-5\-2/](https://openai.com/index/introducing-gpt-5-2/)Released on December 11, 2025Cited by:[§4\.1](https://arxiv.org/html/2608.25398#S4.SS1.p2.1)\.
- Quet al\.\(2026\)Z\. Qu, M\. Zhang, M\. Kong, X\. Li, Z\. Shang, Z\. Wang, Y\. Ban, S\. Qiu, Y\. Shu, and Z\. DaiT\-POP: test\-time personalization with online preference feedback\.InProceedings of the 43rd International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.306,Seoul, South Korea\.Cited by:[§2\.2](https://arxiv.org/html/2608.25398#S2.SS2.p1.1)\.
- Radfordet al\.\(2021\)A\. Radford, J\. W\. Kim, C\. Hallacy, A\. Ramesh, G\. Goh, S\. Agarwal, G\. Sastry, A\. Askell, P\. Mishkin, J\. Clark, G\. Krueger, and I\. SutskeverLearning transferable visual models from natural language supervision\.External Links:2103\.00020,[Link](https://arxiv.org/abs/2103.00020)Cited by:[§1](https://arxiv.org/html/2608.25398#S1.p1.1)\.
- Seed \(2025\)B\. SeedSeed1\.6 – tech introduction\.Note:[https://seed\.bytedance\.com/en/seed1\_6](https://seed.bytedance.com/en/seed1_6)Cited by:[§4\.1](https://arxiv.org/html/2608.25398#S4.SS1.p2.1)\.
- Smith \(2007\)G\. SmithNewton’s philosophiae naturalis principia mathematica\.Cited by:[§1](https://arxiv.org/html/2608.25398#S1.p2.1)\.
- Teamet al\.\(2025\)K\. Team, A\. Du, B\. Yin, B\. Xing, B\. Qu, B\. Wang, C\. Chen, C\. Zhang, C\. Du, C\. Wei, C\. Wang, D\. Zhang, D\. Du, D\. Wang, E\. Yuan, E\. Lu, F\. Li, F\. Sung, G\. Wei, G\. Lai, H\. Zhu, H\. Ding, H\. Hu, H\. Yang, H\. Zhang, H\. Wu, H\. Yao, H\. Lu, H\. Wang, H\. Gao, H\. Zheng, J\. Li, J\. Su, J\. Wang, J\. Deng, J\. Qiu, J\. Xie, J\. Wang, J\. Liu, J\. Yan, K\. Ouyang, L\. Chen, L\. Sui, L\. Yu, M\. Dong, M\. Dong, N\. Xu, P\. Cheng, Q\. Gu, R\. Zhou, S\. Liu, S\. Cao, T\. Yu, T\. Song, T\. Bai, W\. Song, W\. He, W\. Huang, W\. Xu, X\. Yuan, X\. Yao, X\. Wu, X\. Zu, X\. Zhou, X\. Wang, Y\. Charles, Y\. Zhong, Y\. Li, Y\. Hu, Y\. Chen, Y\. Wang, Y\. Liu, Y\. Miao, Y\. Qin, Y\. Chen, Y\. Bao, Y\. Wang, Y\. Kang, Y\. Liu, Y\. Du, Y\. Wu, Y\. Wang, Y\. Yan, Z\. Zhou, Z\. Li, Z\. Jiang, Z\. Zhang, Z\. Yang, Z\. Huang, Z\. Huang, Z\. Zhao, and Z\. ChenKimi\-VL technical report\.External Links:2504\.07491,[Link](https://arxiv.org/abs/2504.07491)Cited by:[§4\.1](https://arxiv.org/html/2608.25398#S4.SS1.p3.1)\.
- Team \(2025a\)P\. TeamP1: mastering physics olympiads with reinforcement learning\.External Links:[Link](https://prime-rl.github.io/P1/)Cited by:[§5](https://arxiv.org/html/2608.25398#S5.p2.1)\.
- Team \(2025b\)Q\. TeamQwen2\.5\-vl\.External Links:[Link](https://qwenlm.github.io/blog/qwen2.5-vl/)Cited by:[§3\.2](https://arxiv.org/html/2608.25398#S3.SS2.p5.1)\.
- Team \(2025c\)Q\. TeamQwen3 technical report\.External Links:2505\.09388,[Link](https://arxiv.org/abs/2505.09388)Cited by:[§3\.2](https://arxiv.org/html/2608.25398#S3.SS2.p4.1)\.
- Wanget al\.\(2025\)W\. Wang, Z\. Gao, L\. Gu, H\. Pu, L\. Cui, X\. Wei, Z\. Liu, L\. Jing, S\. Ye, J\. Shao,et al\.Internvl3\. 5: advancing open\-source multimodal models in versatility, reasoning, and efficiency\.arXiv preprint arXiv:2508\.18265\.Cited by:[§4\.1](https://arxiv.org/html/2608.25398#S4.SS1.p3.1)\.
- Weiet al\.\(2025\)H\. Wei, Y\. Sun, and Y\. LiDeepSeek\-ocr: contexts optical compression\.arXiv preprint arXiv:2510\.18234\.Cited by:[§3\.1](https://arxiv.org/html/2608.25398#S3.SS1.p1.1)\.
- Wuet al\.\(2024\)T\. Wu, G\. Yang, Z\. Li, K\. Zhang, Z\. Liu, L\. Guibas, D\. Lin, and G\. WetzsteinGpt\-4v \(ision\) is a human\-aligned evaluator for text\-to\-3d generation\.InProceedings of the IEEE/CVF conference on computer vision and pattern recognition,pp\. 22227–22238\.Cited by:[§2\.1](https://arxiv.org/html/2608.25398#S2.SS1.p1.1)\.
- Xai \(2025\)XaiGrok\-4 system card\.Note:[https://x\.ai/news/grok\-4](https://x.ai/news/grok-4)Released on December 11, 2025Cited by:[§4\.1](https://arxiv.org/html/2608.25398#S4.SS1.p2.1)\.
- Xianget al\.\(2025\)K\. Xiang, H\. Li, T\. J\. Zhang, Y\. Huang, Z\. Liu, P\. Qu, J\. He, J\. Chen, Y\. Yuan, J\. Han,et al\.SeePhys: does seeing help thinking?–benchmarking vision\-based physics reasoning\.arXiv preprint arXiv:2505\.19099\.Cited by:[Table 1](https://arxiv.org/html/2608.25398#S2.T1.2.1.19.1)\.
- Xiaoet al\.\(2023\)S\. Xiao, Z\. Liu, P\. Zhang, and N\. MuennighoffC\-pack: packaged resources to advance general chinese embedding\.External Links:2309\.07597Cited by:[§3\.2](https://arxiv.org/html/2608.25398#S3.SS2.p3.1)\.
- Xuet al\.\(2025\)X\. Xu, Q\. Xu, T\. Xiao, T\. Chen, Y\. Yan, J\. Zhang, S\. Diao, C\. Yang, and Y\. WangUgphysics: a comprehensive benchmark for undergraduate physics reasoning with large language models\.arXiv preprint arXiv:2502\.00334\.Cited by:[§1](https://arxiv.org/html/2608.25398#S1.p2.1),[§2\.2](https://arxiv.org/html/2608.25398#S2.SS2.p1.1),[Table 1](https://arxiv.org/html/2608.25398#S2.T1.2.1.9.1)\.
- Yanet al\.\(2025\)Y\. Yan, J\. Su, J\. He, F\. Fu, X\. Zheng, Y\. Lyu, K\. Wang, S\. Wang, Q\. Wen, and X\. HuA survey of mathematical reasoning in the era of multimodal large language model: benchmark, method & challenges\.InFindings of the Association for Computational Linguistics: ACL 2025,pp\. 11798–11827\.Cited by:[§2\.1](https://arxiv.org/html/2608.25398#S2.SS1.p1.1)\.
- Yanget al\.\(2024\)A\. Yang, B\. Zhang, B\. Hui, B\. Gao, B\. Yu, C\. Li, D\. Liu, J\. Tu, J\. Zhou, J\. Lin,et al\.Qwen2\. 5\-math technical report: toward mathematical expert model via self\-improvement\.arXiv preprint arXiv:2409\.12122\.Cited by:[§1](https://arxiv.org/html/2608.25398#S1.p1.1)\.
- Yaoet al\.\(2024\)Y\. Yao, T\. Yu, A\. Zhang, C\. Wang, J\. Cui, H\. Zhu, T\. Cai, H\. Li, W\. Zhao, Z\. He,et al\.MiniCPM\-v: a gpt\-4v level mllm on your phone\.arXiv preprint arXiv:2408\.01800\.Cited by:[§3\.2](https://arxiv.org/html/2608.25398#S3.SS2.p5.1)\.
- Yeet al\.\(2023\)J\. Ye, X\. Chen, N\. Xu, C\. Zu, Z\. Shao, S\. Liu, Y\. Cui, Z\. Zhou, C\. Gong, Y\. Shen, J\. Zhou, S\. Chen, T\. Gui, Q\. Zhang, and X\. HuangA comprehensive capability analysis of gpt\-3 and gpt\-3\.5 series models\.External Links:2303\.10420,[Link](https://arxiv.org/abs/2303.10420)Cited by:[§3\.2](https://arxiv.org/html/2608.25398#S3.SS2.p4.1)\.
- Zhanget al\.\(2025a\)M\. Zhang, B\. Jiang, J\. Zhou, Y\. Liu, and X\. LinCoDoL: conditional domain prompt learning for out\-of\-distribution generalization\.InTMLR 2026,Cited by:[§2\.2](https://arxiv.org/html/2608.25398#S2.SS2.p1.1)\.
- Zhanget al\.\(2023\)X\. Zhang, C\. Li, Y\. Zong, Z\. Ying, L\. He, and X\. QiuEvaluating the performance of large language models on gaokao benchmark\.arXiv preprint arXiv:2305\.12474\.Cited by:[§2\.2](https://arxiv.org/html/2608.25398#S2.SS2.p1.1),[Table 1](https://arxiv.org/html/2608.25398#S2.T1.2.1.7.1)\.
- Zhanget al\.\(2025b\)X\. Zhang, Y\. Dong, Y\. Wu, J\. Huang, C\. Jia, B\. Fernando, M\. Z\. Shou, L\. Zhang, and J\. LiuPhysreason: a comprehensive benchmark towards physics\-based reasoning\.arXiv preprint arXiv:2502\.12054\.Cited by:[Table 1](https://arxiv.org/html/2608.25398#S2.T1.2.1.18.1)\.
- Zhenget al\.\(2025\)S\. Zheng, Q\. Cheng, J\. Yao, M\. Wu, H\. He, N\. Ding, Y\. Cheng, S\. Hu, L\. Bai, D\. Zhou,et al\.Scaling physical reasoning with the physics dataset\.arXiv preprint arXiv:2506\.00022\.Cited by:[§2\.2](https://arxiv.org/html/2608.25398#S2.SS2.p1.1),[Table 1](https://arxiv.org/html/2608.25398#S2.T1.2.1.10.1)\.
- Zhonget al\.\(2024\)W\. Zhong, R\. Cui, Y\. Guo, Y\. Liang, S\. Lu, Y\. Wang, A\. Saied, W\. Chen, and N\. DuanAgieval: a human\-centric benchmark for evaluating foundation models\.InFindings of the Association for Computational Linguistics: NAACL 2024,pp\. 2299–2314\.Cited by:[§2\.2](https://arxiv.org/html/2608.25398#S2.SS2.p1.1),[Table 1](https://arxiv.org/html/2608.25398#S2.T1.2.1.8.1)\.
- Zhouet al\.\(2025\)P\. Zhou, F\. Zhang, X\. Peng, Z\. Xu, J\. Ai, Y\. Qiu, C\. Li, Z\. Li, M\. Li, Y\. Feng,et al\.MDK12\-bench: a multi\-discipline benchmark for evaluating reasoning in multimodal large language models\.arXiv preprint arXiv:2504\.05782\.Cited by:[§2\.2](https://arxiv.org/html/2608.25398#S2.SS2.p1.1),[Table 1](https://arxiv.org/html/2608.25398#S2.T1.2.1.21.1)\.
## Appendix APrompt Usage
All prompts presented in this section are translated into English for readability\. In our actual experiments, prompts are administered in Chinese to align with the language of the OmniPhys benchmark, ensuring consistent cross\-modal understanding for both inference and evaluation\.
### A\.1Visual Dependency Annotation
Prompt for Visual Dependency AnnotationYou are an analyst of multimodal physics problem datasets\. You are given access only to the problem text, without seeing the accompanying image\. Based on the textual description, infer the importance of the image for solving the problem\.Image importance levels are defined as follows:Level 1 \(Text\-Only Solvable\): The image is purely decorative or illustrative\. All information required to solve the problem is fully described in the text, and the image is not necessary\.Level 2 \(Text\-Descriptive\): The image contains information, but the text fully restates or describes all visual content\. The image serves only as an aid for understanding and is not strictly required for reasoning\.Level 3 \(Image\-Essential\): The text refers to the image \(e\.g\., "as shown in the figure"\), and critical information such as numerical values, geometric relationships, or measurement readings is available only in the image\. The image is essential for solving the problem\.Please output only a single integer \(1, 2, or 3\) corresponding to the image importance level\.Problem text: \{problem\_text\}
### A\.2Model Inference Prompt
Prompt for Model InferenceI will present you with a physics problem\. Please read and solve it\. Adhere strictly to the following output format:”’ Reasoning Process: \[Solve step\-by\-step logically\. Limit each step to no more than 30 words, focusing only on core deductions without redundant explanations\.\]Answer: \[Output your final answer\. Do not add extra content\.\] ”’\[Input Image\]\{problem\_text\}
### A\.3Automated Evaluation Prompts
We employ a dynamic prompting strategy based on the question type\. The evaluator \(LLM\-as\-a\-Judge\) receives specific instructions for objective and open\-ended tasks respectively\.
Prompt for Objective Tasks \(Dual\-Track Evaluation\)You are a rigorous grader\. Please conduct a dual\-track evaluation for this objective task: assess both the correctness of the final result and the validity of the reasoning process\.\[Question Data\]\[Question\]: \{question\} \[Standard Answer\]: \{std\_answer\} \[Explanation\]: \{std\_explanation\}\[Student Response\]\{model\_answer\}\[Evaluation Tasks\]Please complete the following two parts and merge them into a JSON output:1\.Process Analysis \(process\_eval\): \- Decompose the standard solution intommkey steps\. \- Count how many steps \(nn\) the student correctly completed\. \- Process Score =n/mn/m\. \- If the student provides the correct answer without reasoning steps, treat it as a "reasoning shortcut" \(assign score based on context, typically penalized or full if trivial\)\.2\.Result Analysis \(result\_eval\): \- Ignore the process and strictly check if the final answer matches the standard\. \- For Single Choice / True\-False: 1\.0 for match, 0\.0 for mismatch\. \- For Multi\-Select / Fill\-in\-the\-Blank: Score = \(Count of Correct Slots\) / \(Total Required Slots\)\. \- Ignore format variations \(e\.g\., "A" vs\. "A\."\)\.\[Output Requirement\]Output ONLY in JSON format without Markdown tags: \{ "process\_eval": \{ "reason": "Brief analysis \(Total m steps, Correct n steps\)", "total\_steps": m, "correct\_steps": n, "process\_score": 0\.0 to 1\.0 \}, "result\_eval": \{ "reason": "Brief justification for the result", "score": 0\.0 to 1\.0 \} \}
Prompt for Open\-Ended Tasks \(Process\-Only Evaluation\)You are an expert physics evaluator\. Your task is to assess the student’s reasoning logic step\-by\-step\.\[Question Data\]\[Question\]: \{question\} \[Standard Answer\]: \{std\_answer\} \[Explanation\]: \{std\_explanation\}\[Student Response\]\{model\_answer\}\[Scoring Criteria\]1\. Decompose the complete resolution process intommKey Reasoning Steps\. 2\. Determine how many of these steps \(nn\) are correctly included in the student’s response\. 3\. The final Process Score isn/mn/m\.\[Output Requirement\]Output ONLY in JSON format without Markdown tags: \{ "reason": "Step\-by\-step analysis of the derivation", "total\_steps": m \(integer\), "correct\_steps": n \(integer\), "process\_score": Calculated result of n/m \(0\.0 to 1\.0\) \}
### A\.4Automated Evaluation Prompts for Multimodal Generation
We utilize a Vision\-Language Model \(e\.g\., GPT\-5\.1 or GPT\-4o\) as the evaluator to assess the quality of generated physics diagrams\. The judge receives the problem text, the context image \(if available\), and the generated diagram, then assigns a discrete score based on physical fidelity\.
Prompt for Diagram Editing EvaluationYou are a professional and rigorous physics instructor\. Your task is to grade physics diagrams drawn by students based on specific problem requirements\.\[Input Data\]1\.Problem Description: Describes the physical scenario or process to be drawn\. 2\.Context Image\(Optional\): Provides background information \(if any\)\. 3\.Student Answer Image: The diagram generated by the student\.\[Evaluation Task\]Based on physical principles, comprehensively evaluate whether the student’s image accurately and completely fulfills the problem requirements\. The diagrams may involve mechanics analysis, motion trajectories, optical paths, circuit connections, or electromagnetic field distributions\.\[Scoring Criteria\]•1\.0 \(Fully Correct\): The image perfectly reflects the problem requirements\. All key physical elements \(e\.g\., vector directions, points of application, trajectory shapes, light paths, circuit connections, physical labels\) are accurate and adhere to physical laws\.•0\.5 \(Partially Correct\): The image captures the core physical concept but contains defects in details\. Examples: Main structure is correct but minor labels are missing; key vectors are roughly correct in direction but deviate significantly in angle or proportion; general trend is correct but local errors exist; or unnecessary misleading lines are included\.•0\.0 \(Incorrect\): The image contains fundamental physical errors or omits the most critical information, failing to reflect the problem requirements\. Examples: Depicting the wrong physical phenomenon; missing core elements required by the stem; or generating an image completely irrelevant to the problem\.\[Output Requirement\]Please output strictly in JSON format \(no Markdown tags\): \{ "reasoning": "Concise justification pointing out specific merits or demerits", "score": 1\.0 or 0\.5 or 0\.0 \}\[Input Sequence\]\[Problem Description\]: \{question\_text\}\[Context Image\]: \(Input Image Here\)\[Student Answer Image\]: \(Generated Image Here\)
Figure 4:Screenshot of the Label Studio interface for human evaluation of multimodal outputs\.
## Appendix BAnnotation Interface Details
To ensure the quality and consistency of our evaluation, we utilized the Label Studio platform for manual annotation\. Figure[4](https://arxiv.org/html/2608.25398#A1.F4)illustrates the annotation interface used by human evaluators to assess the model’s multimodal outputs\.
The interface was carefully designed to present all relevant information required for reliable judgment within a single view\. Specifically, it simultaneously displays \(1\) the original problem statement, \(2\) the original input image, \(3\) the ground\-truth reference image provided by the dataset, and \(4\) the image generated by the evaluated model\. This design allows annotators to directly compare the model output with the reference solution in the context of the original task, thereby making the scoring process both efficient and less prone to omission or misinterpretation\.
For each sample, annotators were asked to assign a quality score based on the correctness, completeness, and visual faithfulness of the generated image with respect to the reference answer\. Since all relevant inputs and outputs are visible in a unified interface, the evaluation process is straightforward to operate and reduces unnecessary cognitive load on the annotators\.
The annotation was conducted by three graduate students specializing in physics, all of whom had prior experience with problem solving and diagram interpretation in the target domain\. To improve reliability, each sample was independently annotated by all three evaluators, and the final human evaluation score reported in our experiments is computed as the average of the three scores\. This aggregation strategy helps mitigate individual bias and increases the robustness of the evaluation results\. The annotators were compensated with standard research assistant stipends in accordance with institutional guidelines\.
## Appendix CTest\-Mini Construction and Ablation Details
In this section, we provide a detailed elaboration on the construction process of theTest\-Miniset and the specific settings used for the ablation study\.
### C\.1Hardness\-Based Selection Methodology
To ensure theTest\-Miniset effectively probes the upper limits of multimodal reasoning, we implemented a multi\-dimensional hardness scoring mechanism\. For each candidate questionxix\_\{i\}in the full dataset \(Without multimodal outputs,N=12,885N=12,885\), we computed a hardness scoreH\(xi\)H\(x\_\{i\}\)based on two factors: empirical model failure and reasoning complexity\.
Table 7:Comparison of key statistics between the full dataset and the selected Test\-Mini subset\. The subset demonstrates significantly higher complexity and lower model solvability\.The score is defined as:
H\(xi\)=w1⋅ℱ\(xi\)\+w2⋅𝒞\(xi\)H\(x\_\{i\}\)=w\_\{1\}\\cdot\\mathcal\{F\}\(x\_\{i\}\)\+w\_\{2\}\\cdot\\mathcal\{C\}\(x\_\{i\}\)\(4\)where:
- •ℱ\(xi\)\\mathcal\{F\}\(x\_\{i\}\)represents theEmpirical Failure Rate\. It is calculated asℱ\(xi\)=1−1\|M\|∑m∈MS\(m,xi\)\\mathcal\{F\}\(x\_\{i\}\)=1\-\\frac\{1\}\{\|M\|\}\\sum\_\{m\\in M\}S\(m,x\_\{i\}\), whereMMis the set of five baseline models andS\(m,xi\)∈\[0,1\]S\(m,x\_\{i\}\)\\in\[0,1\]is the normalized score of modelmmon questionxix\_\{i\}\.
- •𝒞\(xi\)\\mathcal\{C\}\(x\_\{i\}\)represents theReasoning Complexity, quantified by the percentile rank of the character length of the ground truth explanation \(Chain\-of\-Thought\)\.
- •We set the weights tow1=0\.75w\_\{1\}=0\.75andw2=0\.25w\_\{2\}=0\.25, prioritizing empirical difficulty while accounting for logical depth\.
The baseline models setMMincludes five state\-of\-the\-art closed\-source models:Doubao\-seed\-1\.6,Gemini\-3\-pro\-preview,GLM\-4\.6v,GPT\-5\.2, andQwen3\-vl\-plus\. We selected the top 10% of samples ranked byH\(xi\)H\(x\_\{i\}\)to form theTest\-Miniset\.
### C\.2Statistics and Performance Gap
The selection process resulted in a subset with significantly higher difficulty\. As shown in Table[7](https://arxiv.org/html/2608.25398#A3.T7), theTest\-Miniset exhibits a sharp increase in the "All\-Model Failure Rate" \(questions where no model answered correctly\) from 1\.6% to 15\.5%, and a substantial increase in the average reasoning chain length\.
Table[5](https://arxiv.org/html/2608.25398#S4.T5)details the performance drop for each baseline model\. The consistent degradation across all models \(ranging from 56\.9% to 72\.5%\) confirms that the difficulty of theTest\-Miniset is not biased towards a specific architecture but stems from the inherent complexity of the physics problems\.
## Appendix DRepresentative Source Materials
To demonstrate the authenticity and complexity of our data, we provide a selection of original screenshots from the standardized Chinese physics examination papers and authoritative textbooks used to curate theOmniPhysbenchmark\. Figure[5](https://arxiv.org/html/2608.25398#A7.F5), Figure[6](https://arxiv.org/html/2608.25398#A7.F6)and Figure[7](https://arxiv.org/html/2608.25398#A7.F7)show 3 examples of original physics problems from Chinese real\-world exams\.
## Appendix EHuman\-Machine Alignment and Quality Control in Data Filtering Process
To address potential biases in the automated components of our pipeline—specifically Visual Dependency Filtering and Difficulty Screening—we conducted a blind correlation study involving two physics experts \(Ph\.D\. candidates\)\. We randomly sampled 300 instances for manual review\.
Table 8:Alignment Statistics for Data Refinement\.The "Human\-Machine" column represents the consistency between our automated pipeline \(using MLLMs and vLLM\) and expert consensus\. Scores above 0\.80 signify strong reliability in our data quality control\.Table 9:Human\-Machine Alignment Results\.Comparison between expert consensus and the LLM\-as\-a\-Judge framework across primary metrics\.Δ\\DeltaScore represents the mean difference \(Machine−\-Human\)\.The high alignment scores in Table[8](https://arxiv.org/html/2608.25398#A5.T8)demonstrate that our adversarial filtering strategy effectively targets non\-trivial problems while maintaining the pedagogical integrity of the physics domains\.
## Appendix FHuman\-Machine Alignment on Evaluation
To validate the reliability of the LLM\-as\-a\-Judge framework \(comprisingDeepSeek\-V3andGPT\-4\) used in our main experiments, we conducted a systematic alignment study\. We randomly sampled 300 model\-generated responses across five physics domains and three task types\. Three physics experts \(graduate students\) independently scored these responses using the same rubric as the LLM judges\.
Table[9](https://arxiv.org/html/2608.25398#A5.T9)presents the correlation and agreement metrics\. Our analysis reveals three key findings:
- •Strong Ranking Consistency: The Pearson correlation \(rr\) forS1S\_\{1\}\(Accuracy\) andS2S\_\{2\}\(Objective Process\) exceeds 0\.85, indicating that the LLM judges and human experts are highly consistent in ranking model performance across different physics domains\.
- •Expert Stringency: We observe that human experts are typically more stringent, resulting in absolute scores marginally lower than the LLM consensus \(with a mean differenceΔ\\Deltaof 1\.2% to 4\.2%\)\. This is primarily due to experts’ higher sensitivity to subtle conceptual inaccuracies in complex reasoning chains\.
- •Robust Mastery Detection: The agreement on strict correctness rates \(PobjP\_\{obj\}andPopenP\_\{open\}\) reaches 91\.2%, confirming that the "perfect performance" threshold used in Table[3](https://arxiv.org/html/2608.25398#S3.T3)provides a reliable signal for identifying model mastery\.
Despite the minor bias in absolute values, the strong correlation across all dimensions justifies the scalability of our automated evaluation protocol for the large\-scale assessment in OmniPhys\.
## Appendix GAI Assistant Usage Disclosure
We used AI assistants \(Gemini/ChatGPT\) to assist with writing script code for data processing and for polishing the language of the manuscript to improve readability\. All scientific claims, experimental designs, and final text were manually verified and revised by the authors\.
Figure 5:A middle school, Mechanics example of Original Physics Problems\.Figure 6:A high school, Electromagnetism example of Original Physics Problems\.Figure 7:A University, Optics example of Original Physics Problems\.Similar Articles
OmniPro: A Comprehensive Benchmark for Omni-Proactive Streaming Video Understanding
OmniPro is the first benchmark for evaluating proactive streaming video understanding in omni-modal large language models, featuring 2,700 samples covering diverse tasks and dual-mode evaluation protocols.
@ziqi_huang_: An interesting work on Physical AI: PhysX-Omni. First unified sim-ready generation framework for rigid, deformable, and…
PhysX-Omni is a unified framework for simulation-ready physical 3D generation covering rigid, deformable, and articulated objects, with a new dataset (PhysXVerse) and benchmark (PhysX-Bench).
S1-Omni: A Unified Multimodal Reasoning Model for Scientific Understanding, Prediction, and Generation
S1-Omni is a unified multimodal reasoning model for scientific tasks including understanding, prediction, and generation. It is trained on a corpus of 200 scientific tasks and outperforms GPT-5.5 and Gemini-3.1-Pro on most benchmarks.
Physics-R1: An Audited Olympiad Corpus and Recipe for Visual Physics Reasoning
This paper audits multimodal physics evaluation pipelines, revealing issues like train-eval contamination, translation drift, and MCQ saturation. It releases new datasets (PhysCorp-A, PhysR1Corp, PhysOlym-A) and a training recipe (Physics-R1) that significantly improves performance on held-out olympiad problems.
MCBench: A Multicontext Safety Assessment Benchmark for Omni Large Language Models
MCBench is a new benchmark for assessing the safety of omnimodal large language models across vision, audio, and text modalities. It includes 1196 scenarios and finds current models struggle with cross-modal safety reasoning.