Science Edge Evaluation: SEE the Missing Step Toward Real Scientific Discovery

arXiv cs.AI Papers

Summary

This paper introduces SEE, a multimodal benchmark of expert-curated questions for scientific discovery in chemistry, biology, and materials science. Evaluation of 19 MLLMs shows the best model reaches only 48.7% accuracy, and even with tool use only 52.7%, revealing that current models lack reliable evidence-bounded scientific reasoning.

arXiv:2608.06931v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly involved in scientific discovery, yet it remains unclear whether they can support complex real laboratory science. Here we introduce Science Edge Evaluation (SEE), a multimodal benchmark of expert-curated questions grounded in peer-reviewed literature and experimental practice in chemistry, biology, and materials science. Evaluation of 19 multimodal large language models (MLLMs) shows that even the best-performing model reaches only 48.7% accuracy. Moreover, general-purpose models outperform science-specialized models on average. In the visual-agent evaluation, the use of tools increases the best accuracy to 52.7%. Tool use can expand the information available to models, but more information does not necessarily lead to reliable scientific reasoning. The key challenge is whether models can manage tool-derived information within the boundaries of the original experimental evidence. Together, these findings reveal that current MLLMs still cannot reliably make justified and evidence-bounded inferences from experimental results, which is an essential capability in real scientific discovery. Bridging this gap requires MLLMs to transition from explaining established scientific concepts to deriving novel and evidence-based insights from experimental data.
Original Article
View Cached Full Text

Cached at: 08/10/26, 07:59 AM

# 1 Introduction
Source: [https://arxiv.org/html/2608.06931](https://arxiv.org/html/2608.06931)
![[Uncaptioned image]](https://arxiv.org/html/2608.06931v1/x1.png)![[Uncaptioned image]](https://arxiv.org/html/2608.06931v1/figures/alibaba_data_logo.jpg)![[Uncaptioned image]](https://arxiv.org/html/2608.06931v1/figures/skylenage.jpg)

Science Edge Evaluation: SEE the Missing Step Toward Real Scientific Discovery

Taolin Han1,3\*, Yuchen Zhang1,4\*, Jinghang Wang1\*, Yun Wu1, Wai Yuet Chiu1, Zhaohai Li2, Yifei Zhang1,5, Jinxin Wang1,4, Yuhao Zhou1,6, Chen Zhao1,4, Jiajia Li1, Jiaxin Li1, Qile Jin1, Kewei Sun1, Shuang Wu1, Weiqi Zhai1, Renquan Lv1,6, Junchao Li1, Ruodan Chen1, Qingteng Chen1, Zhibo Yang2, Hu Wei1, Lin Qu1, Shuai Bai2,†, Bing Zhao1,†

1Alibaba Group,2Qwen Team, Alibaba Group,3University of Chinese Academy of Sciences,4Tsinghua University,5University of Alberta,6Zhejiang University

\*Equal Contribution,†Correspondence

[wangjinghang\.wjh@alibaba\-inc\.com](https://arxiv.org/html/2608.06931v1/mailto:[email protected])

[![[Uncaptioned image]](https://arxiv.org/html/2608.06931v1/x2.png)Code](https://github.com/SKYLENAGE-AI/science-edge-evaluation)

AbstractLarge language models \(LLMs\) are increasingly involved in scientific discovery, yet it remains unclear whether they can support complex real laboratory science\. Here we introduceScience Edge Evaluation\(SEE\), a multimodal benchmark of expert\-curated questions grounded in peer\-reviewed literature and experimental practice in chemistry, biology, and materials science\. Evaluation of 19 multimodal large language models \(MLLMs\) shows that even the best\-performing model reaches only 48\.7% accuracy\. Moreover, general\-purpose models outperform science\-specialized models on average\. In the visual\-agent evaluation, the use of tools increases the best accuracy to 52\.7%\. Tool use can expand the information available to models, but more information does not necessarily lead to reliable scientific reasoning\. The key challenge is whether models can manage tool\-derived information within the boundaries of the original experimental evidence\. Together, these findings reveal that current MLLMs still cannot reliably make justified and evidence\-bounded inferences from experimental results, which is an essential capability in real scientific discovery\. Bridging this gap requires MLLMs to transition from explaining established scientific concepts to deriving novel and evidence\-based insights from experimental data\.

Large language models \(LLMs\) are transforming the scientific discovery process by accelerating some of its main stages, including literature screening, computational simulation, code synthesis, and, more recently, autonomous experimentation\. As these models become more capable and widely adopted, they are beginning to shape the way scientific knowledge is produced, evaluated, and applied across disciplines\. This emerging role positions LLMs not merely as tools for information retrieval or text generation, but as active contributors to the broader research workflow\.

Current evaluation paradigms often prioritize breadth over depth and rely heavily on sanitized, exam\-style problems that reward pattern matching and memorization\. Such settings fail to capture the dynamic, multi\-stage, and evidence\-driven nature of real scientific discovery\. As a result, they provide only a limited picture of whether a model can support realistic research workflows\. Benchmarks grounded in authentic scientific tasks are therefore essential for measuring true analytical ability while reducing the risk of contamination from standard educational data\.

Multimodality is equally indispensable\. Real scientific reasoning rarely depends on text alone\. Instead, it emerges from the joint interpretation of figures, spectra, microscopy images, tables, numerical measurements, and written context\. Without evaluating this convergence of visual, numerical, and textual evidence, benchmark scores remain detached from the actual demands of laboratory science\. A scientifically meaningful benchmark must therefore assess whether models can reason from heterogeneous evidence rather than from linguistic priors alone\.

Scientific benchmarking must also evolve alongside the development of natural science itself\. Historically, subjects such as physics, chemistry, biology, and medicine were often treated as separate domains\. However, in modern research, critical problems increasingly arise at their intersections, where concepts, methods, and data types are deeply entangled\. Thus, evaluating models only within isolated subjects misses a core requirement of real scientific reasoning, which is the ability to integrate knowledge across fields\. Interdisciplinary questions are also critical for exposing brittle reasoning, hallucination triggers, and transfer failures at the boundaries of a model’s knowledge\.

Taken together, these considerations reveal a central gap in current scientific evaluation\. Existing benchmarks often test whether models know scientific facts or solve simplified problems, whereas real laboratory science requires models to reason from incomplete, heterogeneous, and cross\-disciplinary experimental evidence\. This distinction is increasingly important as LLMs enter scientific workflows and agentic research systems\. Scientific evaluation should therefore move beyond knowledge recall toward evidence\-grounded, multimodal, interdisciplinary, and uncertainty\-aware reasoning\.

To address this gap, we introduceScience Edge Evaluation\(SEE\), a multimodal benchmark built from real scientific tasks in chemistry, biology, and materials science\. Our evaluation shows that current MLLMs remain unreliable in realistic experimental settings, and that their limitations cannot be explained by knowledge access alone\. Instead,SEEexposes a fundamental limitation in evidence\-based scientific reasoning\. Our tool\-augmented visual\-agent analysis further shows that tool access can expand evidence acquisition, but does not eliminate the need for reliable evidence management across the interaction trajectory\. Current models struggle to "SEE" multimodal observations across disciplines while maintaining rigorous adherence to the available evidence\. In this sense,SEEevaluates not only whether models know scientific content, but whether they can derive justified, evidence\-bounded insights from multimodal experimental data\.

## 2 Related Work

#### General Academic Benchmarks

Academic benchmarks are essential for evaluating LLM and MLLM capabilities\. General academic benchmarks such as MMLU\[[17](https://arxiv.org/html/2608.06931#bib.bib10)\], MMLU\-Pro\[[55](https://arxiv.org/html/2608.06931#bib.bib11)\], and GPQA\[[42](https://arxiv.org/html/2608.06931#bib.bib12)\]evaluate models across broad academic domains with different levels of difficulty\. The tested capabilities include general academic reasoning, scientific question answering, mathematical reasoning, and code generation\. Region\-specific extensions such as CMMLU\[[23](https://arxiv.org/html/2608.06931#bib.bib16)\]further broaden the evaluation surface\. To challenge models at the frontier of expert\-level reasoning, benchmarks such as HLE\[[39](https://arxiv.org/html/2608.06931#bib.bib13)\]have been specifically designed to incorporate closed\-ended questions of exceptional difficulty\.

#### Multimodal Benchmarks

Multimodal benchmarks such as ScienceQA\[[29](https://arxiv.org/html/2608.06931#bib.bib17)\]and MMMU\[[59](https://arxiv.org/html/2608.06931#bib.bib18)\]represent meaningful advances by extending evaluations beyond text\-only paradigms through the incorporation of visual inputs covering scientific and academic subjects, with MMMU\-Pro\[[60](https://arxiv.org/html/2608.06931#bib.bib72)\]further removing text\-solvable questions to enforce true multimodal reasoning\. Complementary efforts including MM\-Vet\[[58](https://arxiv.org/html/2608.06931#bib.bib14)\], MathVista\[[28](https://arxiv.org/html/2608.06931#bib.bib15)\], and SciFIBench\[[43](https://arxiv.org/html/2608.06931#bib.bib54)\]target integrated multimodal capabilities, mathematical reasoning in visual contexts, and scientific figure interpretation, respectively\. Olympiad\- and contest\-grade benchmarks such as OlympiadBench\[[16](https://arxiv.org/html/2608.06931#bib.bib73)\], the Chinese\-oriented MMSciBench\[[56](https://arxiv.org/html/2608.06931#bib.bib74)\], the multilingual MME\-SCI\[[44](https://arxiv.org/html/2608.06931#bib.bib75)\], and USNCO\[[10](https://arxiv.org/html/2608.06931#bib.bib78)\]built from chemistry\-olympiad exams push difficulty further with bilingual or multilingual multimodal scientific problems\. While current benchmarks are rooted in general academic and exam\-style contexts, they fail to capture the experimental data and workflow\-driven reasoning fundamental to authentic scientific inquiry\.

#### Scientific Benchmarks

Some recent benchmarks are designed to evaluate models in more specialized scientific settings, including chemistry, biology, and materials science\. These efforts focus on domain\-specific reasoning rather than broad academic knowledge, exemplified by ChemBench\[[33](https://arxiv.org/html/2608.06931#bib.bib20)\], SciBench\[[54](https://arxiv.org/html/2608.06931#bib.bib21)\], SciEval\[[49](https://arxiv.org/html/2608.06931#bib.bib53)\], and LAB\-Bench\[[21](https://arxiv.org/html/2608.06931#bib.bib22)\]\. A few of them incorporate multimodal inputs such as figures, molecular structures, spectra, and tables, together with experimental results\. For example, MaCBench\[[1](https://arxiv.org/html/2608.06931#bib.bib19)\]is a multimodal benchmark for chemistry and materials science that is closer to real research scenarios, while Matbench\[[11](https://arxiv.org/html/2608.06931#bib.bib49)\]targets materials property prediction\. More recent efforts further probe modality\-specific limitations, for example, ChemVTS\-Bench\[[18](https://arxiv.org/html/2608.06931#bib.bib77)\]disentangles visual, textual, and symbolic chemical reasoning\. Existing scientific benchmarks lack the disciplinary breadth, experimental realism, and diagnostic depth required to evaluate MLLMs against the complex, interdisciplinary challenges of authentic scientific research\.

#### Scientific LLMs and Agents

Scientific LLMs have been developed through continued pretraining, instruction tuning, and domain adaptation on scientific corpora\. Representative examples include general scientific and biomedical models such as Galactica\[[51](https://arxiv.org/html/2608.06931#bib.bib23)\], BioGPT\[[30](https://arxiv.org/html/2608.06931#bib.bib24)\], BioMedLM\[[5](https://arxiv.org/html/2608.06931#bib.bib25)\], GatorTronGPT\[[38](https://arxiv.org/html/2608.06931#bib.bib26)\], Med\-PaLM\[[47](https://arxiv.org/html/2608.06931#bib.bib27)\], and Meditron\[[9](https://arxiv.org/html/2608.06931#bib.bib28)\]; chemistry\-, molecule\-, and materials\-oriented models or instruction resources such as ChemLLM\[[61](https://arxiv.org/html/2608.06931#bib.bib55)\], ChemDFM\[[62](https://arxiv.org/html/2608.06931#bib.bib56)\], Mol\-Instructions\[[12](https://arxiv.org/html/2608.06931#bib.bib57)\], LlaSMol\[[57](https://arxiv.org/html/2608.06931#bib.bib60)\], HoneyBee\[[48](https://arxiv.org/html/2608.06931#bib.bib58)\], and MatterChat\[[50](https://arxiv.org/html/2608.06931#bib.bib59)\]; and multimodal biomedical or scientific models such as LLaVA\-Med\[[22](https://arxiv.org/html/2608.06931#bib.bib29)\], Med\-PaLM M\[[52](https://arxiv.org/html/2608.06931#bib.bib30)\], BioMedGPT\[[31](https://arxiv.org/html/2608.06931#bib.bib31)\], S1\-VL\[[24](https://arxiv.org/html/2608.06931#bib.bib96)\], Intern\-S2 Preview\[[20](https://arxiv.org/html/2608.06931#bib.bib98)\], Intern\-S2 Preview\-397B\[[19](https://arxiv.org/html/2608.06931#bib.bib99)\], and Intern\-S1 Pro\[[46](https://arxiv.org/html/2608.06931#bib.bib97)\]\. Recent scientific agents further extend LLMs from static question answering to tool\-augmented and workflow\-level scientific assistance, including ChemCrow\[[6](https://arxiv.org/html/2608.06931#bib.bib32)\], Coscientist\[[4](https://arxiv.org/html/2608.06931#bib.bib33)\], The AI Scientist\[[27](https://arxiv.org/html/2608.06931#bib.bib70)\], Agent Laboratory\[[45](https://arxiv.org/html/2608.06931#bib.bib34)\], and Co\-Scientist\[[15](https://arxiv.org/html/2608.06931#bib.bib35)\]\. While these efforts demonstrate the growing role of LLMs in scientific workflows, existing evaluations often focus on knowledge recall, domain\-specific task performance, tool\-use success, or final\-answer accuracy\. In contrast,SEEevaluates whether MLLMs can ground conclusions in multimodal experimental evidence, integrate concepts across disciplines, and reason within the boundaries of available evidence\.

## 3 Method

We organizeSEEthrough a staged construction and validation workflow, from real\-world data sourcing and expert question design to AI\-assisted quality control, expert review, and final benchmark acceptance \(Figure[1](https://arxiv.org/html/2608.06931#S3.F1)\)\.

![Refer to caption](https://arxiv.org/html/2608.06931v1/figures/benchmark_pipeline.png)Figure 1:Benchmark construction pipeline ofSEE\. Questions are derived from peer\-reviewed literature and expert experimental scenarios, standardized and calibrated with AI\-assisted checks, reviewed by domain experts, and accepted only after passing scientific accuracy, information sufficiency, answer uniqueness, and evaluation\-readiness checks\. Of the 1,116 questions, 1,049 are publicly released; the remaining 67 are withheld because they involve unpublished experimental data from contributing experts\.### 3\.1 Data Collection

SEEis a collaborative effort\. The questions are contributed by active frontline experts who hold master’s or PhD degrees or are PhD candidates in chemistry, biology, materials science, or related interdisciplinary areas\.

![Refer to caption](https://arxiv.org/html/2608.06931v1/x3.png)Figure 2:SEEconsists of 1,116 questions spanning 3 disciplines and 17 reported sub\-fields\. Interdisciplinary questions are represented by lines connecting two or more sub\-fields\. Line thickness indicates the number of co\-occurring questions across reported sub\-fields\. Gray lines represent cross\-discipline overlap \(spanning two disciplines\)\.#### Sources

The questions are drawn from real scientific research settings\. Sources include figures and data from peer\-reviewed papers, recent research literature, and first\-hand experimental data contributed by experts\. Visual evidence covers common experimental data structures in laboratory research, including bioactivity measurements, Cryo\-EM structures, Western blot and gel electrophoresis images, spectra such as infrared spectroscopy, nuclear magnetic resonance \(NMR\) and mass spectrometry \(MS\), microscopy images such as SEM, TEM and AFM, X\-ray diffraction patterns, and thermal analysis curves\.

#### Style

The questions inSEEinclude choice\-based, short\-answer, numerical, measurement, and image\-processing tasks\. Some questions contain explicit option lists, while others do not use a separate option field and are evaluated through canonical short answers or numerical answers\. The dataset also records task\-type metadata, including reasoning, measurement, and image\-processing categories\. Each entry contains text and associated visual files\. These associated images may include both visual evidence presented with the question and images used for expert verification\. As a result, the reported image counts are not the exact number of figures displayed in the question itself\. Each entry is reported with a standard answer to support automatic evaluation and expert verification\.

#### Labeling

Each question inSEEis assigned discipline labels spanning biological sciences, chemistry, and materials science\. The reported analysis uses a final taxonomy of 17 fine\-grained labels covering all 1,116 questions\. This labeling scheme captures cross\-disciplinary questions clearly\. Detailed labeling methods and label distributions are included in the Supplementary Information\.

### 3\.2 Data Processing

A multi\-stage review process is used to ensure data quality, difficulty, and reliability\. After collection and label assignment, the questions that do not meet our criteria are filtered out\.

#### Standardization

The raw questions collected are standardized in terms of wording, terminology, answer structures, option numbering, numerical precision, tolerance records, and image organization\. The images are normalized in format and resolution\.

#### General Quality Control

After standardization, questions are checked for adherence to the required format, clear and objective answers, sufficient visual and textual evidence, and correct labels\. We remove semantic duplicates and questions that do not fulfill the requirements\. Each question is then cross\-checked by at least two experts in the relevant domain\.

![Refer to caption](https://arxiv.org/html/2608.06931v1/x4.png)Figure 3:Representative multimodal scientific questions inSEE, showing the associated figure, question stem, answer options, correct answer, discipline/technique tags, and data source\.
#### Expert Quality Control

After AI\-assisted checks, experts review the questions in terms of scientific accuracy\. Questions with problems are returned to their contributors for revision\. The revised questions then undergo the entire standardization and quality control process before inclusion\.

#### Image\-Ablation Checking

Finally, questions with images are checked to verify whether images are indispensable for solving the task\. Questions that can be answered from text alone are removed or revised\. This verification is consistent with previous multimodal benchmark practices\[[59](https://arxiv.org/html/2608.06931#bib.bib18),[43](https://arxiv.org/html/2608.06931#bib.bib54)\]\.

#### Question Evaluation

The evaluation of questions follows the standard scoring protocols established in large\-scale evaluation frameworks\[[25](https://arxiv.org/html/2608.06931#bib.bib80)\]\. For multiple choice questions, MLLM answers must*exactly*match the ground\-truth answers \(including single\- and multiple\-response questions\)\. Numerical fill\-in\-the\-blank questions are evaluated with tolerance level provided by experts when applicable\. Non\-numerical fill\-in\-the\-blank questions are evaluated against standardized canonical answers\.

Representative examples of the questions collected are shown in Figure[3](https://arxiv.org/html/2608.06931#S3.F3), illustrating within\-discipline and cross\-disciplinary multimodal scientific questions\.

### 3\.3 Model Evaluation Environment

We evaluate 19 representative MLLMs withSEE, including 15 general\-purpose models and four science\-specialized models\. The general\-purpose models include Gemini 3\.1 Pro\[[13](https://arxiv.org/html/2608.06931#bib.bib81)\], GPT\-5\.6\-Sol \(Max\)\[[37](https://arxiv.org/html/2608.06931#bib.bib84)\], Claude Opus 5 \(Max\)\[[3](https://arxiv.org/html/2608.06931#bib.bib88)\], Qwen3\.8\-Max\[[41](https://arxiv.org/html/2608.06931#bib.bib86)\], GPT\-5\.5 \(xhigh\)\[[36](https://arxiv.org/html/2608.06931#bib.bib83)\], Kimi K3\[[35](https://arxiv.org/html/2608.06931#bib.bib92)\], Gemini 3\.5 Flash\[[14](https://arxiv.org/html/2608.06931#bib.bib82)\], Claude Opus 4\.8 \(Max\)\[[2](https://arxiv.org/html/2608.06931#bib.bib87)\], Seed2\.1 Pro\[[8](https://arxiv.org/html/2608.06931#bib.bib90)\], Qwen3\.7\-Plus\[[40](https://arxiv.org/html/2608.06931#bib.bib85)\], Seed2\.0 Pro\[[7](https://arxiv.org/html/2608.06931#bib.bib89)\], MiniMax\-M3\[[32](https://arxiv.org/html/2608.06931#bib.bib93)\], Kimi K2\.6\[[34](https://arxiv.org/html/2608.06931#bib.bib91)\], GLM\-5V\-Turbo\[[64](https://arxiv.org/html/2608.06931#bib.bib94)\], and MiMo\-V2\.5\[[26](https://arxiv.org/html/2608.06931#bib.bib95)\]\. The science\-specialized models include S1\-VL\[[24](https://arxiv.org/html/2608.06931#bib.bib96)\], Intern\-S2 Preview\[[20](https://arxiv.org/html/2608.06931#bib.bib98)\], Intern\-S2 Preview\-397B\[[19](https://arxiv.org/html/2608.06931#bib.bib99)\], and Intern\-S1 Pro\[[46](https://arxiv.org/html/2608.06931#bib.bib97)\]\. Only four science\-specialized models are tested because few models in this category possess the multimodal capabilities required for our evaluation\.

Unless otherwise specified, models are evaluated with default inference configurations\. The only non\-default settings are the fixed extended\-thinking configurations used for GPT\-5\.5 \(xhigh\) and Claude Opus 4\.8 \(Max\), as detailed in Supplementary Information[B\.1](https://arxiv.org/html/2608.06931#A2.SS1)\. For image\-containing questions, we apply the preprocessing required by each model interface, including format conversion, resizing, and input organization\. We introduce no manual prompt correction, per\-question parameter tuning, or result filtering\. Claude Opus 4\.8 \(Max\) and Claude Opus 5 \(Max\) have safety\-filtered no\-response cases on biochemistry questions; under the strict binary scoring protocol, these instances are counted as incorrect\.

Model outputs are evaluated with a strict binary LLM\-as\-a\-judge pipeline\[[63](https://arxiv.org/html/2608.06931#bib.bib79)\]\. Each response is extracted and marked as correct or incorrect using Gemini 3\.1 Pro as the primary judge model for downstream accuracy computation and breakdown analysis\. We also test GPT\-5\.5 as an alternative judge model and obtain similar results, indicating that the evaluation is robust to the choice of judge model \(Cohen’sκ≥0\.99\\kappa\\geq 0\.99; Supplementary Information[B\.4](https://arxiv.org/html/2608.06931#A2.SS4)\)\. Refusals, non\-answers, and responses from which no final answer can be identified are counted as incorrect\. Evaluation and judging prompts are detailed in Supplementary Information[B\.2](https://arxiv.org/html/2608.06931#A2.SS2), and subject\-label coverage and cross\-discipline composition ofSEEare reported in Supplementary Information[B\.9](https://arxiv.org/html/2608.06931#A2.SS9)\.

## 4 Results and Discussion

### 4\.1 Overall Performance on SEE

Following the evaluation protocol described in Section[3\.3](https://arxiv.org/html/2608.06931#S3.SS3), we first report the overall accuracy of 19 representative MLLMs onSEE\. All evaluated models achieve low accuracy onSEEin the standard evaluation setting \(Figure[4](https://arxiv.org/html/2608.06931#S4.F4)\)\. The best model, GPT\-5\.6\-Sol \(Max\), reaches 48\.7% accuracy, and no model exceeds 50%\. Across the 19 models, accuracy ranges from 15\.9% to 48\.7%\.

![Refer to caption](https://arxiv.org/html/2608.06931v1/x5.png)Figure 4:Accuracies of different models onSEE\. General\-purpose models are shown as blue bars, while science\-specialized models are shown in brown\. The dashed vertical lines indicate average accuracies in the two categories\. Claude Opus 4\.8 \(Max\) and Claude Opus 5 \(Max\) sometimes give no response to questions involving viral biology and pathogen structural characterization due to safety reasons \(marked with \*\)\. These answers are counted as incorrect\.The result at the discipline level shows similar aggregate precision in all three main disciplines \(Table[8](https://arxiv.org/html/2608.06931#A2.T8)in Supplementary Information[B\.8](https://arxiv.org/html/2608.06931#A2.SS8)\)\. Across all evaluated models, average accuracy is 32\.5% for chemistry, 31\.2% for biology, and 29\.0% for materials science\.

Performance differences become clearer at the sub\-discipline level \(Figure[5](https://arxiv.org/html/2608.06931#S4.F5); full per\-model breakdown in Table[8](https://arxiv.org/html/2608.06931#A2.T8)\)\. The lowest mean accuracies occur in organic and polymer materials \(25\.6%\), polymer chemistry and physics \(28\.7%\), composite materials \(28\.7%\), and inorganic chemistry \(28\.3%\)\. The highest reported sub\-discipline accuracy is organic chemistry \(37\.9%\), followed by metallic materials \(34\.8%\), analytical chemistry \(34\.5%\), and immunology \(34\.0%\)\. Sub\-disciplines with small sample sizes should be interpreted with care\.

The radar plot shows that no single model demonstrates universal dominance throughout scientific spectrum\. However, proficiency remains highly domain\-contingent\. Additionally, we penalized Claude Opus 4\.8 and Claude Opus 5 for a subset of safety\-related refusals \(marked with asterisks in the related figures and tables\), which are triggered exclusively in the biochemistry section by questions involving virus–host interactions, pathogen structural biology, and related experimental techniques \(e\.g\., cryo\-EM, immunoblotting, viral neutralization assays\), to ensure a fair comparison across all models\.

![Refer to caption](https://arxiv.org/html/2608.06931v1/x6.png)Figure 5:Results by discipline of the 6 overall strongest models are summarized in the radar plot\. Each axis reports accuracy on one reported sub\-discipline label\. Axis labels and spokes are colored by broad discipline\. Concentric rings mark 20%, 40%, 60%, 80%, and 100% accuracy respectively\. The red ring highlights the 60% reference level\. The asterisk on Claude Opus 5 \(Max\) indicates safety\-filtered no\-response cases in biochemistry, which are counted as incorrect\.
### 4\.2 Test on Science\-Specialized Models

We next compare general\-purpose and science\-specialized MLLMs to examine whether domain adaptation improves performance onSEE\. Within the tested model set, the four science\-specialized models remain below the general\-purpose average onSEE\(Figure[4](https://arxiv.org/html/2608.06931#S4.F4)\), with an average accuracy of 22\.2% compared to 34\.9% for the general\-purpose models\. However, Intern\-S2 Preview\-397B reaches 29\.0%, outperforming several mid\-tier general\-purpose models and narrowing the gap relative to earlier science\-specialized models\.

We observe that the phenomenon of leading general\-purpose models outperforming domain\-specific models is not unique to our setting, but also appears in other fields such as clinical medicine\[[53](https://arxiv.org/html/2608.06931#bib.bib100)\]\. This suggests that simply specializing a model through post\-training or augmenting it with retrieval\-augmented generation \(RAG\) does not necessarily lead to improved performance\.

These results suggest that success onSEErequires more than domain\-specific factual familiarity\. Science\-specialized training may improve exposure to scientific terminology, concepts, and task formats, butSEErequires models to combine such knowledge with visual experimental evidence and specific experimental context\. The gap between science\-specialized and leading general\-purpose models therefore indicates that robust scientific reasoning in real experimental settings depends not only on domain adaptation, but also on broader multimodal understanding and flexible evidence\-grounded reasoning\.

### 4\.3 Interdisciplinary Analysis

As noted above, greater exposure to a specific scientific domain does not necessarily translate into reliable reasoning across broader scientific contexts\. In practice, scientific discovery often crosses disciplinary boundaries and requires integrating knowledge, assumptions, and observations from multiple fields\. To reflect this reality,SEEexplicitly includes interdisciplinary questions spanning biology, chemistry, and materials science\. This design is central to the benchmark, whose goal is not only to test whether models know isolated scientific facts, but also to evaluate whether they can integrate heterogeneous evidence across disciplinary contexts\.

We first examine the performance gap between single\-discipline and cross\-discipline questions\. Across all models, the average accuracy is 35\.1% on questions within a single main discipline, but drops to 28\.7% on questions that span multiple main disciplines\. This 6\.4% gap suggests that cross\-disciplinary settings impose additional coordination demands beyond those captured by single\-discipline evaluation\.

We next examine cross\-subdisciplinary complexity\. Averaged across all models, performance remains similar for questions annotated with one or two subdisciplinary labels, with accuracies of 36\.2% and 35\.6%, respectively\. However, the accuracy drops to 29\.5% for questions with three labels and further to 28\.2% for questions with four labels\. This pattern suggests that performance degrades as a task requires coordination across a larger number of scientific concepts, experimental techniques, or forms of evidence\.

Together, these results show that interdisciplinary tasks provide a stringent stress test of model robustness\. Such tasks require models to integrate scientific terminology, experimental assumptions, visual evidence, and multistep reasoning within a single response\. The interdisciplinary design ofSEEtherefore assesses whether current MLLMs can coordinate knowledge and evidence across the boundaries commonly encountered in real\-world research\. To be considered robust in a research setting, a model must maintain coherent reasoning even when a problem falls outside the distributions most frequently represented in training\.

### 4\.4 Modality Ablation Study

![Refer to caption](https://arxiv.org/html/2608.06931v1/x7.png)Figure 6:Explicit missing\-image acknowledgment rates in the text\-only ablation test\. The denominator is the number of text\-only instances for each model\. The numerator is the number of responses that explicitly recognize missing image or visual evidence\. Empty outputs, generation errors, generic refusals, and other non\-answer cases are not counted\. The dashed vertical line marks the average rate across the models\. Asterisks denote safety\-filtered no\-response cases counted under the strict evaluation protocol\.To evaluate the level of hallucination of MLLMs by investigating their ability to recognize missing information, we performed a blind\-modality diagnostic by removing image inputs while preserving the original textual queries across 18 models\. This experiment serves as a critical test for grounding\. A truly intelligent system should identify when a scientific conclusion is impossible without visual evidence\. Across 20,088 text\-only instances, models identified missing information in only 921 cases \(4\.6%, Figure[6](https://arxiv.org/html/2608.06931#S4.F6)\)\. This remarkably low rate suggests that current MLLMs are prone to blind hallucination, attempting to derive scientific answers from linguistic priors rather than admitting a lack of empirical evidence\. Removing the image also lowers accuracy for every model, by 12\.2 percentage points on average; this and further text\-only diagnostics are reported in Supplementary Information[B\.7](https://arxiv.org/html/2608.06931#A2.SS7)\(Figure[9](https://arxiv.org/html/2608.06931#A2.F9)\)\.

This result reveals limited awareness of evidential boundaries in current MLLMs\. When visual inputs are removed, models often continue to answer based on textual cues, prior knowledge, or common scientific patterns, rather than recognizing that the available evidence is insufficient\. This behavior is especially concerning in real\-world scientific settings, where decisions are often made under incomplete or ambiguous information and where unsupported but confident answers may pose greater risks than appropriate abstention\. Thus, this ablation analysis provides a diagnostic signal that current MLLMs do not yet consistently reason within the limits of experimental evidence\.

### 4\.5 Tool\-Augmented Visual\-Agent Evaluation

![Refer to caption](https://arxiv.org/html/2608.06931v1/x8.png)Figure 7:Comparison between standard multimodal inference and the tool\-augmented visual\-agent setting, in which models can iteratively use web search and a code interpreter before returning a final answer\. Paired bars show the accuracy of each model under the two settings, and numbers on the right report the corresponding difference in percentage points\. The asterisk on Claude Opus 5 \(Max\) denotes no\-response cases, which are counted as incorrect\.Finally, we test whether tool access is sufficient to narrow the capability gap exposed bySEE\. We evaluate the six top\-performing MLLMs that support both web search and code\-interpreter tools in a tool\-augmented visual\-agent setting \(Figure[7](https://arxiv.org/html/2608.06931#S4.F7)\)\. Models receive the same original multimodal inputs as in the standard baseline, but can invoke a web search and a code interpreter before producing a final answer\. The code interpreter supports programmatic inspection, measurement, and computation over the input image, whereas the web search provides external scientific background information\. In this setting, all six models improve, with gains ranging from 1\.9% to 7\.6% and a mean improvement of 4\.6%\. GPT\-5\.5 \(xhigh\) shows the largest improvement, increasing from 41\.3% to 48\.9%\. GPT\-5\.6\-Sol \(Max\) achieves the highest tool\-augmented accuracy of 52\.7%\. Nevertheless, significant errors remain, indicating that tool access alone does not make MLLMs reliable onSEE\. Detailed configuration is provided in the Supplementary Information[B\.5](https://arxiv.org/html/2608.06931#A2.SS5)\.

![Refer to caption](https://arxiv.org/html/2608.06931v1/x9.png)Figure 8:Effects of using tools in tool\-augmented visual agents\. \(a\) Illustration of tool effects on stages of tool\-mediated visual\-agent reasoning\. Corrections are made when models select relevant actions, use tools to inspect or quantify visual evidence, retrieve key external facts, or verify intermediate interpretations\. Failures can arise from wrong action selection, misinterpretation of observations, inappropriate integration of retrieved or computed information, or failure to control and terminate the interaction\. The dashed loop indicates that unresolved uncertainty may trigger further tool use\. \(b\) Cases for every failed scenario\.Trajectory analysis shows that tool use has bidirectional effects on model outputs \(Figure[8](https://arxiv.org/html/2608.06931#S4.F8)\)\. Across the six models, 411 corrections are substantively attributed to tool use\. Fact retrieval accounts for 47\.7% of these cases, code\-assisted image analysis for 41\.6%, information cross\-checking for 7\.1%, and code\-assisted computation for 3\.6%\. These results show that tools can expand the models’ ability to inspect visual evidence and access background information, partially compensating for the limitations of static multimodal inference\.

Tools can also introduce new errors\. Among 145 tool\-attributed errors, failure modes gather in evidence integration failures \(40\.0%\), action selection failures \(26\.9%\), observation interpretation failures \(21\.4%\), and endless overthinking \(11\.7%\)\. Both beneficial and detrimental effects of tools occur in the tool\-use trajectory \(Figure[8](https://arxiv.org/html/2608.06931#S4.F8)\)\. This indicates that tool augmentation does not simply provide models with additional information, but turns static multimodal reasoning into an evidence\-management problem along the tool\-use trajectory\. Models must decide which actions are relevant to the current problem, whether tool outputs are reliable, how retrieved or computed information should be integrated with the original experimental observation, and when to stop further interaction\.

## 5 Model Failure Analysis

### 5\.1 Weakness in Perception

#### Limited Visual Quantitative Precision

A substantial proportion of model failures are due to inadequate visual reading precision \(tasks requiring the extraction of exact numerical values from spectra, standard curves, or instrument readouts\)\. These failures expose that the visual encoders of current MLLMs are pattern recognizers rather than measurement instruments\. They excel at categorical judgments, but perform poorly at fine\-grained quantitative extraction\. This perceptual inaccuracy can propagate downstream, causing the model to arrive at incorrect conclusions even when it possesses sound domain knowledge and applies valid reasoning logic\. For scientific applications where quantitative precision is non\-negotiable, this limitation constitutes a fundamental bottleneck that cannot be resolved by improvements to reasoning alone\.

#### Prior Knowledge Suppressing Visual Input

When visual evidence presented in an image conflicts with patterns prevalent in the training data, models systematically defer to prior knowledge rather than faithfully attending to the actual visual input\. Rather than genuinely "reading" an image, models tend to infer its content by reverse\-mapping from the most frequently encountered concepts in training\. This pattern is particularly noticeable in tasks involving molecular structure interpretation typical in biology: models directly match surface\-level visual features to high\-frequency terms encountered during training instead of decomposing the structure, identifying functional fragments, and progressively deriving the corresponding name or property as a human expert would\. This pattern matching shortcut fails when confronted with atypical or non\-canonical structures\. Current alignment methods produce models that match patterns rather than observe structures\. They excel at associating frequent labels with common visuals but lack the bottom\-up logic necessary to parse the unknown\.

### 5\.2 Weakness in Inference

Beyond perceptual limitations, our analysis identifies three distinct inferential failure modes that reflect deeper architectural constraints in the way current MLLMs construct and evaluate logical representations\.

#### Neglect of Global Logical Coherence

A recurring failure mode is the model’s inability to maintain a coherent logical framework, often processing question components as disparate units rather than an integrated system\. There are two typical patterns\. First, models frequently overlook intra\-option contradictions\. If an option pairs two factually correct but logically incompatible premises, the model tends to evaluate them independently, missing the overall information\. Second, there is a clear deficiency in multi\-modal integration\. When a problem relies on the interplay between several subfigures, models often perform a localized search within one subfigure while neglecting the broader logical constraints established by the remaining figures\. This fragmented processing prevents the model from forming a comprehensive relational structure, leading to conclusions that are locally plausible but globally invalid\.

#### Overextrapolation of Experimental Conclusions

Another critical failure mode is the tendency to over\-infer from partial experimental data\. For example, an inquiry requires the synthesis of several independent experiments to reach a conclusion, the current models often treat preliminary or isolated results as conclusive evidence\. Models frequently interpolate "missing" experimental steps and invent evidence to justify their final claims\. This tendency to overextrapolate from individual observations undermines the rigorous evidentiary standards essential to valid scientific inference\.

#### Multi\-Capability Coordination Bottleneck

Many scientific challenges require the coordinated operation of multiple cognitive skills, and failure rates escalate sharply when faced with these demands\. NMR spectral interpretation, for instance, requires a model to utilize precise visual perception, deep domain knowledge, and complex logical reasoning simultaneously\. Because current MLLMs possess individual deficiencies in each area, the requirement for their serial execution leads to a multiplicative compounding of error\. This explains why NMR\-related tasks show disproportionately high failure rates across all architectures\. In summary, failures on these tasks do not arise from a single weakness, but from the model’s inability to reliably coordinate multiple interdependent capabilities\.

Collectively, these failure modes underscore a fundamental contradiction: while the essence of scientific inquiry lies in the ability to make rigorous judgments regarding out\-of\-distribution \(OOD\) phenomena, current models systematically regress to high\-frequency patterns when confronted with unfamiliar evidence\. Whether by prioritizing prior knowledge over visual observation or substituting "textbook" conclusions for incomplete evidence chains, models gravitate toward the path of least statistical resistance\. This suggests that merely expanding the training database is insufficient, as the frontier of scientific discovery, by definition, resides at the boundary of available data\.

## 6 Conclusion

We introducedSEEto evaluate if current MLLMs can support real laboratory science by reasoning from multimodal experimental evidence\. Unlike benchmarks that primarily measure scientific knowledge recall or isolated domain competence,SEEis based on peer\-reviewed literature and experimental practice in chemistry, biology, and materials science\. Under the standard inference setting, the best\-performing model among the 19 MLLMs evaluated reaches only 48\.7% accuracy, and no model exceeds 50%\. This low performance should not be interpreted simply as another hard benchmark result\. Instead, it reveals a mismatch between the abilities captured by many existing scientific benchmarks and the abilities required for real reasoning process in experimental research\.

The additional analyses clarify the nature of this mismatch\. Science\-specialized models do not close the performance gap\. Meanwhile, interdisciplinary analysis shows thatSEErequires models to coordinate concepts, methods, and different types of evidence across disciplinary boundaries\. Ablation analysis reveals that when visual evidence is removed, MLLMs rarely recognize that essential information is missing\. Together, these findings suggest that the central limitation is not simply what the models know but whether they can extract and organize information from the evidence that is actually available in the problem\.

The tool\-augmented visual\-agent experiment extends this conclusion from static multimodal inference to interactive scientific workflows\. Tool access boosts accuracy to 52\.7% \(strongest model\), showing that code interpreter and web search can partially compensate for the limitations of multimodal reasoning\. However, trajectory analysis reveals that tool use is not always beneficial\. The interaction process can help models inspect visual evidence through code\-assisted image analysis, retrieve missing scientific facts, cross\-check information, or perform code\-assisted computation\. Yet it can also introduce new errors through wrong action selection, misinterpretation of tool\-derived observations, evidence integration failure, or endless overthinking\. The diagnosis of tool\-use trajectory provides an explicit view of tool\-using in scientific reasoning rather than only the overall accuracy\. The main challenge is therefore whether a model can manage tool\-derived information by selecting appropriate actions, judging the reliability and relevance of tool outputs, integrating retrieved or computed result with the original information, and stopping when the available evidence is sufficient\.

The failure analysis explains why this limitation matters for scientific reasoning\. Current MLLMs often miss critical visual evidence, allow prior knowledge to override experimental observations, fail to maintain global logical coherence, and draw conclusions beyond what the evidence supports\. These errors are not trivial\. These deficiencies highlight a fundamental gap between pattern recognition and the rigorous, evidence\-based logic required for authentic scientific inquiry\.

These findings suggest a new stage for scientific evaluation where the focus shifts from single\-response MLLMs toward visual agents and broader agentic scientific systems\. Real discovery is rarely completed by answering one question from a fixed input\. Instead, it requires an iterative process of proposing hypotheses, selecting tools, transforming or measuring observations, analyzing experimental outputs, and deciding which evidence should be collected next\. Future benchmarks for scientific agents should therefore evaluate not only accuracy, but also the quality of the interaction trajectory\. This includes assessing whether each action is justified by the available evidence, whether image operations, retrieval, and data analysis are appropriate, whether uncertainty is recognized before the next step, and whether conclusions are revised when new observations contradict prior assumptions\. Therefore, agent\-level evaluation should extend the evidence\-based principle ofSEEfrom static experimental reasoning to dynamic scientific decision\-making\.

Overall,SEEshows that the main barrier to scientifically useful AI is not simply more knowledge, stronger domain specialization, or broader tool access, but the ability to manage multimodal evidence and make justified evidence\-bounded inferences from experimental results\. The tool\-augmented visual\-agent analysis shows that this challenge persists in interactive workflows, where tools expand evidence access but also require reliable action selection, output evaluation, evidence integration, and termination control\. The failure analysis reveals the same limitations\. Current MLLMs miss critical observations, let prior knowledge override evidence, lose global coherence, and overextend conclusions beyond the data\. Future evaluations should therefore test not only what models know, but whether they can transform multimodal experimental observations into new, evidence\-supported scientific understanding\. This is the missing step toward real scientific discovery\.

## Data Availability

To facilitate the benchmarking and reproducibility of our work, the accompanying code and public datasets are available on[GitHub](https://github.com/SKYLENAGE-AI/science-edge-evaluation)\. The main benchmark contains 1,116 questions, of which 1,049 are publicly released\. To support reproducibility, we also report headline results on the public subset in Supplementary Tables[5](https://arxiv.org/html/2608.06931#A2.T5)and[6](https://arxiv.org/html/2608.06931#A2.T6)\.

## References

## References

- \[1\]\(2025\)Probing the limitations of multimodal language models for chemistry and materials research\.Nature Computational Science\.Note:MaCBenchExternal Links:[Document](https://dx.doi.org/10.1038/s43588-025-00836-3)Cited by:[§2](https://arxiv.org/html/2608.06931#S2.SS0.SSS0.Px3.p1.1)\.
- \[2\]Anthropic\(2026\)Claude Opus 4\.8 \(Max\)\.Note:Official announcementExternal Links:[Link](https://www.anthropic.com/news/claude-opus-4-8)Cited by:[§3\.3](https://arxiv.org/html/2608.06931#S3.SS3.p1.1)\.
- \[3\]Anthropic\(2026\)Claude Opus 5\.Note:Official announcementExternal Links:[Link](https://www.anthropic.com/news/claude-opus-5)Cited by:[§3\.3](https://arxiv.org/html/2608.06931#S3.SS3.p1.1)\.
- \[4\]D\. A\. Boiko, R\. MacKnight, B\. Kline, and G\. Gomes\(2023\)Autonomous chemical research with large language models\.Nature624\(7992\),pp\. 570–578\.External Links:[Document](https://dx.doi.org/10.1038/s41586-023-06792-0)Cited by:[§2](https://arxiv.org/html/2608.06931#S2.SS0.SSS0.Px4.p1.1)\.
- \[5\]E\. Bolton, A\. Venigalla, M\. Yasunaga, D\. Hall, B\. Xiong, T\. Lee, R\. Daneshjou, and J\. Frankle\(2024\)BioMedLM: a 2\.7b parameter language model trained on biomedical text\.External Links:2403\.18421,[Link](https://arxiv.org/abs/2403.18421)Cited by:[§2](https://arxiv.org/html/2608.06931#S2.SS0.SSS0.Px4.p1.1)\.
- \[6\]A\. M\. Bran, S\. Cox, O\. Schilter, C\. Baldassari, A\. D\. White, and P\. Schwaller\(2024\)Augmenting large language models with chemistry tools\.Nature Machine Intelligence6\(5\),pp\. 525–535\.Note:ChemCrowExternal Links:[Document](https://dx.doi.org/10.1038/s42256-024-00832-8)Cited by:[§2](https://arxiv.org/html/2608.06931#S2.SS0.SSS0.Px4.p1.1)\.
- \[7\]ByteDance Seed\(2026\)Seed2\.0 Pro\.Note:Official product pageExternal Links:[Link](https://seed.bytedance.com/zh/seed2)Cited by:[§3\.3](https://arxiv.org/html/2608.06931#S3.SS3.p1.1)\.
- \[8\]ByteDance Seed\(2026\)Seed2\.1 Pro\.Note:Official product pageExternal Links:[Link](https://seed.bytedance.com/en/seed2_1)Cited by:[§3\.3](https://arxiv.org/html/2608.06931#S3.SS3.p1.1)\.
- \[9\]Z\. Chen, A\. Hernández Cano, A\. Romanou, A\. Bonnet, K\. Matoba, F\. Salvi, M\. Pagliardini, S\. Fan,et al\.\(2023\)MEDITRON\-70B: scaling medical pretraining for large language models\.External Links:2311\.16079,[Link](https://arxiv.org/abs/2311.16079)Cited by:[§2](https://arxiv.org/html/2608.06931#S2.SS0.SSS0.Px4.p1.1)\.
- \[10\]Y\. Cui, X\. Yao, Y\. Qin, X\. Li, S\. Wang, and G\. Hu\(2025\)Evaluating large language models on multimodal chemistry olympiad exams\.Communications Chemistry8,pp\. 402\.Note:USNCO\-V benchmarkExternal Links:[Document](https://dx.doi.org/10.1038/s42004-025-01782-x)Cited by:[§2](https://arxiv.org/html/2608.06931#S2.SS0.SSS0.Px2.p1.1)\.
- \[11\]A\. Dunn, Q\. Wang, A\. Ganose, D\. Dopp, and A\. Jain\(2020\)Benchmarking materials property prediction methods: the Matbench test set and Automatminer reference algorithm\.npj Computational Materials6\(1\),pp\. 138\.External Links:[Document](https://dx.doi.org/10.1038/s41524-020-00406-3)Cited by:[§2](https://arxiv.org/html/2608.06931#S2.SS0.SSS0.Px3.p1.1)\.
- \[12\]Y\. Fang, X\. Liang, N\. Zhang, K\. Liu, R\. Huang, Z\. Chen, X\. Fan, and H\. Chen\(2024\)Mol\-Instructions: a large\-scale biomolecular instruction dataset for large language models\.InInternational Conference on Learning Representations \(ICLR\),External Links:[Link](https://openreview.net/forum?id=Tlsdsb6l9n),2306\.08018Cited by:[§2](https://arxiv.org/html/2608.06931#S2.SS0.SSS0.Px4.p1.1)\.
- \[13\]Google DeepMind\(2026\)Gemini 3\.1 Pro\.Note:Official model pageExternal Links:[Link](https://deepmind.google/models/gemini/pro/)Cited by:[§3\.3](https://arxiv.org/html/2608.06931#S3.SS3.p1.1)\.
- \[14\]Google DeepMind\(2026\)Gemini 3\.5 Flash\.Note:Official model pageExternal Links:[Link](https://deepmind.google/models/gemini/flash/)Cited by:[§3\.3](https://arxiv.org/html/2608.06931#S3.SS3.p1.1)\.
- \[15\]J\. Gottweis, W\. Weng, A\. Daryin, T\. Tu, P\. Sirkovic, A\. Myaskovsky, G\. Glowaty, F\. Weissenberger, A\. Orlandi,et al\.\(2026\)Accelerating scientific discovery with Co\-Scientist\.Nature\.External Links:[Document](https://dx.doi.org/10.1038/s41586-026-10644-y),[Link](https://www.nature.com/articles/s41586-026-10644-y)Cited by:[§2](https://arxiv.org/html/2608.06931#S2.SS0.SSS0.Px4.p1.1)\.
- \[16\]C\. He, R\. Luo, Y\. Bai, S\. Hu, Z\. L\. Thai, J\. Shen, J\. Hu, X\. Han, Y\. Huang, Y\. Zhang, J\. Liu, L\. Qi, Z\. Liu, and M\. Sun\(2024\)OlympiadBench: a challenging benchmark for promoting AGI with olympiad\-level bilingual multimodal scientific problems\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(ACL\),pp\. 3828–3850\.External Links:[Link](https://aclanthology.org/2024.acl-long.211/)Cited by:[§2](https://arxiv.org/html/2608.06931#S2.SS0.SSS0.Px2.p1.1)\.
- \[17\]D\. Hendrycks, C\. Burns, S\. Basart, A\. Zou, M\. Mazeika, D\. Song, and J\. Steinhardt\(2021\)Measuring massive multitask language understanding\.InInternational Conference on Learning Representations \(ICLR\),External Links:[Link](https://openreview.net/forum?id=d7KBjmI3GmQ),2009\.03300Cited by:[§2](https://arxiv.org/html/2608.06931#S2.SS0.SSS0.Px1.p1.1)\.
- \[18\]Z\. Huang, B\. Yang, Z\. He, Y\. Wu, H\. Fang, Z\. Liu, D\. Lin, and B\. Su\(2025\)ChemVTS\-Bench: evaluating visual–textual–symbolic reasoning of multimodal large language models in chemistry\.External Links:2511\.17909,[Link](https://arxiv.org/abs/2511.17909)Cited by:[§2](https://arxiv.org/html/2608.06931#S2.SS0.SSS0.Px3.p1.1)\.
- \[19\]InternLM\(2026\)Intern\-S2 Preview\-397B\.Note:Official model pageExternal Links:[Link](https://huggingface.co/internlm/Intern-S2-Preview-397B)Cited by:[§2](https://arxiv.org/html/2608.06931#S2.SS0.SSS0.Px4.p1.1),[§3\.3](https://arxiv.org/html/2608.06931#S3.SS3.p1.1)\.
- \[20\]InternLM\(2026\)Intern\-S2 Preview\.Note:Official model pageExternal Links:[Link](https://huggingface.co/internlm/Intern-S2-Preview)Cited by:[§2](https://arxiv.org/html/2608.06931#S2.SS0.SSS0.Px4.p1.1),[§3\.3](https://arxiv.org/html/2608.06931#S3.SS3.p1.1)\.
- \[21\]J\. M\. Laurent, J\. D\. Janizek, M\. Ruzo, M\. M\. Hinks, M\. J\. Hammerling, S\. Narayanan, M\. Ponnapati, A\. D\. White, and S\. G\. Rodriques\(2024\)LAB\-Bench: measuring capabilities of language models for biology research\.External Links:2407\.10362,[Link](https://arxiv.org/abs/2407.10362)Cited by:[§2](https://arxiv.org/html/2608.06931#S2.SS0.SSS0.Px3.p1.1)\.
- \[22\]C\. Li, C\. Wong, S\. Zhang, N\. Usuyama, H\. Liu, J\. Yang, T\. Naumann, H\. Poon, and J\. Gao\(2023\)LLaVA\-Med: training a large language\-and\-vision assistant for biomedicine in one day\.InAdvances in Neural Information Processing Systems,Vol\.36,pp\. 28541–28564\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2023/hash/5abcdf8ecdcacba028c6662789194572-Abstract-Datasets_and_Benchmarks.html),2306\.00890Cited by:[§2](https://arxiv.org/html/2608.06931#S2.SS0.SSS0.Px4.p1.1)\.
- \[23\]H\. Li, Y\. Zhang, F\. Koto, Y\. Yang, H\. Zhao, Y\. Gong, N\. Duan, and T\. Baldwin\(2023\)CMMLU: measuring massive multitask language understanding in Chinese\.External Links:2306\.09212,[Link](https://arxiv.org/abs/2306.09212)Cited by:[§2](https://arxiv.org/html/2608.06931#S2.SS0.SSS0.Px1.p1.1)\.
- \[24\]Q\. Li, L\. Xu, Q\. Wang, Y\. Bai, M\. Ou, S\. Hu, and N\. Xu\(2026\)S1\-VL: scientific multimodal reasoning model with thinking\-with\-images\.Note:We evaluate the S1\-VL\-32B\-RL checkpointExternal Links:2604\.21409,[Link](https://arxiv.org/abs/2604.21409)Cited by:[§2](https://arxiv.org/html/2608.06931#S2.SS0.SSS0.Px4.p1.1),[§3\.3](https://arxiv.org/html/2608.06931#S3.SS3.p1.1)\.
- \[25\]P\. Liang, R\. Bommasani, T\. Lee, D\. Tsipras, D\. Soylu, M\. Yasunaga, Y\. Zhang, D\. Narayanan, Y\. Wu, A\. Kumar, B\. Newman, B\. Yuan, B\. Yan, C\. Zhang, C\. Cosgrove, C\. D\. Manning, C\. Ré, D\. Acosta\-Navas, D\. A\. Hudson,et al\.\(2023\)Holistic evaluation of language models\.Transactions on Machine Learning Research \(TMLR\)\.External Links:[Link](https://openreview.net/forum?id=iO4LZibEqW),2211\.09110Cited by:[§3\.2](https://arxiv.org/html/2608.06931#S3.SS2.SSS0.Px5.p1.1)\.
- \[26\]LLM\-Core Team, Xiaomi\(2026\)MiMo\-V2\.5\.Note:Official model pageExternal Links:[Link](https://mimo.xiaomi.com/mimo-v2-5)Cited by:[§3\.3](https://arxiv.org/html/2608.06931#S3.SS3.p1.1)\.
- \[27\]C\. Lu, C\. Lu, R\. T\. Lange, J\. Foerster, J\. Clune, and D\. Ha\(2024\)The AI Scientist: towards fully automated open\-ended scientific discovery\.External Links:2408\.06292,[Link](https://arxiv.org/abs/2408.06292)Cited by:[§2](https://arxiv.org/html/2608.06931#S2.SS0.SSS0.Px4.p1.1)\.
- \[28\]P\. Lu, H\. Bansal, T\. Xia, J\. Liu, C\. Li, H\. Hajishirzi, H\. Cheng, K\. Chang, M\. Galley, and J\. Gao\(2024\)MathVista: evaluating mathematical reasoning of foundation models in visual contexts\.InInternational Conference on Learning Representations \(ICLR\),External Links:[Link](https://openreview.net/forum?id=KUNzEQMWU7),2310\.02255Cited by:[§2](https://arxiv.org/html/2608.06931#S2.SS0.SSS0.Px2.p1.1)\.
- \[29\]P\. Lu, S\. Mishra, T\. Xia, L\. Qiu, K\. Chang, S\. Zhu, O\. Tafjord, P\. Clark, and A\. Kalyan\(2022\)Learn to explain: multimodal reasoning via thought chains for science question answering\.InAdvances in Neural Information Processing Systems \(NeurIPS\),External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2022/hash/11332b6b6cf4485b84afadb1352d3a9a-Abstract-Conference.html),2209\.09513Cited by:[§2](https://arxiv.org/html/2608.06931#S2.SS0.SSS0.Px2.p1.1)\.
- \[30\]R\. Luo, L\. Sun, Y\. Xia, T\. Qin, S\. Zhang, H\. Poon, and T\. Liu\(2022\)BioGPT: generative pre\-trained transformer for biomedical text generation and mining\.Briefings in Bioinformatics23\(6\),pp\. bbac409\.External Links:[Document](https://dx.doi.org/10.1093/bib/bbac409)Cited by:[§2](https://arxiv.org/html/2608.06931#S2.SS0.SSS0.Px4.p1.1)\.
- \[31\]Y\. Luo, J\. Zhang, S\. Fan, K\. Yang, Y\. Wu, M\. Qiao, and Z\. Nie\(2023\)BioMedGPT: open multimodal generative pre\-trained transformer for biomedicine\.External Links:2308\.09442,[Link](https://arxiv.org/abs/2308.09442)Cited by:[§2](https://arxiv.org/html/2608.06931#S2.SS0.SSS0.Px4.p1.1)\.
- \[32\]MiniMax\(2026\)MiniMax\-M3: frontier coding, 1m context, native multimodality\.Note:Official blog postExternal Links:[Link](https://www.minimax.io/blog/minimax-m3)Cited by:[§3\.3](https://arxiv.org/html/2608.06931#S3.SS3.p1.1)\.
- \[33\]A\. Mirza, N\. Alampara, S\. Kunchapu, M\. Ríos\-García, B\. Emoekabu, A\. Krishnan, K\. M\. Jablonka,et al\.\(2025\)A framework for evaluating the chemical knowledge and reasoning abilities of large language models against the expertise of chemists\.Nature Chemistry\.Note:ChemBenchExternal Links:[Document](https://dx.doi.org/10.1038/s41557-025-01815-x)Cited by:[§2](https://arxiv.org/html/2608.06931#S2.SS0.SSS0.Px3.p1.1)\.
- \[34\]Moonshot AI\(2026\)Kimi K2\.6: advancing open\-source coding\.Note:Official blog postExternal Links:[Link](https://www.kimi.com/blog/kimi-k2-6)Cited by:[§3\.3](https://arxiv.org/html/2608.06931#S3.SS3.p1.1)\.
- \[35\]Moonshot AI\(2026\)Kimi K3: open frontier intelligence\.Note:Technical reportExternal Links:[Link](https://arxiv.org/abs/2607.24653)Cited by:[§3\.3](https://arxiv.org/html/2608.06931#S3.SS3.p1.1)\.
- \[36\]OpenAI\(2026\)GPT\-5\.5 \(xhigh\)\.Note:Official announcementExternal Links:[Link](https://openai.com/zh-Hans-CN/index/introducing-gpt-5-5/)Cited by:[§3\.3](https://arxiv.org/html/2608.06931#S3.SS3.p1.1)\.
- \[37\]OpenAI\(2026\)GPT\-5\.6\-Sol\.Note:Official announcementExternal Links:[Link](https://openai.com/index/gpt-5-6/)Cited by:[§3\.3](https://arxiv.org/html/2608.06931#S3.SS3.p1.1)\.
- \[38\]C\. Peng, X\. Yang, A\. Chen, K\. E\. Smith, N\. PourNejatian, A\. B\. Costa, C\. Martin, M\. G\. Flores, Y\. Zhang,et al\.\(2023\)A study of generative large language model for medical research and healthcare\.npj Digital Medicine6,pp\. 210\.Note:GatorTronGPTExternal Links:[Document](https://dx.doi.org/10.1038/s41746-023-00958-w),[Link](https://www.nature.com/articles/s41746-023-00958-w)Cited by:[§2](https://arxiv.org/html/2608.06931#S2.SS0.SSS0.Px4.p1.1)\.
- \[39\]L\. Phan, A\. Gatti, Z\. Han, N\. Li,et al\.\(2025\)Humanity’s last exam\.External Links:2501\.14249,[Link](https://arxiv.org/abs/2501.14249)Cited by:[§2](https://arxiv.org/html/2608.06931#S2.SS0.SSS0.Px1.p1.1)\.
- \[40\]Qwen Team, Alibaba Cloud\(2026\)Qwen3\.7\-Plus: multimodal agent intelligence\.Note:Official announcementExternal Links:[Link](https://www.alibabacloud.com/blog/qwen3-7-plus-multimodal-agent-intelligence_603206)Cited by:[§3\.3](https://arxiv.org/html/2608.06931#S3.SS3.p1.1)\.
- \[41\]Qwen Team, Alibaba Cloud\(2026\)Qwen3\.8\-Max\.Note:Official announcementExternal Links:[Link](https://qwen.ai/blog?id=qwen3.8)Cited by:[§3\.3](https://arxiv.org/html/2608.06931#S3.SS3.p1.1)\.
- \[42\]D\. Rein, B\. L\. Hou, A\. C\. Stickland, J\. Petty, R\. Y\. Pang, J\. Dirani, J\. Michael, and S\. R\. Bowman\(2023\)GPQA: a graduate\-level google\-proof Q&A benchmark\.External Links:2311\.12022,[Link](https://arxiv.org/abs/2311.12022)Cited by:[§2](https://arxiv.org/html/2608.06931#S2.SS0.SSS0.Px1.p1.1)\.
- \[43\]J\. Roberts, K\. Han, N\. Houlsby, and S\. Albanie\(2024\)SciFIBench: benchmarking large multimodal models for scientific figure interpretation\.External Links:2405\.08807,[Link](https://arxiv.org/abs/2405.08807)Cited by:[§2](https://arxiv.org/html/2608.06931#S2.SS0.SSS0.Px2.p1.1),[§3\.2](https://arxiv.org/html/2608.06931#S3.SS2.SSS0.Px4.p1.1)\.
- \[44\]J\. Ruan, D\. Jiang, X\. Gao, T\. Liu, Y\. Fu, and Y\. Kang\(2026\)MME\-SCI: a comprehensive and challenging science benchmark for multimodal large language models\.InProceedings of the AAAI Conference on Artificial Intelligence \(AAAI\),Vol\.40,pp\. 8760–8768\.External Links:2508\.13938,[Document](https://dx.doi.org/10.1609/aaai.v40i11.37829),[Link](https://ojs.aaai.org/index.php/AAAI/article/view/37829)Cited by:[§2](https://arxiv.org/html/2608.06931#S2.SS0.SSS0.Px2.p1.1)\.
- \[45\]S\. Schmidgall, Y\. Su, Z\. Wang, X\. Sun, J\. Wu, X\. Yu, J\. Liu, M\. Moor, Z\. Liu, and E\. Barsoum\(2025\)Agent Laboratory: using LLM agents as research assistants\.InFindings of the Association for Computational Linguistics: EMNLP 2025,Suzhou, China,pp\. 5977–6043\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.findings-emnlp.320),[Link](https://aclanthology.org/2025.findings-emnlp.320/),2501\.04227Cited by:[§2](https://arxiv.org/html/2608.06931#S2.SS0.SSS0.Px4.p1.1)\.
- \[46\]Shanghai AI Laboratory\(2026\)Intern\-S1 Pro\.External Links:2603\.25040,[Link](https://arxiv.org/abs/2603.25040)Cited by:[§2](https://arxiv.org/html/2608.06931#S2.SS0.SSS0.Px4.p1.1),[§3\.3](https://arxiv.org/html/2608.06931#S3.SS3.p1.1)\.
- \[47\]K\. Singhal, T\. Tu, J\. Gottweis, R\. Sayres, E\. Wulczyn, A\. Mohamed, L\. Hou, K\. Clark, S\. R\. Pfohl,et al\.\(2025\)Toward expert\-level medical question answering with large language models\.Nature Medicine31,pp\. 943–950\.Note:Med\-PaLMExternal Links:[Document](https://dx.doi.org/10.1038/s41591-024-03423-7),[Link](https://www.nature.com/articles/s41591-024-03423-7)Cited by:[§2](https://arxiv.org/html/2608.06931#S2.SS0.SSS0.Px4.p1.1)\.
- \[48\]Y\. Song, S\. Miret, H\. Zhang, and B\. Liu\(2023\)HoneyBee: progressive instruction finetuning of large language models for materials science\.InFindings of the Association for Computational Linguistics: EMNLP 2023,Singapore,pp\. 5724–5739\.External Links:[Document](https://dx.doi.org/10.18653/v1/2023.findings-emnlp.380),[Link](https://aclanthology.org/2023.findings-emnlp.380/),2310\.08511Cited by:[§2](https://arxiv.org/html/2608.06931#S2.SS0.SSS0.Px4.p1.1)\.
- \[49\]L\. Sun, Y\. Han, Z\. Zhao, D\. Ma, Z\. Shen, B\. Chen, L\. Chen, and K\. Yu\(2024\)SciEval: a multi\-level large language model evaluation benchmark for scientific research\.InProceedings of the AAAI Conference on Artificial Intelligence \(AAAI\),External Links:[Link](https://ojs.aaai.org/index.php/AAAI/article/view/29872)Cited by:[§2](https://arxiv.org/html/2608.06931#S2.SS0.SSS0.Px3.p1.1)\.
- \[50\]Y\. Tang, W\. Xu, J\. Cao, W\. Gao, S\. Farrell, B\. Erichson, M\. W\. Mahoney, A\. Nonaka, and Z\. Yao\(2026\)A multimodal large language model for materials science\.Nature Machine Intelligence\.Note:MatterChatExternal Links:[Document](https://dx.doi.org/10.1038/s42256-026-01214-y),[Link](https://www.nature.com/articles/s42256-026-01214-y)Cited by:[§2](https://arxiv.org/html/2608.06931#S2.SS0.SSS0.Px4.p1.1)\.
- \[51\]R\. Taylor, M\. Kardas, G\. Cucurull, T\. Scialom, A\. Hartshorn, E\. Saravia, A\. Poulton, V\. Kerkez, and R\. Stojnic\(2022\)Galactica: a large language model for science\.External Links:2211\.09085,[Link](https://arxiv.org/abs/2211.09085)Cited by:[§2](https://arxiv.org/html/2608.06931#S2.SS0.SSS0.Px4.p1.1)\.
- \[52\]T\. Tu, S\. Azizi, D\. Driess, M\. Schaekermann, M\. Amin, P\. Chang, A\. Carroll, C\. Lau, R\. Tanno,et al\.\(2023\)Towards generalist biomedical AI\.External Links:2307\.14334,[Link](https://arxiv.org/abs/2307.14334)Cited by:[§2](https://arxiv.org/html/2608.06931#S2.SS0.SSS0.Px4.p1.1)\.
- \[53\]K\. Vishwanath, A\. Alyakin, M\. Ghosh, A\. Hage,et al\.\(2026\)General\-purpose large language models outperform specialized clinical AI tools on medical benchmarks\.Nature Medicine\.External Links:[Document](https://dx.doi.org/10.1038/s41591-026-04431-5),[Link](https://www.nature.com/articles/s41591-026-04431-5)Cited by:[§4\.2](https://arxiv.org/html/2608.06931#S4.SS2.p2.1)\.
- \[54\]X\. Wang, Z\. Hu, P\. Lu, Y\. Zhu, J\. Zhang, S\. Subramaniam, A\. R\. Loomba, S\. Zhang, Y\. Sun, and W\. Wang\(2024\)SciBench: evaluating college\-level scientific problem\-solving abilities of large language models\.InInternational Conference on Machine Learning \(ICML\),External Links:[Link](https://proceedings.mlr.press/v235/wang24z.html)Cited by:[§2](https://arxiv.org/html/2608.06931#S2.SS0.SSS0.Px3.p1.1)\.
- \[55\]Y\. Wang, X\. Ma, G\. Zhang, Y\. Ni, A\. Chandra, S\. Guo, W\. Ren,et al\.\(2024\)MMLU\-Pro: a more robust and challenging multi\-task language understanding benchmark\.InAdvances in Neural Information Processing Systems \(NeurIPS\), Datasets and Benchmarks Track,External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2024/hash/ad236edc564f3e3156e1b2feafb99a24-Abstract-Datasets_and_Benchmarks_Track.html),2406\.01574Cited by:[§2](https://arxiv.org/html/2608.06931#S2.SS0.SSS0.Px1.p1.1)\.
- \[56\]X\. Ye, C\. Li, S\. Chen, W\. Wei, and R\. Tang\(2025\)MMSciBench: benchmarking language models on Chinese multimodal scientific problems\.InFindings of the Association for Computational Linguistics: ACL 2025,Vienna, Austria,pp\. 14621–14663\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.findings-acl.755),[Link](https://aclanthology.org/2025.findings-acl.755/)Cited by:[§2](https://arxiv.org/html/2608.06931#S2.SS0.SSS0.Px2.p1.1)\.
- \[57\]B\. Yu, F\. N\. Baker, Z\. Chen, X\. Ning, and H\. Sun\(2024\)LlaSMol: advancing large language models for chemistry with a large\-scale, comprehensive, high\-quality instruction tuning dataset\.External Links:2402\.09391,[Link](https://arxiv.org/abs/2402.09391)Cited by:[§2](https://arxiv.org/html/2608.06931#S2.SS0.SSS0.Px4.p1.1)\.
- \[58\]W\. Yu, Z\. Yang, L\. Li, J\. Wang, K\. Lin, Z\. Liu, X\. Wang, and L\. Wang\(2024\)MM\-Vet: evaluating large multimodal models for integrated capabilities\.InInternational Conference on Machine Learning \(ICML\),Vol\.235,pp\. 57730–57754\.External Links:[Link](https://proceedings.mlr.press/v235/yu24o.html),2308\.02490Cited by:[§2](https://arxiv.org/html/2608.06931#S2.SS0.SSS0.Px2.p1.1)\.
- \[59\]X\. Yue, Y\. Ni, K\. Zhang, T\. Zheng, R\. Liu, G\. Zhang, S\. Stevens, D\. Jiang, W\. Ren, Y\. Sun,et al\.\(2024\)MMMU: a massive multi\-discipline multimodal understanding and reasoning benchmark for expert AGI\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),pp\. 9556–9567\.External Links:[Link](https://openaccess.thecvf.com/content/CVPR2024/html/Yue_MMMU_A_Massive_Multi-discipline_Multimodal_Understanding_and_Reasoning_Benchmark_for_CVPR_2024_paper.html),2311\.16502,[Document](https://dx.doi.org/10.1109/CVPR52733.2024.00913)Cited by:[§2](https://arxiv.org/html/2608.06931#S2.SS0.SSS0.Px2.p1.1),[§3\.2](https://arxiv.org/html/2608.06931#S3.SS2.SSS0.Px4.p1.1)\.
- \[60\]X\. Yue, T\. Zheng, Y\. Ni, Y\. Wang, K\. Zhang, S\. Tong, Y\. Sun, B\. Yu, G\. Zhang, H\. Sun, Y\. Su, W\. Chen, and G\. Neubig\(2024\)MMMU\-Pro: a more robust multi\-discipline multimodal understanding benchmark\.External Links:2409\.02813,[Link](https://arxiv.org/abs/2409.02813)Cited by:[§2](https://arxiv.org/html/2608.06931#S2.SS0.SSS0.Px2.p1.1)\.
- \[61\]D\. Zhang, W\. Liu, Q\. Tan, J\. Chen, H\. Yan, Y\. Yan, J\. Li, W\. Huang, X\. Yue,et al\.\(2024\)ChemLLM: a chemical large language model\.External Links:2402\.06852,[Link](https://arxiv.org/abs/2402.06852)Cited by:[§2](https://arxiv.org/html/2608.06931#S2.SS0.SSS0.Px4.p1.1)\.
- \[62\]Z\. Zhao, D\. Ma, L\. Chen, L\. Sun, Z\. Li, Y\. Xia, B\. Chen, H\. Xu,et al\.\(2025\)Developing ChemDFM as a large language foundation model for chemistry\.Cell Reports Physical Science6\(4\),pp\. 102523\.External Links:[Document](https://dx.doi.org/10.1016/j.xcrp.2025.102523),[Link](https://www.cell.com/cell-reports-physical-science/fulltext/S2666-3864(25)00122-5)Cited by:[§2](https://arxiv.org/html/2608.06931#S2.SS0.SSS0.Px4.p1.1)\.
- \[63\]L\. Zheng, W\. Chiang, Y\. Sheng, S\. Zhuang, Z\. Wu, Y\. Zhuang, Z\. Lin, Z\. Li, D\. Li, E\. P\. Xing, H\. Zhang, J\. E\. Gonzalez, and I\. Stoica\(2023\)Judging LLM\-as\-a\-Judge with MT\-Bench and Chatbot Arena\.InAdvances in Neural Information Processing Systems \(NeurIPS\), Datasets and Benchmarks Track,External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2023/hash/91f18a1287b398d378ef22505bf41832-Abstract-Datasets_and_Benchmarks.html),2306\.05685Cited by:[§3\.3](https://arxiv.org/html/2608.06931#S3.SS3.p3.1)\.
- \[64\]Zhipu AI \(Z\.ai\)\(2026\)GLM\-5V\-Turbo\.Note:Official documentationExternal Links:[Link](https://docs.z.ai/guides/vlm/glm-5v-turbo)Cited by:[§3\.3](https://arxiv.org/html/2608.06931#S3.SS3.p1.1)\.

Supplementary Information

## Appendix ADataset Details

SEEcontains 1,116 multimodal scientific questions\. Each released entry includes the question text, associated visual files, a standard answer, source information, a UUID, discipline labels, and knowledge\-point labels\. The associated visual files are entry\-level references: they may include images required by the question, but may also include images used in the solution rationale or expert verification, and therefore should not be read as a direct count of question\-side figures\.

The released question file includes both entries with explicit option lists and entries without a separate option field\. The latter include short\-answer, numerical, and visually grounded selection tasks whose answers are represented directly as canonical answers\. The dataset metadata records three task\-type categories: reasoning, measurement, and image processing\. The source field is normalized into DOI\-like literature sources, personal or experimental sources, and other source text or URL records\.

## Appendix BEvaluation

### B\.1 Evaluation Setup

We evaluateSEEon 19 representative MLLMs: Gemini 3\.1 Pro, GPT\-5\.6\-Sol \(Max\), Claude Opus 5 \(Max\), Qwen3\.8\-Max, GPT\-5\.5 \(xhigh\), Kimi K3, Gemini 3\.5 Flash, Claude Opus 4\.8 \(Max\), Seed2\.1 Pro, Qwen3\.7\-Plus, Seed2\.0 Pro, MiniMax\-M3, Kimi K2\.6, GLM\-5V\-Turbo, MiMo\-V2\.5, S1\-VL, Intern\-S2 Preview, Intern\-S2 Preview\-397B, and Intern\-S1 Pro\.

All models are evaluated using default inference settings, with two exceptions: GPT\-5\.5 is configured with extended thinking set toxhigh, and Claude Opus 4\.8 is configured with extended thinking set tomax\. For S1\-VL, we evaluate the S1\-VL\-32B\-RL checkpoint\. We do not introduce additional manual prompt optimization, response filtering, or post\-hoc correction during evaluation\. For questions containing image inputs, images are preprocessed according to the requirements of each model interface, including necessary format conversion, resizing, and input organization\. Claude Opus 4\.8 \(Max\) and Claude Opus 5 \(Max\) have safety\-filtered no\-response cases on biochemistry questions involving viral biology and pathogen structural characterization; these cases are counted as incorrect under the strict binary scoring protocol\.

SEEcontains 1,116 multimodal questions with both explicit\-option and open\-answer formats\. The raw dataset records associated visual files at the entry level; in evaluation, models are given the image inputs retained in the evaluation payload for each question\. Model performance is measured by accuracy, where a question is counted as correct only if the final model answer matches the ground\-truth answer under the scoring rules described below\.

### B\.2 Prompting Protocol

We use separate prompts for answer generation and answer judging\. The generation prompt asks each evaluated model to analyze the question and return a JSON object containing both the analysis and the final answer\. The judging prompt is executed by Gemini 3\.1 Pro; it compares the candidate response with the reference answer and returns a binary correctness label\.

Listing 1:Default generation prompt used in the evaluation pipeline\.\[System\]

Youareanintelligentassistant\.Pleasereadthequestionandimages

carefully,andprovidethecorrectanswer\.

\[User\]

Pleaseanswerthefollowingquestion\.Ifitisachoicequestion,itmay

beeithersingle\-choiceormultiple\-choice\.

\[Questiontext\]:

\{question\}

\[Options\]:

\{options\}

Pleasefirstprovidetheanalysisprocess,andthenprovidethefinal

answer\.StrictlyoutputthefollowingJSONformat,anddonotinclude

Markdownformatting:

\{

"analysis":"youranalysisprocess",

"answer":"finalanswer"

\}

Listing 2:Default judging prompt used in the evaluation pipeline\.\[System\]

Youareastrictexaminer\.Pleasejudgewhetherthestudent’sansweris

consistentwiththereferenceanswer\.

\[User\]

Pleasejudgewhetherthefollowingansweriscorrect:

\[Question\]:

\{question\}

\[Referenceanswer\]:

\{reference\}

\[Studentanswer\]\(itmaybeashortanswer,orafullresponsecontaining

ananalysisprocess\):

\{candidate\}

\[Task\]:

1\.Ifthestudent’sanswerisafullresponsecontainingananalysis

process,firstidentifythefinalanswerfromit,andthencompareit

withthereferenceanswer\.

2\.Judgewhetherthecoremeaningisconsistent\.Forchoicequestions,

thelettersmustbeidentical\.Forfill\-in\-the\-blankorshort\-answer

questions,thenumericalvalueorkeyphrasemustbeidentical\.Format

differencessuchas"5"and"5\.0"areallowed\.

3\.Ifthereferenceanswerisanumericalrangeorcontainsanerror

tolerance,suchas"5\-10","\[0\.8,1\.2\]","greaterthan100",

"5\+/\-0\.5","about0\.05",or"<3\.2",thestudent’sansweriscorrect

aslongasthenumericalvaluefallswithintherange\.Range

boundariesarealsotreatedascorrect\(closedinterval\)\.The

student’sanswermayalsobearange;inthiscase,judgewhetherthe

tworangesaresubstantiallyconsistent,withsimilarcentersand

widths\.

4\.Ifthestudentrefusestoanswerornoanswercanbeidentified,mark

itasincorrect\.

StrictlyoutputthefollowingJSONformat,anddonotincludeMarkdown

formatting:

\{

"correct":trueorfalse,

"reason":"judgmentreasonwithin50characters"

\}

For models that support multimodal inputs, images are provided together with the question text in the same evaluation instance\. No additional tool use, retrieval, or external browsing is allowed in the standard evaluation setting\.

### B\.3 Answer Extraction and Scoring

Candidate responses are scored by Gemini 3\.1 Pro using the judging prompt in Listing[2](https://arxiv.org/html/2608.06931#LST2)\. For entries with explicit option lists, including both single\-choice and multiple\-choice questions, the judge first extracts the final answer from the model response and then checks whether the predicted option letters exactly match the ground\-truth option letters\. Partial matches are not counted as correct\.

For entries without a separate option field, answers are evaluated against canonical short answers or numerical answers\. Numerical answers are treated as correct when they match the reference value or fall within the reference range or tolerance specified in the answer\. Model outputs that refuse to answer, state that the question cannot be determined, fail to provide a final answer, or otherwise cannot be judged as matching the ground\-truth answer are counted as incorrect\. Thus, all final evaluation results follow a strict binary scoring protocol: each question is either correct or incorrect\.

### B\.4 Judge Sensitivity Analysis

Because Gemini 3\.1 Pro serves as both the default judge model and one of the evaluated models, we conduct a sensitivity analysis to verify that scores are not biased by same\-model alignment\. We re\-judge all responses of a representative subset of five evaluated models using GPT\-5\.5 as an independent judge, covering both standard and text\-only \(no\-image\) evaluation settings \(10 runs,∼\\sim11,000 judgments in total\)\.

Table[1](https://arxiv.org/html/2608.06931#A2.T1)reports inter\-judge agreement\. Across all runs, raw agreement exceeds 99\.5% and Cohen’sκ\\kappaexceeds 0\.989, indicating near\-perfect concordance\. The maximum accuracy difference between the two judges is 0\.46 percentage points\. Notably, for Gemini 3\.1 Pro’s own responses, the alternative judge assigns a marginally*higher*accuracy \(\+0\.18 pp\), ruling out self\-favoring bias in the original judge\.

Manual inspection of the disagreement cases reveals that Gemini 3\.1 Pro’s judgments are more consistent with the intended scoring protocol\. The original judge applies stricter matching criteria aligned with our evaluation rules \(e\.g\., requiring identifiable final answers rather than accepting truncated reasoning traces, and enforcing format requirements specified in the judging prompt\), while GPT\-5\.5 more readily accepts semantically plausible but formally non\-conforming outputs\. Based on this human verification, we retain Gemini 3\.1 Pro as the default judge\. Its conservative tendency works against, rather than in favor of, any hypothetical same\-model bias\.

Table 1:Inter\-judge agreement between Gemini 3\.1 Pro \(default\) and GPT\-5\.5 \(independent\)\. Agreement is the fraction of identically scored questions\.κ\\kappais Cohen’s kappa\.Δ\\DeltaAcc\. is GPT\-5\.5 judge accuracy minus Gemini 3\.1 Pro judge accuracy, in percentage points\. The asterisk on Claude Opus 4\.8 \(Max\) denotes safety\-filtered no\-response cases on biochemistry questions, counted as incorrect under the strict binary scoring protocol\.Evaluated ModelAgreement \(%\)Cohen’sκ\\kappaΔ\\DeltaAcc\. \(pp\)Claude Opus 4\.8 \(Max\)\*99\.810\.996\+0\.46Seed2\.0 Pro99\.640\.991\+0\.36Gemini 3\.1 Pro99\.820\.996\+0\.18GPT\-5\.5 \(xhigh\)99\.820\.996\+0\.14Kimi K2\.699\.720\.993\+0\.09
### B\.5 Tool\-Augmented Visual\-Agent Evaluation Protocol

To test whether iterative, tool\-augmented inference improves performance under more pragmatic research\-assistance conditions, we evaluate six models in a visual\-agent environment with both web search and a code interpreter: GPT\-5\.6\-Sol \(Max\), GPT\-5\.5 \(xhigh\), Claude Opus 5 \(Max\), Gemini 3\.1 Pro, Qwen3\.8\-Max, and Gemini 3\.5 Flash\. For each model, we use the tool capabilities officially provided in its evaluation environment rather than third\-party or custom\-built substitutes\. Each tool\-augmented visual\-agent run uses the same 1,116\-question evaluation set as its corresponding standard baseline and receives the same question text, options, and image inputs\.

The tool\-augmented visual\-agent setting differs from the standard baseline by allowing each model to form a multi\-step interaction trajectory at inference time\. At each step, the model may invoke web search to gather external context or use the code interpreter to inspect and manipulate the visual input, perform computation, or verify an intermediate interpretation\. The resulting text, numerical output, or processed visual observation is returned to the model and may inform its next action or final answer\. No local database, curated retrieval corpus, or domain\-specific scientific tool is provided\. Tool use is optional rather than forced, and the model may terminate the trajectory and answer directly whenever it judges the available evidence sufficient\. This design distinguishes the tool\-augmented visual\-agent setting from the standard baseline, which requires a final answer from the original multimodal input without external retrieval or code execution\.

Table 2:Configuration of the tool\-augmented visual\-agent setting\. Each model receives the same benchmark inputs as in the standard multimodal baseline and uses the web\-search and code\-interpreter capabilities officially provided in its evaluation environment for iterative evidence acquisition and verification during inference\. The asterisk on Claude Opus 5 \(Max\) denotes no\-response cases, which are counted as incorrect under the strict binary scoring protocol\.ComponentConfigurationEvaluated modelsSix MLLMs with complete paired standard and tool\-augmented visual\-agent runs: GPT\-5\.6\-Sol \(Max\), GPT\-5\.5 \(xhigh\), Claude Opus 5 \(Max\)\*, Gemini 3\.1 Pro, Qwen3\.8\-Max, and Gemini 3\.5 Flash\.Input payloadSame question text, options, and evaluation image inputs as the standard baseline evaluation\.Added toolsThe web\-search and code\-interpreter capabilities officially provided in each model’s evaluation environment; no third\-party or custom\-built substitutes are used\. No local database, curated retrieval corpus, or specialized scientific tool is enabled\.Tool policyThe evaluated model decides whether to call available tools; tool use is optional rather than required for every question\.ScoringSame binary judging protocol as the baseline; stage failures, refusals, unanswered outputs, and outputs without an identifiable final answer are counted as incorrect\.The tool\-augmented visual\-agent evaluation contains two stages\. In the first stage, the evaluated model analyzes the multimodal question and may construct a variable\-length visual\-agent trajectory: it selects a tool action, observes the execution result, and either continues gathering or verifying evidence or terminates with the structured answer object described above\. In the second stage, only the final answer produced by the trajectory is evaluated, using the same binary correctness protocol as in the baseline experiments\. Intermediate tool outputs are retained for trajectory\-level diagnostics but are not scored independently\.

For multiple\-choice questions, the final option letters must match the ground\-truth answer exactly\. For short\-answer and numerical questions, the answer is compared against the standardized canonical answer and any expert\-defined tolerance\. Tool\-augmented execution failures, refusals, unanswered outputs, and outputs from which no final answer can be identified are counted as incorrect, matching the conservative treatment used in the standard evaluation\.

#### Human\-in\-the\-loop trajectory attribution\.

To diagnose how tool use changes an answer, we compare each paired standard and tool\-augmented visual\-agent trajectory for which correctness changes between the two settings\. We use a human\-in\-the\-loop procedure in which an attribution model first performs a structured preliminary analysis using the question, reference answer, standard response, tool\-augmented visual\-agent response, and available tool\-call trace\. It identifies the key step associated with the improvement or deterioration, determines whether and how web search or code execution plays a substantive role, and proposes a mechanism category with supporting evidence\. Domain experts then inspect the original responses and tool trace, verify the proposed attribution against the task\-specific experimental context, and revise the category or rationale when necessary\. Improvement mechanisms include code\-assisted figure analysis, retrieval of key scientific facts, search\-based confirmation, and code\-based computation\. Tool\-related detrimental mechanisms are organized into four recurring failure categories: action\-selection failure, observation\-interpretation failure, evidence\-integration failure, and endless overthinking\. Their relationship to the fine\-grained attribution types is summarized in Table[3](https://arxiv.org/html/2608.06931#A2.T3)\. The resulting expert\-validated attribution is used as a diagnostic analysis of the interaction trajectory rather than as an additional correctness score\.

Table 3:Mapping between the recurring categories of tool\-related failure and their fine\-grained manifestations in the trajectory attribution analysis\.Failure categoryFine\-grained manifestationsAction\-selection failureInappropriate use of code, including an unsuitable image\-processing method, data\-processing assumption, calculation target, or retrieval direction\.Observation\-interpretation failureCode\-logic or computation error; tool\-mediated misreading of figures, crops, measurements, or computed outputs\.Evidence\-integration failureOver\-reliance on retrieved information; noisy or misleading retrieval; inappropriate weighting of tool\-derived evidence relative to the question\-specific experimental evidence\.Endless overthinkingExcessive tool use and context degradation; unproductive repetition; failure to stop or produce a final answer\.Across the six models, 990 paired instances change correctness between the standard and tool\-augmented settings: 648 change from incorrect to correct and 342 from correct to incorrect, yielding a net gain of 306 correct predictions\. Trajectory attribution identifies substantive tool contributions in 411 improvements and 145 regressions\. Among the remaining changes, some occur without tool calls, while others follow tool use but lack evidence that a specific tool output decisively changes the final answer; these cases are therefore classified primarily as variation in reasoning or visual interpretation\.

Table 4:Paired correctness transitions and trajectory attribution by model\. Wrong→\\rightarrowCorrect and Correct→\\rightarrowWrong report all answer changes between the standard and tool\-augmented visual\-agent settings\. Tool\-attributed improvements and regressions report the subsets for which web search or code execution is identified as playing a substantive role\. Percentages are calculated relative to the corresponding number of improvements or regressions for each model\.ModelWrong→\\rightarrowCorrectCorrect→\\rightarrowWrongNetchangeTool\-attributedimprovementsTool\-attributedregressionsClaude Opus 5 \(Max\)8038\+4224 \(30\.0%\)8 \(21\.1%\)Gemini 3\.1 Pro9574\+2124 \(25\.3%\)15 \(20\.3%\)Gemini 3\.5 Flash13277\+5575 \(56\.8%\)28 \(36\.4%\)GPT\-5\.5 \(xhigh\)12742\+85103 \(81\.1%\)19 \(45\.2%\)GPT\-5\.6\-Sol \(Max\)11368\+45101 \(89\.4%\)50 \(73\.5%\)Qwen3\.8\-Max10143\+5884 \(83\.2%\)25 \(58\.1%\)Overall648342\+306411 \(63\.4%\)145 \(42\.4%\)

### B\.6 Public\-Subset Reproducibility

The main evaluation uses all 1,116 questions inSEE, including 67 withheld questions based on unpublished experimental data\. To support reproducibility on the released benchmark, we recompute headline results on the 1,049 publicly released questions marked as public in theopen\_sourcefield of the dataset metadata\. Public\-subset accuracies closely track the full\-set results, indicating that the reported conclusions are not driven by the withheld subset\.

Table 5:Full\-set and public\-subset accuracy under the standard no\-tool evaluation setting\. The public subset contains the 1,049 released questions; the full set contains all 1,116 questions\. Delta is public\-subset accuracy minus full\-set accuracy, in percentage points\. Asterisks on Claude Opus 4\.8 \(Max\) and Claude Opus 5 \(Max\) denote safety\-filtered no\-response cases on biochemistry questions, counted as incorrect under the strict binary scoring protocol\.ModelFullCorrectFullAcc\. \(%\)PublicCorrectPublicAcc\. \(%\)Delta\(pp\)Gemini 3\.1 Pro504/1,11645\.2478/1,04945\.6\+0\.4GPT\-5\.6\-Sol \(Max\)543/1,11648\.7503/1,04948\.0\-0\.7Claude Opus 5 \(Max\)\*495/1,11644\.4465/1,04944\.3\-0\.0Qwen3\.8\-Max466/1,11641\.8439/1,04941\.8\+0\.1GPT\-5\.5 \(xhigh\)461/1,11641\.3433/1,04941\.3\-0\.0Kimi K3432/1,11638\.7403/1,04938\.4\-0\.3Gemini 3\.5 Flash431/1,11638\.6410/1,04939\.1\+0\.5Claude Opus 4\.8 \(Max\)\*406/1,11636\.4377/1,04935\.9\-0\.4Seed2\.1 Pro377/1,11633\.8350/1,04933\.4\-0\.4Qwen3\.7\-Plus365/1,11632\.7346/1,04933\.0\+0\.3Intern\-S2 Preview\-397B324/1,11629\.0302/1,04928\.8\-0\.2Seed2\.0 Pro315/1,11628\.2298/1,04928\.4\+0\.2MiniMax\-M3314/1,11628\.1302/1,04928\.8\+0\.7Kimi K2\.6286/1,11625\.6266/1,04925\.4\-0\.3Intern\-S2 Preview266/1,11623\.8252/1,04924\.0\+0\.2GLM\-5V\-Turbo259/1,11623\.2245/1,04923\.4\+0\.1Intern\-S1 Pro225/1,11620\.2215/1,04920\.5\+0\.3MiMo\-V2\.5196/1,11617\.6185/1,04917\.6\+0\.1S1\-VL177/1,11615\.9169/1,04916\.1\+0\.3Table 6:Public\-subset accuracy under the combined web\-search\-and\-code\-interpreter setting\. Results are recomputed on the same 1,049 publicly released questions used in Table[5](https://arxiv.org/html/2608.06931#A2.T5)\. The asterisk on Claude Opus 5 \(Max\) denotes no\-response cases, which are counted as incorrect under the strict binary scoring protocol\.ModelCorrectPublic Acc\. \(%\)GPT\-5\.6\-Sol \(Max\)547/1,04952\.1GPT\-5\.5 \(xhigh\)511/1,04948\.7Claude Opus 5 \(Max\)\*502/1,04947\.9Gemini 3\.1 Pro497/1,04947\.4Qwen3\.8\-Max488/1,04946\.5Gemini 3\.5 Flash459/1,04943\.8
### B\.7 Text\-Only Ablation

To examine whether models recognize missing visual evidence, we conduct a text\-only ablation by removing all evaluation image inputs from the questions while retaining the textual question content\. This setting is designed to test whether models can avoid unsupported conclusions when key information is missing\.

Figure[9](https://arxiv.org/html/2608.06931#A2.F9)reports the paired accuracies for each model; alongside the mean drop, the spread across models narrows from 31\.1 to 14\.0 points\.

![Refer to caption](https://arxiv.org/html/2608.06931v1/x10.png)Figure 9:Accuracy with and without the question image for the 18 models with paired runs\. Each model is scored twice on the same question set; the text\-only run receives the identical question with the image removed\. Accuracy follows the strict binary protocol, in which every instance stays in the denominator and refusals, non\-answers, and generation errors count as incorrect\. Delta is the text\-only accuracy minus the with\-image accuracy\.Across the models with available no\-image runs, explicit missing\-image acknowledgment is consistently rare \(Table[7](https://arxiv.org/html/2608.06931#A2.T7)\)\. Across 18 text\-only runs, only 921 out of 20,088 instances \(4\.6%\) are classified as explicit acknowledgments that necessary image evidence is missing\. This pattern indicates that current MLLMs generally lack robust evidence\-boundary awareness under visually under\-specified conditions\.

In this ablation setting, only explicit missing\-image acknowledgments are treated as evidence\-aware behavior for the main analysis\. Empty outputs, generation errors, generic refusals, and other non\-answer cases are not counted toward this metric because they do not directly show recognition that visual evidence is unavailable\. Unsupported final answers are analyzed as evidence\-insensitive attempts\.

Table 7:Explicit missing\-image acknowledgments in the text\-only ablation\. The denominator is all text\-only instances for each model; the numerator includes only responses withrefusal\_type=no\_image, i\.e\., responses explicitly classified as acknowledging missing image or visual evidence\. Empty outputs, generation errors, generic refusals, and other non\-answer cases are not counted\.ModelTotalExplicit Missing\-Image Ack\.Rate \(%\)GPT\-5\.6\-Sol \(Max\)1,11616715\.0Kimi K2\.61,11616314\.6GPT\-5\.5 \(xhigh\)1,11614613\.1Intern\-S2 Preview\-397B1,116908\.1Qwen3\.8\-Max1,116756\.7MiniMax\-M31,116615\.5Intern\-S2 Preview1,116554\.9MiMo\-V2\.51,116373\.3Qwen3\.7\-Plus1,116302\.7Kimi K31,116272\.4Seed2\.0 Pro1,116201\.8GLM\-5V\-Turbo1,116171\.5Gemini 3\.1 Pro1,116141\.3Gemini 3\.5 Flash1,11670\.6Intern\-S1 Pro1,11660\.5Claude Opus 4\.8 \(Max\)\*1,11640\.4Seed2\.1 Pro1,11620\.2Claude Opus 5 \(Max\)\*1,11600\.0Total / Mean20,0889214\.6
\*Safety\-filtered no\-response cases, counted as incorrect under the strict binary scoring protocol\.

### B\.8 Per\-Discipline Results

We report model performance across the three main disciplines \(chemistry, biology, materials\) and the final set of 17 fine\-grained discipline labels used for reported sub\-discipline analysis\. BecauseSEEuses multi\-label annotation, a single question may contribute to multiple labels; counts and accuracies follow label\-membership semantics\.

Table[8](https://arxiv.org/html/2608.06931#A2.T8)reports per\-sub\-discipline accuracy for all 19 evaluated models, ordered by overall accuracy\.

Table 8:Per\-model accuracy by reported sub\-discipline for all 19 evaluated models, ordered by overall accuracy\. Values are percentages; bold indicates the best model for each label\. Asterisks on Claude Opus 4\.8 \(Max\) and Claude Opus 5 \(Max\) denote no\-response cases triggered by safety filters on biochemistry questions involving viral biology and pathogen structural characterization; under the strict binary scoring protocol, these instances are counted as incorrect\.MainDisciplineSub\-DisciplineModel Accuracy \(%\)GPT\-5\.6\-Sol \(Max\) Gemini 3\.1 Pro Claude Opus 5\(Max\)\* Qwen3\.8\-Max GPT\-5\.5 \(xhigh\) Kimi K3 Gemini 3\.5 Flash Claude Opus 4\.8\(Max\)\* Seed2\.1 Pro Qwen3\.7\-Plus Intern\-S2 Preview\-397B Seed2\.0 Pro MiniMax\-M3 Kimi K2\.6 Intern\-S2 Preview GLM\-5V\-Turbo Intern\-S1 Pro MiMo\-V2\.5 S1\-VL ChemistryAnalytical Chemistry50\.548\.249\.542\.342\.539\.542\.540\.936\.433\.930\.228\.431\.125\.928\.624\.123\.417\.519\.3Physical Chemistry41\.846\.042\.131\.835\.332\.639\.231\.225\.827\.628\.521\.127\.024\.627\.919\.623\.113\.916\.6Organic Chemistry53\.050\.953\.946\.146\.140\.047\.449\.640\.938\.735\.230\.937\.027\.031\.725\.723\.520\.422\.2Polymer Chemistry and Physics37\.942\.540\.832\.235\.137\.435\.129\.924\.121\.824\.721\.329\.928\.227\.614\.928\.715\.517\.8Inorganic Chemistry44\.539\.145\.535\.532\.730\.932\.729\.131\.826\.420\.016\.430\.023\.630\.020\.923\.610\.014\.5BiologyBiochemistry47\.444\.842\.742\.743\.640\.434\.035\.833\.433\.729\.132\.323\.523\.817\.722\.714\.216\.911\.0Molecular Biology48\.344\.139\.147\.543\.342\.132\.634\.536\.435\.628\.733\.721\.524\.913\.823\.813\.019\.510\.7Cell Biology53\.843\.439\.048\.446\.244\.534\.636\.339\.036\.328\.639\.626\.923\.617\.024\.713\.222\.511\.0Structural Biology48\.842\.234\.342\.241\.041\.028\.333\.131\.333\.126\.530\.121\.127\.114\.525\.321\.117\.512\.7Biophysics47\.444\.938\.536\.538\.539\.137\.230\.130\.832\.729\.530\.120\.523\.121\.826\.318\.618\.616\.7Genetics47\.261\.144\.458\.352\.838\.925\.027\.825\.041\.719\.441\.722\.227\.82\.819\.42\.816\.72\.8Immunology56\.231\.231\.262\.543\.843\.834\.434\.440\.637\.525\.040\.628\.112\.531\.231\.231\.218\.812\.5Physiology42\.142\.152\.647\.436\.831\.636\.826\.331\.631\.621\.126\.30\.015\.831\.621\.110\.526\.315\.8MaterialsOrganic and Polymer Materials37\.239\.934\.625\.030\.329\.832\.427\.122\.317\.020\.214\.925\.026\.125\.515\.428\.716\.018\.6Inorganic Non\-metallic Materials46\.643\.744\.737\.935\.933\.039\.833\.034\.035\.024\.322\.328\.224\.330\.126\.219\.414\.616\.5Composite Materials39\.741\.247\.129\.432\.435\.335\.332\.423\.522\.132\.425\.027\.925\.026\.514\.726\.513\.216\.2Metallic Materials57\.144\.942\.955\.136\.732\.749\.032\.736\.736\.734\.730\.634\.726\.534\.726\.518\.412\.218\.4

Table 9:Mean model accuracy across 19 evaluated models for each reported fine\-grained discipline label\. Labels with smaller sample sizes \(e\.g\., Immunology, Physiology, Genetics\) should be interpreted with care due to their limited support\.Sub\-disciplineMain Discipline\# QuestionsMean Acc\. \(%\)Organic ChemistryChemistry23037\.9Metallic MaterialsMaterials4934\.8Analytical ChemistryChemistry44034\.5ImmunologyBiology3234\.0Cell BiologyBiology18233\.1Molecular BiologyBiology26131\.2BiochemistryBiology34431\.0Inorganic Non\-metallic MaterialsMaterials10331\.0BiophysicsBiology15630\.6GeneticsBiology3630\.4Structural BiologyBiology16630\.1Physical ChemistryChemistry33729\.3PhysiologyBiology1928\.8Composite MaterialsMaterials6828\.7Polymer Chemistry and PhysicsChemistry17428\.7Inorganic ChemistryChemistry11028\.3Organic and Polymer MaterialsMaterials18825\.6
### B\.9 Subject Label Coverage

SEEexhibits pronounced multi\-label and cross\-disciplinary characteristics\. Under the final 17\-label reporting taxonomy, all 1,116 questions are covered by the reported fine\-grained labels\. The overall model evaluation is therefore based on the same 1,116\-question set used for reported label coverage\.

Table 10:Fine\-grained discipline\-label coverage under the 17\-label reporting taxonomy\.Fine\-Grained LabelMain Discipline\# QuestionsAnalytical ChemistryChemistry440BiochemistryBiology344Physical ChemistryChemistry337Molecular BiologyBiology261Organic ChemistryChemistry230Organic and Polymer MaterialsMaterials188Cell BiologyBiology182Polymer Chemistry and PhysicsChemistry174Structural BiologyBiology166BiophysicsBiology156Inorganic ChemistryChemistry110Inorganic Non\-metallic MaterialsMaterials103Composite MaterialsMaterials68Metallic MaterialsMaterials49GeneticsBiology36ImmunologyBiology32PhysiologyBiology19Table 11:Distribution of reported subject\-label cardinality among questions covered by the 17\-label taxonomy\. Most questions span multiple sub\-disciplines, reflecting the cross\-concept, cross\-direction, and cross\-discipline nature of real experimental scenarios rather than isolated single\-topic knowledge\.Labels per Question\# QuestionsShare \(%\)1968\.6241637\.3345841\.0413712\.3590\.8Table 12:Main\-discipline coverage by question count and by label annotation count\. Chemistry and biology coverage is high; materials coverage is comparatively lower\.Main Discipline\# QuestionsCoverage \(%\)\# LabelAnnotationsShare ofAnnotations \(%\)Chemistry73265\.61,29144\.6Biology53848\.21,19641\.3Materials36332\.540814\.1#### Cross\-Discipline Composition

Under the 17\-label reporting taxonomy, questions labeled with biology only total 370 \(33\.2%\); questions jointly labeled with chemistry and materials total 346 \(31\.0%\); chemistry\-only questions total 219 \(19\.6%\); and chemistry\-with\-biology questions total 164 \(14\.7%\)\. Materials\-only questions total just 13 \(1\.2%\)\. A small number of questions combine all three main disciplines \(3 questions, 0\.3%\) or biology with materials only \(1 question, 0\.1%\)\. This distribution matches the strong overlap between materials science and chemistry in real experimental practice, and accordingly model accuracy on materials labels should be interpreted as reflecting joint chemistry–materials reasoning rather than isolated materials knowledge\.

#### Long\-Tail Fine\-Grained Labels

Analytical chemistry \(440 questions\) is the most frequent reported fine\-grained label, followed by biochemistry \(344\) and physical chemistry \(337\)\. At the tail, metallic materials, genetics, immunology, and physiology have only 49, 36, 32, and 19 questions respectively\. Accuracy estimates on tail labels are more sensitive to a small number of questions and should be read with sample size in mind\.

Similar Articles

SciR: A Controllable Benchmark for Scientific Reasoning in LLMs

arXiv cs.AI

SciR is a new controllable benchmark for evaluating LLMs on scientific reasoning including deduction, induction, and causal abduction, with parametric control over extraction and inference difficulty. Tests show both axes degrade performance across models, with reasoning models like DeepSeek-R1 outperforming instruct models on inference.

COMPOSITE-Stem

arXiv cs.CL

COMPOSITE-STEM introduces a benchmark of 70 expert-curated agentic tasks across physics, biology, chemistry, and mathematics, designed to evaluate AI agents on scientific workflows beyond saturated benchmarks. The top-performing model (Claude Opus 4.6) achieves only 21.4%, demonstrating significant capability gaps in scientific reasoning.

SEAM: Global consistency beyond local accuracy in scientific machine learning

arXiv cs.LG

SEAM is a generator-agnostic framework that audits global consistency of explanations in scientific machine learning, detecting incompatible local explanations even when predictions are locally accurate and attributing failures to specific channels and overlaps. The paper presents theory and experiments across PDE systems, neural operators, and four open datasets.

SciPaths: Forecasting Pathways to Scientific Discovery

arXiv cs.CL

Introduces SciPaths, a benchmark for forecasting the enabling contributions required to realize a target scientific discovery, and evaluates frontier and open-weight language models, finding significant room for improvement in reasoning backward from contributions to enabling building blocks.