BioPhys-Bridge: A Benchmark for Interdisciplinary Scientific Reasoning in Physics-Grounded Biological Research

arXiv cs.AI Papers

Summary

BioPhys-Bridge introduces a benchmark dataset for evaluating language models on evidence-grounded scientific reasoning in biophysics, focusing on interdisciplinary tasks with quantitative and mechanistic grounding.

arXiv:2609.19180v1 Announce Type: new Abstract: Language models face unique challenges in analyzing interdisciplinary scientific research literature. In biophysics research, faithful answers require grounding observed data in source evidence, interpreting it through a quantitative physics model, and linking it to a biological mechanism. To address this challenge, we introduce BioPhys-Bridge, a novel benchmark dataset for evidence-grounded scientific reasoning over biophysical literature. Each case contains evidence blocks, stable evidence IDs, quantitative values, units, equations, assumptions, mechanisms, and next decisions as grounding targets for question answering (QA) and retrieval-augmented generation (RAG). The initial release contains 500 cases, 1,517 agent-facing tasks, and covers six biological domains and nine physical model families, including three sparse families reserved for future expansion. We enforce strict quality gates for all cases in schema, evidence-integrity, quantitative-grounding, source-license, duplicate, unit-normalization, with domain expert review and annotation for 81 cases. Preliminary evaluations show that DeepSeek-V4-Flash obtain the highest evidence-ID $F_1$ score (0.360), followed by Qwen3.7-Max (0.316) and GPT-4o-mini (0.294). BioPhys-Bridge is an interdisciplinary benchmark for evaluating attribution, faithfulness, hallucination reduction, and biological experiment design with complex, multi-step scientific reasoning. Future works will increase the size and complexity of the dataset and perform comprehensive evaluations. Code and data are available in the GitHub repository and on Hugging Face.
Original Article
View Cached Full Text

Cached at: 09/18/26, 09:16 AM

# BioPhys-Bridge: A Benchmark for Interdisciplinary Scientific Reasoning in Physics-Grounded Biological Research
Source: [https://arxiv.org/html/2609.19180](https://arxiv.org/html/2609.19180)
Qingyang Xu

###### Abstract

Language models face unique challenges in analyzing interdisciplinary scientific research literature\. In biophysics research, faithful answers require grounding observed data in source evidence, interpreting it through a quantitative physics model, and linking it to a biological mechanism\. To address this challenge, we introduce BioPhys\-Bridge, a novel benchmark dataset for evidence\-grounded scientific reasoning over biophysical literature\. Each case contains evidence blocks, stable evidence IDs, quantitative values, units, equations, assumptions, mechanisms, and next decisions as grounding targets for question answering \(QA\) and retrieval\-augmented generation \(RAG\)\. The initial release contains 500 cases, 1,517 agent\-facing tasks, and covers six biological domains and nine physical model families, including three sparse families reserved for future expansion\. We enforce strict quality gates for all cases in schema, evidence\-integrity, quantitative\-grounding, source\-license, duplicate, unit\-normalization, with domain expert review and annotation for 81 cases\. Preliminary evaluations show that DeepSeek\-V4\-Flash obtain the highest evidence\-IDF1F\_\{1\}score \(0\.360\), followed by Qwen3\.7\-Max \(0\.316\) and GPT\-4o\-mini \(0\.294\)\. BioPhys\-Bridge is an interdisciplinary benchmark for evaluating attribution, faithfulness, hallucination reduction, and biological experiment design with complex, multi\-step scientific reasoning\. Future works will increase the size and complexity of the dataset and perform comprehensive evaluations\. Code and data are available in the[GitHub repository](https://github.com/qyxu1994/BioPhys-Bridge)and on[Hugging Face](https://huggingface.co/datasets/qyxu1994/BioPhys-Bridge)\.

## 1Introduction

Language models \(LM\) are increasingly used to analyze scientific research papers in retrieval\-augmented generation \(RAG\) systems and other agentic AI frameworks\. However, retrieving a relevant paragraph does not guarantee the faithfulness of the final answer: the AI agent can cite loosely related evidence, omit key assumptions, mix units, or produce a plausible scientific interpretation that the cited paper does not support\. Scientific grounding requires finding faithful connections among source evidence, quantitative measurements, model assumptions, and the leveraging domain expertise to derive scientific conclusions and hypotheses\.

Biophysics is an epitome of this challenge due to its interdisciplinary nature\. The LM needs to both perform semantic text matching and draw quantitative predictions from a biophysical model, in order to connect a quantitative prediction \(e\.g\., binding affinity, reaction rate, phase\-separation threshold\) to the model assumptions under which it is meaningful, explain the underlying biological mechanism, and suggest the next experiment or computation\. Existing scientific QA and RAG benchmarks often focus on gauging whether an answer is supported by textual evidence\. In contrast, BioPhys\-Bridge gauges whether LM can perform evidence\-grounded reasoning across modalities \(equations and semantics\) and scientific disciplines\.

We propose a novel framework for grounding LM which requires LM to identify source evidence, use quantitative values and units consistently, respect physics model assumptions, and produce an actionable response supported by evidence IDs\. Evidence blocks are the atomic citation units, while equations, assumptions, quantitative evidence, and mechanism annotations provide intermediate structure for auditing whether a model’s answer is faithful to the research question and biological reasoning rather than superficial semantic relations\.

The initial release of BioPhys\-Bridge contains 500 cases extracted from open\-access peer\-reviewed research papers in biophysics\. This release contains 1,517 distinct tasks for biological research\. Every case passed rigorous schema validation, evidence\-integrity checks, quantitative\-grounding checks, source\-license coverage checks, unit normalization, duplicate checks, and semantic content filters\. To validate the correctness and scientific value of the curated samples, we asked domain experts to review and annotate 81 cases, including the full held\-out test set \(with 50 cases\) and additional golden examples\.

This works makes five important contributions for grounded scientific LM\. First, we introduce a novel benchmark dataset for literature\-grounded scientific reasoning over biophysical papers\. Second, we define a structured case schema that treats evidence blocks, quantitative values, equations, assumptions, mechanisms, and next decisions as first\-class grounding objects\. Third, we provide a task suite for derivation, mechanism\-from\-evidence, discrepancy explanation, and next\-experiment design, each requiring supporting evidence IDs\. Fourth, we describe a validated release pipeline with schema, evidence\-integrity, quantitative\-grounding, license, unit\-normalization, duplicate, and content\-quality gates\. Finally, we provide a de\-leaked evaluation harness with preliminary baselines that separate deterministic evidence retrieval from answer generation on the same held\-out tasks\.

As far as we know, BioPhys\-Bridge is the first benchmark for grounding LMs in interdisciplinary scientific reasoning in physics\-grounded biological research, where LM must connect evidence IDs to physics equations, model assumptions, biological mechanisms, and guide experiment design\. It is useful for studying faithfulness, attribution, and hallucination in real\-world research scenarios\.

## 2Related Work

### 2\.1Grounded Language Models and Scientific RAG

Scientific RAG systems aim to answer research questions by retrieving and synthesizing evidence from papers\. PaperQA and LitQA study retrieval\-augmented question answering over scientific literature\([Lala et al\., 2023](https://arxiv.org/html/2609.19180#bib.bib5)\)\. LitSearch evaluates realistic scientific literature search queries over recent papers\([Ajith et al\., 2024](https://arxiv.org/html/2609.19180#bib.bib7)\)\. LAB\-Bench measures biology research capabilities, including literature reasoning, database use, figure interpretation, and sequence manipulation\([Laurent et al\., 2024](https://arxiv.org/html/2609.19180#bib.bib4)\)\. SPIQA targets multimodal question answering over scientific paper figures\([Pramanick et al\., 2024](https://arxiv.org/html/2609.19180#bib.bib11)\)\. These resources make scientific documents central to evaluation, but their primary supervision is usually an answer, paper\-level evidence, retrieval target, or document QA context\. BioPhys\-Bridge instead makes evidence blocks, evidence IDs, quantitative values, units, equations, model assumptions, and mechanism links explicit fields in each case, so attribution can be evaluated beyond passage retrieval\.

BioPhys\-Bridge also builds on scientific document parsing and evidence extraction\. MinerU provides open\-source document content extraction for complex PDFs, including layout, table, formula, and OCR processing\([Wang et al\., 2024a](https://arxiv.org/html/2609.19180#bib.bib6)\)\. In our pipeline, MinerU is the parsing layer from which normalized evidence blocks are derived, audited, and linked to structured scientific reasoning fields\.

### 2\.2Scientific Reasoning and Multimodal Benchmarks

Scientific reasoning benchmarks increasingly test difficult science questions\. ScienceQA introduced a multimodal science question\-answering benchmark with explanations\([Lu et al\., 2022](https://arxiv.org/html/2609.19180#bib.bib1)\)\. SciBench evaluates college\-level scientific problem solving across mathematics, chemistry, and physics\([Wang et al\., 2024b](https://arxiv.org/html/2609.19180#bib.bib2)\)\. GPQA uses expert\-written graduate\-level biology, physics, and chemistry questions to stress scalable oversight\([Rein et al\., 2024](https://arxiv.org/html/2609.19180#bib.bib3)\)\. MathVista and MMMU combine visual understanding with expert or college\-level reasoning\([Lu et al\., 2024](https://arxiv.org/html/2609.19180#bib.bib8);[Yue et al\., 2024](https://arxiv.org/html/2609.19180#bib.bib9)\), while LongBench tests long\-context understanding over extended inputs\([Bai et al\., 2024](https://arxiv.org/html/2609.19180#bib.bib10)\)\. These benchmarks are valuable tests of scientific knowledge, reasoning, and multimodal understanding, but they generally do not represent a source\-paper case as an evidence\-linked chain from measurement to physical model to biological mechanism and scientific decision\.

### 2\.3AI\-for\-Science and Physical\-Model Benchmarks

Scientific machine\-learning resources such as PDEBench and BubbleML provide physics\-grounded data, benchmark tasks, and baselines for simulation and surrogate modeling\([Takamoto et al\., 2022](https://arxiv.org/html/2609.19180#bib.bib12);[Hassan et al\., 2023](https://arxiv.org/html/2609.19180#bib.bib13)\)\. Other realistic agent benchmarks, including SWE\-bench, WebArena, AgentBench, MINT, ToolBench, and MLE\-bench, evaluate software repair, web navigation, tool use, and machine\-learning engineering with functional, human\-referenced, or environment\-based scoring\([Jimenez et al\., 2024](https://arxiv.org/html/2609.19180#bib.bib14);[Zhou et al\., 2024](https://arxiv.org/html/2609.19180#bib.bib15);[Liu et al\., 2024](https://arxiv.org/html/2609.19180#bib.bib16);[Wang et al\., 2024c](https://arxiv.org/html/2609.19180#bib.bib17);[Qin et al\., 2024](https://arxiv.org/html/2609.19180#bib.bib18);[Chan et al\., 2025](https://arxiv.org/html/2609.19180#bib.bib19)\)\. BioPhys\-Bridge is complementary: it does not benchmark a PDE solver, surrogate model, or general\-purpose tool agent directly\. Instead, it evaluates source\-paper grounding for quantitative and mechanistic scientific interpretation\. BioPhys\-Bridge sits between literature\-grounded QA and AI\-for\-science benchmarks: it evaluates whether grounded language models can use source evidence to support quantitatively and mechanistically meaningful scientific interpretations\.

Table 1:Positioning BioPhys\-Bridge against recent benchmark families\. The distinguishing unit is not a question alone, but a structured case linking source evidence, physical model, quantitative evidence, biological mechanism, and scientific decision\.

## 3Dataset Design

### 3\.1Evidence\-Grounded Scientific Reasoning Cases

A BioPhys\-Bridge record is a structured grounding object rather than a standalone QA item\. The core reasoning pattern is:

evidence\\displaystyle\\text\{evidence\}→quantitative value→physical model\\displaystyle\\rightarrow\\text\{quantitative value\}\\rightarrow\\text\{physical model\}→mechanism→next decision\.\\displaystyle\\rightarrow\\text\{mechanism\}\\rightarrow\\text\{next decision\}\.Figure[1](https://arxiv.org/html/2609.19180#S3.F1)summarizes this structure\. The schema separates the biological setting from the physical model family\. The fielddomaincaptures the scientific area, whilebiophysical\_model\.model\_familyrecords the primary physical modeling strategy used by the main equation and decision\. Evidence blocks are the atomic citation units: all supporting evidence IDs in tasks must resolve to these blocks, and quantitative evidence must be traceable to cited evidence text\.

![Refer to caption](https://arxiv.org/html/2609.19180v1/figure1_case_structure.png)Figure 1:Each case as an evidence\-linked bridge from source\-paper evidence to quantitative values, physical\-model assumptions, mechanistic interpretation, and an agent\-facing scientific decision task\.Appendix[A](https://arxiv.org/html/2609.19180#A1)gives a concrete expert\-annotated gold\-case schematic for the same evidence\-to\-decision structure\.

### 3\.2Record Fields

Each case includes paper provenance, evidence blocks, quantitative evidence, model annotation, mechanism annotation, a staged trajectory, agent tasks, and quality metadata\. Table[2](https://arxiv.org/html/2609.19180#S3.T2)gives the main fields\. The source\-of\-truth schema is implemented in Pydantic and exported as JSON Schema; unknown fields are forbidden and cross\-field invariants are checked during validation\. The expected model output for each task includes both an answer and supporting evidence IDs, making evidence attribution part of the task rather than post\-hoc explanation\.

Table 2:Dataset schema of BioPhys\-Bridge\. Evidence IDs, quantitative values, equations, assumptions, mechanisms, and decisions are explicit grounding targets\.
### 3\.3Task Types

The release contains four task types, each designed to test a different grounding challenge\.derivationtasks ask a model to use a physical relation or equation while citing the relevant evidence IDs\.mechanism\_from\_evidencetasks require a biological mechanism inferred from cited quantitative or textual evidence\.discrepancy\_explanationtasks require an explanation of tension between expected physical behavior and observed evidence, grounded in the case’s evidence blocks\.next\_experiment\_designtasks require a follow\-up experiment or computation grounded in cited evidence and model assumptions\.

## 4Data Curation Pipeline

The data curation pipeline converts scientific research papers into grounding benchmark records\. It is designed to be traceable, resumable, and compliant with source licensing\. Figure[2](https://arxiv.org/html/2609.19180#S4.F2)sketches the main stages; the archived release audit funnel is reported in Appendix[B](https://arxiv.org/html/2609.19180#A2)\.

![Refer to caption](https://arxiv.org/html/2609.19180v1/figure3_pipeline.png)Figure 2:Curation pipeline\. Open\-access papers are parsed with MinerU, normalized into evidence blocks, structured into case records, and filtered by validation and quality gates before release\.#### Source acquisition and batching\.

Candidate papers are drawn from open\-access scientific sources with DOI and/or PMCID provenance and release\-compatible licenses\. Candidate metadata includes a title, source URL, license, abstract, domain guess, model\-family guess, and model or quantitative keywords\. Papers are selected in batches while tracking coverage across domains and physical model families\.

#### Document parsing and evidence block construction\.

Each PDF is parsed with MinerU\. The raw PDF files and MinerU intermediate payloads are retained outside the public release, while normalized evidence artifacts are used to build release records\. Normalization converts parsed outputs into canonical document folders with markdown, structured JSON, evidence blocks, and parse metadata\. Evidence blocks are preserved as the atomic citation units used by model prompts and evidence\-ID scoring\.

#### Grounded candidate extraction and structuring\.

A regex\-first pass extracts candidate numeric values, equations, units, model keywords, and evidence\-rich snippets from prose, tables, formulas, and figure captions\. An evidence\-only LLM pass may then structure fields such as quantitative evidence, physical interpretation, mechanism text, and agent tasks\. The shipped release used OpenAIgpt\-4ofor this evidence\-only structuring pass\. The prompt instructs the model to use only supplied evidence and to leave unsupported fields empty rather than hallucinate measurements, equations, assumptions, or mechanisms\. Unsupported or fabricated evidence IDs are removed by release validation\.

#### Release export\.

Final export writes JSONL records, deterministic splits, metadata, schema, a dataset card, gold samples, and a Sci\-Evo view\. Source PDFs, raw MinerU payloads, and LLM responses are not redistributed\.

#### Expert annotation layer\.

Expert annotation is tracked separately from release\-gate review\. The full held\-out test split, 10 contest gold samples, 30 extended\-gold samples, and one legacy reviewed record include physics reasoning, biological reasoning, uncertainty, and reviewer notes\. This separation prevents the release\-gate fieldquality\.manual\_review\_statusfrom being mistaken for a claim that all 500 cases have full expert biological and physical review\.

## 5Release Statistics

Table[3](https://arxiv.org/html/2609.19180#S5.T3)summarizes the current release\. The dataset contains 500 cases and 1,517 agent tasks\. The split is deterministic and stratified by domain, with 400 training cases, 50 validation cases, and 50 test cases\. It is also grouped by source paper under DOI/PMCID/title keys: the train, validation, and test splits share zero source papers\. The release includes 107 cases with an explicit failure or revision stage, and 81 cases with expert annotation\. The expert annotation count covers the 50 held\-out test cases, 10 contest gold samples, 30 extended\-gold samples, and one legacy reviewed record\.

Table 3:BioPhys\-Bridge v1 release at a glance\.Table[4](https://arxiv.org/html/2609.19180#S5.T4)shows the weighted domain coverage\. The release is not intended to be perfectly balanced\. Instead, it emphasizes coverage across biophysical discovery settings while preserving enough examples in the larger protein\-ligand and systems\-biology groups to support model training and evaluation\.

Table 4:Biological\-domain coverage in the 500\-case release\.Table 5:Agent\-task distribution\.The primary physical model families include binding thermodynamics \(185\), systems stochastic dynamics \(128\), conformational/allosteric energy landscapes \(93\), polymer phase\-separation statistical mechanics \(41\), enzyme reaction kinetics \(23\), folding stability thermodynamics \(20\), spatial transport and electrostatics \(6\), evolutionary fitness landscapes \(2\), and mechanical force response \(2\)\. The last three families are included for coverage and future expansion, not for standalone per\-family evaluation in this release\.

## 6Quality Control and Validation

Quality control is central to the release\. Table[6](https://arxiv.org/html/2609.19180#S6.T6)summarizes the main gates\. A shipped case must validate against the schema, cite valid evidence IDs, ground quantitative values in cited evidence text, carry source license metadata, pass duplicate checks, and avoid unresolved templates or weak task prompts\. Unit normalization retains both raw and normalized units when normalization is supported\. The content gate blocks examples with unsupported measurements, fabricated evidence IDs, evidence\-ID\-only prompts, unresolved template strings, malformed tool vocabularies, and missing next\-step stages\.

Table 6:Release quality gates\. The deterministic audit notes are a supporting traceability aid; only about 2% of cases have an implemented closed\-form relation\-level check, so it is not used as a global correctness metric\.The fieldquality\.manual\_review\_status = reviewedshould be read as a release\-gate status, not as a claim that all 500 cases have full expert physics and biology annotation\. Full expert annotations are tracked separately through theexpert\_annotationfield\.

## 7Evaluation Protocol and Baselines

#### Evaluation setting\.

Given a task prompt and candidate evidence blocks, the model must produce a JSON answer and a set of supporting evidence IDs\. The scorer computes evidence\-ID F1 against the gold supporting evidence IDs and answer token F1 against the gold answer\. Evidence\-ID F1 evaluates whether the model grounds its answer in the correct evidence blocks\. Answer token F1 is a rough sanity\-check metric for this initial release, not a final measure of scientific correctness\. The released harness also computesoverall\_scoreas0\.70\.7answer token F1 plus0\.30\.3evidence\-ID F1 for diagnostics, but we do not use that composite metric for the main paper because token F1 is weak for open\-ended scientific answers\.

#### De\-leaked prompt construction\.

The evaluation prompt hides gold answers, gold supporting evidence IDs, and expert annotations\. The main de\-leaked prompt exposes the task domain, task type, question, and ranked candidate evidence blocks, but not the structured equation, physical directionality, or biological mechanism fields\. We use this no\-scaffold prompt for the main results because those structured fields can be answer\-bearing for derivation or mechanism tasks\. Candidate evidence selection uses a lexical ranker over only the task type, task question, research question, domain, and source evidence text; it does not accesstask\.supporting\_evidence\_idsor structured target annotations such as equations, physical directionality, biological mechanisms, or quantitative value records\. Gold IDs are used only by the scorer after prediction\. In the released harness, prompt construction and candidate selection operate on public case/task views that excludegold\_answer,supporting\_evidence\_ids, andexpert\_annotation; regression tests fail if these fields reach the model\-visible evaluation path\. Each test task is shown 48 candidate evidence blocks selected from a median 206 evidence blocks per case; the deterministic lexical baseline cites the top 3 ranked blocks\. This candidate generator contains 234 of 267 gold supporting evidence IDs on the test set \(0\.876 gold\-ID recall\), with all gold IDs present for 127 of 154 tasks; the mean maximum achievable evidence\-ID F1 under the candidate set is 0\.916\. Thus the reported scores evaluate attribution and re\-ranking within a lexically pre\-filtered candidate set, not unconstrained retrieval over every evidence block\. Regression tests verify that prompts and the lexical baseline do not copy gold evidence IDs\. All LLM rows use temperature 0 and JSON\-mode decoding\. Confidence intervals are nonparametric bootstrap 95% intervals over tasks\.

Table 7:Preliminary de\-leaked no\-scaffold evaluation on the identical 154\-task held\-out test set\. Values in brackets are bootstrap 95% confidence intervals over tasks\. Empty EID outputs counts predictions with no parsedsupporting\_evidence\_ids; this is an output\-compliance measure as well as an attribution failure mode\.Paired bootstrap over the same 154 tasks gives evidence\-ID F1 deltas over lexical retrieval of\+0\.172\+0\.172for DeepSeek v4 Flash \(95% CI\[0\.117,0\.227\]\[0\.117,0\.227\]\),\+0\.127\+0\.127for Qwen 3\.7 Max \(95% CI\[0\.076,0\.179\]\[0\.076,0\.179\]\), and\+0\.106\+0\.106for GPT\-4o\-mini \(95% CI\[0\.061,0\.152\]\[0\.061,0\.152\]\)\. Claude Opus 4\.8 shows a smaller positive delta \(\+0\.050\+0\.050, 95% CI\[0\.005,0\.094\]\[0\.005,0\.094\]\), while Claude Sonnet 5 is not significantly above lexical retrieval under this paired bootstrap \(\+0\.035\+0\.035, 95% CI\[−0\.010,0\.081\]\[\-0\.010,0\.081\]\)\. DeepSeek v4 Flash also outperforms Qwen 3\.7 Max by\+0\.045\+0\.045evidence\-ID F1 \(95% CI\[0\.012,0\.079\]\[0\.012,0\.079\]\)\. We retain a scaffolded prompt mode for diagnostic audits, but exclude it from the main baseline because structured fields can be answer\-bearing, including equations that match derivation gold answers\.

Table 8:Held\-out evidence\-ID F1 by task type for the lexical floor and the three strongest LLM rows in Table[7](https://arxiv.org/html/2609.19180#S7.T7)\. Each task\-type cell contains only 31–46 tasks, so these point estimates should be read as noisier diagnostics rather than stable per\-type rankings\.These baselines are intentionally preliminary\. The same\-test comparison shows that several LLMs improve over lexical retrieval in aggregate evidence attribution, with the largest gains onderivation,discrepancy\_explanation, andmechanism\_from\_evidencetasks\.next\_experiment\_designremains harder: even the strongest model reaches only 0\.216 evidence\-ID F1, consistent with the task’s open\-ended decision structure\. The panel also exposes a grounded\-generation failure mode: some models produce plausible prose but fail to emit usable evidence IDs\. Gemini 2\.5 Pro has nontrivial answer\-token overlap \(0\.099\) but empty parsed evidence IDs for 151 of 154 tasks, making its attribution score near zero\. DeepSeek v4 Pro similarly leaves 127 of 154 outputs without parsed evidence IDs\. Answer token F1 therefore remains a weak proxy for grounded scientific answer quality, while evidence\-ID F1 captures whether the answer is actually attributable to the supplied evidence blocks\.

Because scientific answers are open\-ended, future releases should add rubric\-based expert scoring for physical\-model use, quantitative consistency, mechanism correctness, uncertainty handling, and experimental feasibility\. This is especially important for grounded scientific agents, where a cited but mechanistically wrong answer can still receive partial evidence overlap\. The 50 test cases have expert annotation notes, but they are not independent parallel task\-level evidence\-ID labels, so we do not report inter\-annotator agreement for supporting evidence IDs in this release\.

## 8Dataset Access and Reproducibility

For peer review, the dataset, schema, aggregate reports, and evaluation outputs are supplied in the supplementary materials\. The public dataset and code repository will be released after peer review\. The repository will contain code, schema, reports, validation tests, small samples, aggregate summaries, and release files\. Code will be released under the MIT License, while curated dataset metadata, reports, and public release files will be released under CC\-BY\-4\.0\. Upstream source papers retain their original licenses, recorded per case\. Raw PDFs, raw MinerU payloads, and LLM responses are not redistributed\.

The public release will expose validation and evaluation commands through thebiophysevoPython package\. The source\-of\-truth schema lives insrc/biophysevo/schemas/case\_schema\.py, and the generated JSON Schema is included in the supplementary materials\.

## 9Limitations and Responsible Use

BioPhys\-Bridge is an initial benchmark release, not a substitute for reading the original papers and not a complete solution to biological reasoning\. Users should verify important claims against the cited source literature before using them in research, engineering, clinical, or safety\-critical settings\. Expert annotation covers 81 cases, not all 500\. Current baselines are preliminary: under the fully de\-leaked public\-query setting, they include a lexical retrieval floor and seven LLMs, but not a human ceiling, hidden test server, repeated stochastic runs, or full rubric\-based expert grading\. Evidence\-ID F1 captures attribution to evidence blocks but not full scientific correctness, and token F1 is a weak proxy for open\-ended answers\. The expert layer checks physical and biological reasoning for the test cases, but it does not provide independent parallel task\-level labels for supporting evidence IDs; future versions should add inter\-annotator agreement and expert scoring for mechanism correctness\. Because the release structuring pass used GPT\-4o, shared\-provider bias remains a caveat for the GPT\-4o\-mini row, although the strongest reported evidence\-ID scores come from non\-OpenAI systems\. Model routing and output\-format behavior also affect measured attribution: systems that omit parseablesupporting\_evidence\_idscan receive very low evidence\-ID F1 even when their prose answer overlaps the reference answer\.

The test set is public in this release, which supports reuse and workshop discussion but limits long\-term contamination control; future versions should include hidden or contamination\-aware evaluation\. The source pool is open\-access and release\-compatible, which induces domain and source bias\. Domain coverage is weighted rather than balanced\. OCR and table parsing artifacts can remain in evidence text, because preserving traceability to parsed source artifacts is prioritized over silently rewriting source evidence\. The archived release metadata begins at the reviewed candidate set rather than the full exploratory source pool, so future versions should report complete rejection and parser\-failure counts\.

## 10Conclusion

We introduce BioPhys\-Bridge, an initial benchmark for evidence\-grounded scientific reasoning over biophysical literature\. The current 500\-case release is reusable and validated enough for community exploration: it represents evidence blocks, quantitative values, equations, assumptions, mechanisms, and next scientific decisions in a shared case schema, and it provides de\-leaked baselines for evidence\-ID attribution\. Preliminary same\-test results show that several LLMs improve over lexical retrieval for evidence attribution, with DeepSeek v4 Flash, Qwen 3\.7 Max, and GPT\-4o\-mini producing the strongest evidence\-ID F1 in our initial panel\. The results also show that grounded answering is not only answer generation: models can produce plausible prose while failing to emit usable evidence IDs\. This workshop version is intended to invite feedback on evaluation design, expert\-rubric scoring, contamination\-aware protocols, and future grounded scientific\-agent benchmarks that test not only what a paper says, but how source evidence supports quantitative and mechanistic scientific decisions\.

## Acknowledgments

Acknowledgments are omitted for anonymous review\.

## Declaration on Generative AI

During the preparation of this work, the author used ChatGPT for brainstorming, paper drafting and editing, and code scaffolding, and used Claude Code for code implementation\. The author reviewed and edited the entire content and takes full responsibility for the publication’s content\.

## References

- Ajithet al\.\(2024\)A\. Ajith, M\. Xia, A\. Chevalier, T\. Goyal, D\. Chen, and T\. GaoLitSearch: a retrieval benchmark for scientific literature search\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,Miami, Florida, USA,pp\. 15068–15083\.External Links:[Link](https://aclanthology.org/2024.emnlp-main.840/),[Document](https://dx.doi.org/10.18653/v1/2024.emnlp-main.840)Cited by:[§2\.1](https://arxiv.org/html/2609.19180#S2.SS1.p1.1)\.
- Baiet al\.\(2024\)Y\. Bai, X\. Lv, J\. Zhang, H\. Lyu, J\. Tang, Z\. Huang, Z\. Du, X\. Liu, A\. Zeng, L\. Hou, Y\. Dong, J\. Tang, and J\. LiLongBench: a bilingual, multitask benchmark for long context understanding\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),Bangkok, Thailand,pp\. 3119–3137\.External Links:[Link](https://aclanthology.org/2024.acl-long.172/),[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.172)Cited by:[§2\.2](https://arxiv.org/html/2609.19180#S2.SS2.p1.1)\.
- Chanet al\.\(2025\)J\. S\. Chan, N\. Chowdhury, O\. Jaffe, J\. Aung, D\. Sherburn, E\. Mays, G\. Starace, K\. Liu, L\. Maksin, T\. Patwardhan, A\. Madry, and L\. WengMLE\-bench: evaluating machine learning agents on machine learning engineering\.InInternational Conference on Learning Representations,Cited by:[§2\.3](https://arxiv.org/html/2609.19180#S2.SS3.p1.1)\.
- Hassanet al\.\(2023\)S\. M\. S\. Hassan, A\. Feeney, A\. Dhruv, J\. Kim, Y\. Suh, J\. Ryu, Y\. Won, and A\. ChandramowlishwaranBubbleML: a multiphase multiphysics dataset and benchmarks for machine learning\.InAdvances in Neural Information Processing Systems, Datasets and Benchmarks Track,Cited by:[§2\.3](https://arxiv.org/html/2609.19180#S2.SS3.p1.1),[Table 1](https://arxiv.org/html/2609.19180#S2.T1.2.5.1.1.1)\.
- Jimenezet al\.\(2024\)C\. E\. Jimenez, J\. Yang, A\. Wettig, S\. Yao, K\. Pei, O\. Press, and K\. NarasimhanSWE\-bench: can language models resolve real\-world github issues?\.InInternational Conference on Learning Representations,Cited by:[§2\.3](https://arxiv.org/html/2609.19180#S2.SS3.p1.1)\.
- Lalaet al\.\(2023\)J\. Lala, O\. O’Donoghue, A\. Shtedritski, S\. Cox, S\. G\. Rodriques, and A\. D\. WhitePaperQA: retrieval\-augmented generative agent for scientific research\.External Links:2312\.07559Cited by:[§2\.1](https://arxiv.org/html/2609.19180#S2.SS1.p1.1),[Table 1](https://arxiv.org/html/2609.19180#S2.T1.2.2.1.1.1)\.
- Laurentet al\.\(2024\)J\. M\. Laurent, J\. D\. Janizek, M\. Ruzo, M\. M\. Hinks, M\. J\. Hammerling, S\. Narayanan, M\. Ponnapati, A\. D\. White, and S\. G\. RodriquesLAB\-bench: measuring capabilities of language models for biology research\.External Links:2407\.10362Cited by:[§2\.1](https://arxiv.org/html/2609.19180#S2.SS1.p1.1),[Table 1](https://arxiv.org/html/2609.19180#S2.T1.2.2.1.1.1)\.
- Liuet al\.\(2024\)X\. Liu, H\. Yu, H\. Zhang, Y\. Xu, X\. Lei, H\. Lai, Y\. Gu, H\. Ding, K\. Men, K\. Yang,et al\.AgentBench: evaluating llms as agents\.InInternational Conference on Learning Representations,Cited by:[§2\.3](https://arxiv.org/html/2609.19180#S2.SS3.p1.1)\.
- Luet al\.\(2024\)P\. Lu, H\. Bansal, T\. Xia, J\. Liu, C\. Li, H\. Hajishirzi, H\. Cheng, K\. Chang, M\. Galley, and J\. GaoMathVista: evaluating mathematical reasoning of foundation models in visual contexts\.InInternational Conference on Learning Representations,Cited by:[§2\.2](https://arxiv.org/html/2609.19180#S2.SS2.p1.1),[Table 1](https://arxiv.org/html/2609.19180#S2.T1.2.4.1.1.1)\.
- Luet al\.\(2022\)P\. Lu, S\. Mishra, T\. Xia, L\. Qiu, K\. Chang, S\. Zhu, O\. Tafjord, P\. Clark, and A\. KalyanLearn to explain: multimodal reasoning via thought chains for science question answering\.InAdvances in Neural Information Processing Systems,Vol\.35,pp\. 2507–2521\.External Links:2209\.09513Cited by:[§2\.2](https://arxiv.org/html/2609.19180#S2.SS2.p1.1),[Table 1](https://arxiv.org/html/2609.19180#S2.T1.2.3.1.1.1)\.
- Pramanicket al\.\(2024\)S\. Pramanick, R\. Chellappa, and S\. VenugopalanSPIQA: a dataset for multimodal question answering on scientific papers\.InAdvances in Neural Information Processing Systems, Datasets and Benchmarks Track,Cited by:[§2\.1](https://arxiv.org/html/2609.19180#S2.SS1.p1.1),[Table 1](https://arxiv.org/html/2609.19180#S2.T1.2.4.1.1.1)\.
- Qinet al\.\(2024\)Y\. Qin, S\. Liang, Y\. Ye, K\. Zhu, L\. Yan, Y\. Lu, Y\. Lin, X\. Cong, X\. Tang, B\. Qian, S\. Zhao, L\. Hong, R\. Tian, R\. Xie, J\. Zhou, M\. Gerstein, D\. Li, Z\. Liu, and M\. SunToolLLM: facilitating large language models to master 16000\+ real\-world apis\.InInternational Conference on Learning Representations,Cited by:[§2\.3](https://arxiv.org/html/2609.19180#S2.SS3.p1.1)\.
- Reinet al\.\(2024\)D\. Rein, B\. L\. Hou, A\. C\. Stickland, J\. Petty, R\. Y\. Pang, J\. Dirani, J\. Michael, and S\. R\. BowmanGPQA: a graduate\-level google\-proof q&a benchmark\.InFirst Conference on Language Modeling,External Links:[Link](https://openreview.net/forum?id=Ti67584b98),2311\.12022Cited by:[§2\.2](https://arxiv.org/html/2609.19180#S2.SS2.p1.1),[Table 1](https://arxiv.org/html/2609.19180#S2.T1.2.3.1.1.1)\.
- Takamotoet al\.\(2022\)M\. Takamoto, T\. Praditia, R\. Leiteritz, D\. MacKinlay, F\. Alesiani, D\. Pflüger, and M\. NiepertPDEBench: an extensive benchmark for scientific machine learning\.InAdvances in Neural Information Processing Systems, Datasets and Benchmarks Track,Cited by:[§2\.3](https://arxiv.org/html/2609.19180#S2.SS3.p1.1),[Table 1](https://arxiv.org/html/2609.19180#S2.T1.2.5.1.1.1)\.
- Wanget al\.\(2024a\)B\. Wang, C\. Xu, X\. Zhao, L\. Ouyang, F\. Wu, Z\. Zhao, R\. Xu, K\. Liu, Y\. Qu, F\. Shang,et al\.MinerU: an open\-source solution for precise document content extraction\.External Links:2409\.18839Cited by:[§2\.1](https://arxiv.org/html/2609.19180#S2.SS1.p2.1)\.
- Wanget al\.\(2024b\)X\. Wang, Z\. Hu, P\. Lu, Y\. Zhu, J\. Zhang, S\. Subramaniam, A\. R\. Loomba, S\. Zhang, Y\. Sun, and W\. WangSciBench: evaluating college\-level scientific problem\-solving abilities of large language models\.InProceedings of the 41st International Conference on Machine Learning,pp\. 50622–50649\.External Links:2307\.10635Cited by:[§2\.2](https://arxiv.org/html/2609.19180#S2.SS2.p1.1),[Table 1](https://arxiv.org/html/2609.19180#S2.T1.2.3.1.1.1)\.
- Wanget al\.\(2024c\)X\. Wang, Z\. Wang, J\. Liu, Y\. Chen, L\. Yuan, H\. Peng, and H\. JiMINT: evaluating llms in multi\-turn interaction with tools and language feedback\.InInternational Conference on Learning Representations,Cited by:[§2\.3](https://arxiv.org/html/2609.19180#S2.SS3.p1.1)\.
- Yueet al\.\(2024\)X\. Yue, Y\. Ni, K\. Zhang, T\. Zheng, R\. Liu, G\. Zhang, S\. Stevens, D\. Jiang, W\. Ren, Y\. Sun,et al\.MMMU: a massive multi\-discipline multimodal understanding and reasoning benchmark for expert agi\.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition,Cited by:[§2\.2](https://arxiv.org/html/2609.19180#S2.SS2.p1.1),[Table 1](https://arxiv.org/html/2609.19180#S2.T1.2.4.1.1.1)\.
- Zhouet al\.\(2024\)S\. Zhou, F\. F\. Xu, H\. Zhu, X\. Zhou, R\. Lo, A\. Sridhar, X\. Cheng, T\. Ou, Y\. Bisk, D\. Fried, U\. Alon, and G\. NeubigWebArena: a realistic web environment for building autonomous agents\.InInternational Conference on Learning Representations,Cited by:[§2\.3](https://arxiv.org/html/2609.19180#S2.SS3.p1.1)\.

## Appendix AGold Case Example

![Refer to caption](https://arxiv.org/html/2609.19180v1/figure2_golden_case.png)Figure 3:A real BioPhys\-Bridge case illustrates the benchmark object: source provenance, evidence\-linked quantitative measurements, a physical model, mechanism interpretation, caveats, and an agent\-facing task are stored together rather than flattened into a standalone question\.
## Appendix BRelease Audit Funnel

Table 9:Release audit funnel recorded in the public metadata\. Future dataset versions should additionally archive exploratory source\-pool and parser\-failure counts to support a full rejection funnel\.
## Appendix CPrompt Template

The evaluation prompt is generated by the released harness as follows:

> You are answering an evidence\-grounded scientific question for biological research\. Use only the evidence below\. Return JSON with keys answer and supporting\_evidence\_ids\. Choose supporting\_evidence\_ids from the candidate evidence IDs below; gold IDs are not provided\. Domain: <domain\>\. Task type: <task\_type\>\. Evidence: <ranked evidence blocks\>\. Question: <task input\>\.

For evaluation, gold answers, gold supporting evidence IDs, and expert annotations are hidden from the prompt; gold labels are used only by the scorer\. The harness also retains a scaffolded diagnostic mode that addsPhysical model,Equation,Physical directionality, andBiological mechanismlines; it is not used for the main de\-leaked results because those fields can be answer\-bearing\.

Similar Articles

PlantMarkerBench: A Multi-Species Benchmark for Evidence-Grounded Plant Marker Reasoning

Hugging Face Daily Papers

This paper introduces PlantMarkerBench, a multi-species benchmark for evaluating language models' ability to interpret evidence for plant marker genes from scientific literature across four species. It highlights that while frontier models perform well on direct evidence, they struggle with functional and indirect evidence types.