PlantMarkerBench: A Multi-Species Benchmark for Evidence-Grounded Plant Marker Reasoning
Summary
This paper introduces PlantMarkerBench, a multi-species benchmark for evaluating language models' ability to interpret evidence for plant marker genes from scientific literature across four species. It highlights that while frontier models perform well on direct evidence, they struggle with functional and indirect evidence types.
View Cached Full Text
Cached at: 05/13/26, 12:20 AM
Paper page - PlantMarkerBench: A Multi-Species Benchmark for Evidence-Grounded Plant Marker Reasoning
Source: https://huggingface.co/papers/2605.10032
Abstract
PlantMarkerBench presents a multi-species benchmark for evaluating literature-based plant marker evidence interpretation, assessing models on identifying valid marker evidence and categorizing evidence types across four plant species.
Cell-type-specific marker genes are fundamental to plant biology, yet existing resources primarily rely on curated databases or high-throughput studies without explicitly modeling the supporting evidence found in scientific literature. We introduce PlantMarkerBench, a multi-species benchmark for evaluating literature-grounded plant marker evidence interpretation from full-text biological papers. PlantMarkerBench is constructed using a modular curation pipeline integrating large-scale literature retrieval,hybrid search, species-awarebiological grounding,structured evidence extraction, and targeted human review. The benchmark spans four plant species -- Arabidopsis, maize, rice, and tomato -- and contains 5,550sentence-level evidence instancesannotated formarker-evidence validity, evidence type, and support strength. We define two benchmark tasks: determining whether a candidate sentence provides valid marker evidence for a gene-cell-type pair, and classifying the evidence into expression, localization, function, indirect, or negative categories. We benchmark diverse open-weight andclosed-source language modelsacross species andprompting strategies. Although frontier models achieve relatively strong performance on direct expression evidence, performance drops substantially on functional, indirect, and weak-support evidence, with evidence-type confusion emerging as a dominant failure mode.Open-weight modelsadditionally exhibit elevatedfalse-positive ratesunder ambiguous biological contexts. PlantMarkerBench provides a challenging and reproducible evaluation framework forliterature-grounded biological evidenceattribution and supports future research on trustworthy scientific information extraction and AI-assisted plant biology.
View arXiv pageView PDFProject pageAdd to collection
Get this paper in your agent:
hf papers read 2605\.10032
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2605.10032 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2605.10032 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2605.10032 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
M$^3$R-Bench: A Unified Benchmark for Evidence-Grounded Multimodal Metaphor Understanding
This paper introduces M3R-Bench, a unified evidence-grounded benchmark for multimodal metaphor understanding with 1,000 image-text instances, and proposes M3R-Reasoner, an 8B-parameter model combining curriculum-based reasoning supervision and reinforcement learning that outperforms larger proprietary models.
BenchBench-Protocol: Evaluating Real-World Wet-Lab Protocol Reasoning and Modification
The article introduces BenchBench-Protocol, a benchmark of 149 real-world wet-lab protocol-modification tasks to evaluate large language models' reasoning, with Claude Opus 5 scoring highest at 59.2%.
EvalDetectBench: A Benchmark for Measuring Evaluation Awareness in Frontier Language Models
This paper introduces EvalDetectBench, an open benchmark and pipeline for measuring evaluation awareness in frontier language models, addressing biases in existing methods to improve AI safety assessments.
BEAR-Bench: A Bilingual Enterprise and Academic Reasoning Benchmark for Multimodal Models
BEAR-Bench introduces a bilingual English-and-Russian benchmark for evaluating multimodal models' reasoning on text-rich professional documents, assessing 16 models and highlighting performance gaps.
Can AI Reason Like an Urban Planner? Benchmarking Large Language Models Against Professional Judgment
This paper introduces UPBench, a benchmark to evaluate large language models on urban planning knowledge across four knowledge pillars and five cognitive levels, finding that models perform better on higher-order analysis than factual recall, and identifying epistemic limitations such as regulatory hallucination and phronetic deficit.