Tag
This paper introduces RENDEQ, a generator of render-equivalence sets for scientific figures, and measures how well model agreement across semantics-preserving re-renderings tracks correctness in open-weight VLMs. It finds that agreement certifies correctness only above a threshold and that fine-tuning on self-consensus can hurt accuracy.
This paper presents the ICDAR 2026 Competition on Information Extraction from ALD/E Scientific Figures, introducing the Sci-ImageMiner benchmark with four complementary tasks. Results show SOTA multimodal models perform well on classification and summarization but struggle with data extraction and scientific reasoning, especially visual question answering.
Introduces SciDraw-Bench, a benchmark for evaluating scientific figure generation by text-to-image and multimodal models, with a four-dimensional evaluation protocol. Findings show domain-specific systems outperform general-purpose models, with text fidelity remaining the hardest challenge.
Introduces MINARD, a pipeline for generating narrated, region-grounded walkthrough videos from scientific figures and their papers, along with the FigTalk benchmark and new grounding metrics.
This is a tool that automatically converts scientific paper charts into executable Python plotting code, using the Qwen vision model and Codex agent for panel segmentation, code generation, and refinement.