Tag
WuYuEval is a multi-level benchmark for evaluating large language models in solid waste management, covering foundational knowledge, domain reasoning, and expert decision-making. It includes 4,590 multiple-choice and 247 scenario-based open-ended questions, and reports performance across 33 LLMs.
This paper presents the ICDAR 2026 Competition on Information Extraction from ALD/E Scientific Figures, introducing the Sci-ImageMiner benchmark with four complementary tasks. Results show SOTA multimodal models perform well on classification and summarization but struggle with data extraction and scientific reasoning, especially visual question answering.