Tag
Fig 和 Georgia Tech 的技术报告评估了 Astra、Opus 5.5 等前沿模型在网页自动化与物理任务(操作、装配、工业流程、驾驶)上的表现,发现模型性能呈'锯齿状'分布,没有明确的性能领先者,并发布 item 级结果数据集 RIDGE。
Introduces Uncheatable Eval, a dynamic benchmark using compression rates to evaluate language models and mitigate data contamination.
The paper proposes Tempered Evidence Fusion (TEF), a method for long-text value measurement that weights sentence-level LLM judgments by information gain, and introduces the MIND benchmark dataset for evaluation.
BioPhys-Bridge introduces a benchmark dataset for evaluating language models on evidence-grounded scientific reasoning in biophysics, focusing on interdisciplinary tasks with quantitative and mechanistic grounding.
FINESSE is an agent-based simulation framework and benchmark dataset for generating synthetic multimodal financial event sequences, addressing data scarcity and supporting tasks like fraud detection and balance forecasting.
This paper introduces DasanCallDial, a large-scale Korean benchmark dataset for dialogue-level ASR error correction, and proposes the DCSC framework, achieving state-of-the-art performance in text-only post-editing.
STAIR introduces a structure-aware retrieval system that leverages Table of Contents for enhanced document retrieval in Large Language Models, achieving high performance on the new SearchTome benchmark dataset.
This paper investigates the structural sensitivity of multilingual large language models to semantics-preserving perturbations in Hindi and Malayalam, showing significant degradation in mathematical reasoning performance and introducing the IndicReStruct benchmark for evaluation.
This paper introduces a benchmark dataset for Bangla idioms and evaluates recent large language models on idiom-related tasks, revealing substantial variability in performance across models.
The paper presents Chiaro, a new benchmark dataset for contrastive emotion recognition where two individuals experience opposing emotions from a shared event, grounded in appraisal theory. It evaluates seven LLMs and four emotion classifiers, revealing that current models fall short of human performance.
OmniPhys is a large-scale multimodal benchmark for physics understanding and generation, covering middle school to university-level problems from Chinese educational corpora, aimed at evaluating and advancing multimodal large language models in scientific domains.
This paper introduces grounded glossary generation for Classical Sanskrit, a task involving recovering Sanskrit phrases and producing translation-grounded meanings from sloka-translation pairs. It constructs a benchmark from Hindu texts and evaluates various AI models, finding that instruction fine-tuning improves performance, with morphological modeling identified as a key challenge.
This paper introduces FaceVid-Forensics-100K, a large-scale deepfake video dataset with fine-grained annotations, and proposes a multi-agent forensic reasoning framework (ARGUS) that uses four specialized expert agents and a judge agent to outperform closed-source models on deepfake detection.
A new mixture-of-experts approach for reconstructing handwriting trajectories from IMU sensor data, with separate experts for touching and hovering phases, and a new public benchmark dataset.
PhononBench-MP40 is a benchmark dataset of phonon stability labels and spectra for over 46,000 Materials Project-derived crystals, designed to evaluate workflow-defined phonon stability and support materials screening.
This paper introduces a benchmark dataset of 1,516 expert-verified Bangla sentences for disambiguating culturally entangled homographs (words that are both names and common nouns). It shows that LLMs suffer from dominant-meaning bias and proposes contrastive chain-of-thought prompting and distillation to reduce this bias.
Introduces DECODEM, benchmark datasets for evaluating automated extraction of corporate governance variables from legal documents using large language models, showing high accuracy for many provisions.
Introducing GPIC (Giant Permissive Image Corpus), a large-scale dataset of 100M VLM-captioned image-text pairs for training and 1M pairs for benchmarking, fully permissive for research and commercial use.
This article introduces EmoS, a high-fidelity multimodal benchmark designed for fine-grained streaming emotional understanding, addressing limitations in ecological validity and labeling reliability found in existing datasets.
This paper introduces SynopticBench, a dataset of 1.3M+ weather forecast discussions paired with meteorological images, and SPACE, a novel evaluation framework for assessing VLM-generated weather forecasts.