benchmark-dataset

Tag

Cards List
#benchmark-dataset

Astra, Opus 5.5, and other Frontier Models Demonstrate Jagged Performance Across SoTA Agentic Tasks from Web Browsing to Robotics (32 minute read)

TLDR AI ↗ · 4d ago Cached

Fig 和 Georgia Tech 的技术报告评估了 Astra、Opus 5.5 等前沿模型在网页自动化与物理任务(操作、装配、工业流程、驾驶)上的表现,发现模型性能呈'锯齿状'分布,没有明确的性能领先者,并发布 item 级结果数据集 RIDGE。

0 favorites 0 likes
#benchmark-dataset

Uncheatable Eval: Dynamic Compression-Based Evaluation of Language Models

arXiv cs.CL ↗ · 2026-09-24 Cached

Introduces Uncheatable Eval, a dynamic benchmark using compression rates to evaluate language models and mitigate data contamination.

0 favorites 0 likes
#benchmark-dataset

Count Evidence, Not Sentences: Tempered Evidence Fusion of LLM Judgments for Long-Text Value Measurement

arXiv cs.CL ↗ · 2026-09-24 Cached

The paper proposes Tempered Evidence Fusion (TEF), a method for long-text value measurement that weights sentence-level LLM judgments by information gain, and introduces the MIND benchmark dataset for evaluation.

0 favorites 0 likes
#benchmark-dataset

BioPhys-Bridge: A Benchmark for Interdisciplinary Scientific Reasoning in Physics-Grounded Biological Research

arXiv cs.AI ↗ · 2026-09-18 Cached

BioPhys-Bridge introduces a benchmark dataset for evaluating language models on evidence-grounded scientific reasoning in biophysics, focusing on interdisciplinary tasks with quantitative and mechanistic grounding.

0 favorites 0 likes
#benchmark-dataset

FINESSE: An Agent-Based Simulator and Benchmark Dataset for Multimodal Financial Event Sequences

arXiv cs.LG ↗ · 2026-09-14 Cached

FINESSE is an agent-based simulation framework and benchmark dataset for generating synthetic multimodal financial event sequences, addressing data scarcity and supporting tasks like fraud detection and balance forecasting.

0 favorites 0 likes
#benchmark-dataset

Leveraging Fine-grained Error Correction in Korean Speech Recognition for Consultation Services

arXiv cs.CL ↗ · 2026-09-10 Cached

This paper introduces DasanCallDial, a large-scale Korean benchmark dataset for dialogue-level ASR error correction, and proposes the DCSC framework, achieving state-of-the-art performance in text-only post-editing.

0 favorites 0 likes
#benchmark-dataset

STAIR (STructure Aware Information Retriever): A novel dataset and LLM based retriever for document structure augmentation

arXiv cs.AI ↗ · 2026-09-04 Cached

STAIR introduces a structure-aware retrieval system that leverages Table of Contents for enhanced document retrieval in Large Language Models, achieving high performance on the new SearchTome benchmark dataset.

0 favorites 0 likes
#benchmark-dataset

Lost in Reordering: Structural Sensitivity of Multilingual LLMs under Semantics-Preserving Perturbations

arXiv cs.CL ↗ · 2026-09-04 Cached

This paper investigates the structural sensitivity of multilingual large language models to semantics-preserving perturbations in Hindi and Malayalam, showing significant degradation in mathematical reasoning performance and introducing the IndicReStruct benchmark for evaluation.

0 favorites 0 likes
#benchmark-dataset

To What Extent Do Large Language Models Understand Bangla Idioms?

arXiv cs.CL ↗ · 2026-09-04 Cached

This paper introduces a benchmark dataset for Bangla idioms and evaluates recent large language models on idiom-related tasks, revealing substantial variability in performance across models.

0 favorites 0 likes
#benchmark-dataset

Chiaroscuro for Emotions: A Contrastive Emotion Benchmark Grounded in Appraisal Theory

arXiv cs.CL ↗ · 2026-09-04 Cached

The paper presents Chiaro, a new benchmark dataset for contrastive emotion recognition where two individuals experience opposing emotions from a shared event, grounded in appraisal theory. It evaluates seven LLMs and four emotion classifiers, revealing that current models fall short of human performance.

0 favorites 0 likes
#benchmark-dataset

OmniPhys: A Unified Multimodal Benchmark for Physics Understanding and Generation from Chinese Educational Corpora

arXiv cs.CL ↗ · 2026-08-27 Cached

OmniPhys is a large-scale multimodal benchmark for physics understanding and generation, covering middle school to university-level problems from Chinese educational corpora, aimed at evaluating and advancing multimodal large language models in scientific domains.

0 favorites 0 likes
#benchmark-dataset

Padamitra: Grounded Glossary Generation for Classical Sanskrit

arXiv cs.CL ↗ · 2026-08-27 Cached

This paper introduces grounded glossary generation for Classical Sanskrit, a task involving recovering Sanskrit phrases and producing translation-grounded meanings from sloka-translation pairs. It constructs a benchmark from Hindu texts and evaluates various AI models, finding that instruction fine-tuning improves performance, with morphological modeling identified as a key challenge.

0 favorites 0 likes
#benchmark-dataset

Multi-Agent Forensic Reasoning for Generalizable Deepfake Video Detection

Hugging Face Daily Papers ↗ · 2026-08-07 Cached

This paper introduces FaceVid-Forensics-100K, a large-scale deepfake video dataset with fine-grained annotations, and proposes a multi-agent forensic reasoning framework (ARGUS) that uses four specialized expert agents and a judge agent to outperform closed-source models on deepfake detection.

0 favorites 0 likes
#benchmark-dataset

Mixture-of-experts for handwriting trajectory reconstruction from IMU sensors

arXiv cs.LG ↗ · 2026-07-30 Cached

A new mixture-of-experts approach for reconstructing handwriting trajectories from IMU sensor data, with separate experts for touching and hovering phases, and a new public benchmark dataset.

0 favorites 0 likes
#benchmark-dataset

PhononBench-MP40: a spectrum-resolved benchmark dataset for phonon stability

arXiv cs.AI ↗ · 2026-07-28 Cached

PhononBench-MP40 is a benchmark dataset of phonon stability labels and spectra for over 46,000 Materials Project-derived crystals, designed to evaluate workflow-defined phonon stability and support materials screening.

0 favorites 0 likes
#benchmark-dataset

When a Name Is Not a Name: A Benchmark Dataset and Distilled Reasoning for Culturally Entangled Bangla Homographs in Low-Resource LLMs

arXiv cs.CL ↗ · 2026-07-21 Cached

This paper introduces a benchmark dataset of 1,516 expert-verified Bangla sentences for disambiguating culturally entangled homographs (words that are both names and common nouns). It shows that LLMs suffer from dominant-meaning bias and proposes contrastive chain-of-thought prompting and distillation to reduce this bias.

0 favorites 0 likes
#benchmark-dataset

DECODEM: Data Extraction from Corporate Organizational Documents via Enhanced Methods

arXiv cs.CL ↗ · 2026-07-20 Cached

Introduces DECODEM, benchmark datasets for evaluating automated extraction of corporate governance variables from legal documents using large language models, showing high accuracy for many provisions.

0 favorites 0 likes
#benchmark-dataset

@drfeifei: I’m very excited by this new benchmark dataset for visual generation that is suitable for the modern era of large scale…

X AI KOLs Following ↗ · 2026-05-29 Cached

Introducing GPIC (Giant Permissive Image Corpus), a large-scale dataset of 100M VLM-captioned image-text pairs for training and 1M pairs for benchmarking, fully permissive for research and commercial use.

0 favorites 0 likes
#benchmark-dataset

EmoS: A High-Fidelity Multimodal Benchmark for Fine-grained Streaming Emotional Understanding

arXiv cs.CL ↗ · 2026-05-12 Cached

This article introduces EmoS, a high-fidelity multimodal benchmark designed for fine-grained streaming emotional understanding, addressing limitations in ecological validity and labeling reliability found in existing datasets.

0 favorites 0 likes
#benchmark-dataset

SynopticBench: Evaluating Vision-Language Models on Generating Weather Forecast Discussions of the Future

arXiv cs.CL ↗ · 2026-04-21 Cached

This paper introduces SynopticBench, a dataset of 1.3M+ weather forecast discussions paired with meteorological images, and SPACE, a novel evaluation framework for assessing VLM-generated weather forecasts.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback