dataset

Tag

Cards List
#dataset

TutorMoments: Do AI tutors know when to help and when to hold back?

Hugging Face Blog · 2d ago Cached

Ai2 introduces TutorMoments, a replay-based evaluation framework and dataset for measuring whether LLMs can balance when to help and when to hold back in one-on-one math tutoring. Preliminary results show models tend to over-help, and prompt engineering only partially closes the gap to human tutors.

0 favorites 0 likes
#dataset

Measuring and Detecting Harmful AI Sycophancy

arXiv cs.AI · 3d ago Cached

This paper introduces Contrastive Anchor Probing (CAP) to study and detect preference-induced stance reversal sycophancy (PSRS) in LLMs, analyzing 290,460 labeled responses across 17 models and showing detection is possible from response text alone.

0 favorites 0 likes
#dataset

@vanstriendaniel: Uploaded a dataset of 1,080,814 public domain images, mostly from 19th-century books, to the Hub. https://huggingface.c…

X AI KOLs Following · 3d ago Cached

Daniel van Strien uploaded a dataset of 1,080,814 public domain images from 49,455 digitised books (c.1510–1900) from the British Library to Hugging Face Hub, organised into four configs by image type.

0 favorites 0 likes
#dataset

IslamicTurathBench: A Multi-Task, Multi-Discipline Benchmark for Evaluating Large Language Models on the Islamic Scholarly Tradition (turath)

arXiv cs.CL · 4d ago Cached

IslamicTurathBench (ISTB) is a new multi-task, multi-discipline benchmark for evaluating large language models on classical Islamic scholarship, containing 3,465 expert-reviewed questions across 35 works and seven fields.

0 favorites 0 likes
#dataset

MMLongBench-Doc-V2: A Corrected-Annotation, Semantics-Aware Revision of MMLongBench-Doc

arXiv cs.AI · 5d ago Cached

This paper presents MMLongBench-Doc-V2, a corrected and semantics-aware revision of the MMLongBench-Doc long-document QA benchmark, fixing annotation errors and replacing string matching with an LLM judge, along with a decision procedure for empty-set keys.

0 favorites 0 likes
#dataset

The methodology of Constructing the Large-Scale Dataset for Detecting Presuicidal and Anti-Suicidal Signals in Social Media Texts in Russian

arXiv cs.CL · 6d ago Cached

The paper presents a methodology for building a large-scale Russian dataset from social media texts to detect presuicidal and anti-suicidal signals, including annotation guidelines and baseline classification experiments.

0 favorites 0 likes
#dataset

LegalPincite: Multi-level Legal Information Retrieval Dataset

Hugging Face Daily Papers · 6d ago Cached

Introduces LegalPincite, a large-scale legal information retrieval dataset built from CJEU judgments, featuring masked queries, full corpora, and paragraph-level citation annotations to enable multi-level retrieval evaluation.

0 favorites 0 likes
#dataset

ARB: A Matched Authorship-Rewriting Benchmark Dataset for AI-Text Detector Evaluation

arXiv cs.CL · 2026-08-03 Cached

Introduces ARB, a matched authorship-rewriting benchmark for evaluating AI-text detectors, showing that detector performance drops significantly when human text is rewritten by an LLM despite high recall on direct LLM-generated text.

0 favorites 0 likes
#dataset

Ego2Robot: Scalable Robot Data Synthesis from Egocentric Human Data

Hugging Face Daily Papers · 2026-08-03 Cached

Ego2Robot is a scalable pipeline that converts egocentric human manipulation videos into robot training data via action retargeting and visual synthesis, producing 18,561 hours of data across 15 robot morphologies. Experiments show that joint pretraining on this synthesized data improves out-of-distribution generalization for vision-language-action models, including on real-robot deployment.

0 favorites 0 likes
#dataset

GEOID-Flood: A Large-Scale Multi-Modal Benchmark Dataset for Flood Segmentation

Hugging Face Daily Papers · 2026-08-03 Cached

Introduces GEOID-Flood, a large-scale multi-modal benchmark dataset for flood segmentation with over 14,000 tiles from 219 events across 65 countries, evaluating foundation models against conventional encoders across single-image, multi-temporal, and multi-modal protocols.

0 favorites 0 likes
#dataset

FactorJEPA: Factorizing Monolithic Futures into Layout-Agent-Interaction Channels for Crowded and Chaotic Global South Urban Worlds

Hugging Face Daily Papers · 2026-08-02 Cached

Introduces DENSEWORLD, a 1,000-hour dataset of crowded Global South urban scenes, and FactorJEPA, a JEPA variant that factorizes future prediction into layout, agents, and interactions, improving accuracy and robustness under occlusion and heterogeneity.

0 favorites 0 likes
#dataset

Decoding Children's Gait Behavior

Hugging Face Daily Papers · 2026-08-01 Cached

Introduces a new problem domain for fine-grained analysis of children's gait behaviors from standard RGB video, along with a new dataset of over 1,100 high-frame-rate sequences and a unified framework, demonstrating that current SOTA methods and MLLMs fail on this clinical task.

0 favorites 0 likes
#dataset

ICLE++: Modeling Fine-Grained Traits for Holistic Essay Scoring

arXiv cs.CL · 2026-07-31 Cached

This paper introduces ICLE++, a new corpus of persuasive student essays annotated with both holistic and trait-specific scores, aimed at improving generalization and multi-trait scoring in automated essay scoring (AES) research.

0 favorites 0 likes
#dataset

AHA-Memes: A Fine-Grained Multimodal Benchmark for Understanding Hate in Arabic Memes

arXiv cs.CL · 2026-07-31 Cached

This paper introduces AHA-Memes, the first large-scale Arabic hateful meme benchmark with fine-grained multi-label annotations, covering 5K manually annotated and ~66K silver-labeled memes, and benchmarks various multimodal models for culturally grounded hate detection.

0 favorites 0 likes
#dataset

CG-World: A Large-Scale World-State Dataset and Protocol for World Models

arXiv cs.AI · 2026-07-31 Cached

CG-World is a large-scale world-state dataset and protocol derived from industrial computer graphics pipelines, explicitly recording multimodal world states, interventions, and counterfactual branches to support world model research. It demonstrates improvements in geometry-conditioned video generation, action prediction, and closed-loop transfer of vision-language-action policies.

0 favorites 0 likes
#dataset

Making a synthetic dataset for fine-tuning

Reddit r/LocalLLaMA · 2026-07-30

The author proposes a pipeline for generating diverse synthetic reasoning training data for LLMs using formal solvers, and asks for existing work and advice on avoiding repetitive templates.

0 favorites 0 likes
#dataset

A large-scale corpus of religious radio broadcast transcripts from webstream recordings in the United States

arXiv cs.CL · 2026-07-30 Cached

This Data Descriptor presents a large-scale corpus of transcribed religious radio broadcasts captured from live webstreams over one month in July 2025, comprising over 700,000 recordings and 60 million transcript lines, annotated using LLMs for program format and topic. It enables descriptive study of religious broadcasting and analysis of social/political issues in religious media.

0 favorites 0 likes
#dataset

ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine

Hugging Face Daily Papers · 2026-07-30 Cached

Introduces ACE-Data-0, a large-scale embodied AI dataset with 150 hours of synchronized multimodal human demonstrations across 200 task categories, captured by the Ambient Capture Engine (ACE) in real home environments. Includes a hierarchical benchmark exposing gaps in current methods under contact, occlusion, and long horizons.

0 favorites 0 likes
#dataset

Indic DiarBench: A Multilingual Joint Diarization and ASR Benchmark for Indian Languages

arXiv cs.CL · 2026-07-28 Cached

Indic DiarBench is a multilingual joint diarization and ASR benchmark covering all 22 scheduled languages of India with 108 hours of human-corrected multi-speaker audio, capturing conversational nuances like code-mixing and speaker overlap.

0 favorites 0 likes
#dataset

GEMCo: A Validated, Ethically Releasable Proxy for Inaccessible Counselling Data

arXiv cs.CL · 2026-07-28 Cached

This paper introduces GEMCo, the first publicly available German corpus of multi-turn e-mail counselling conversations, validated against real counselling data to serve as an ethically releasable proxy for inaccessible sensitive data.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback