Tag
Ai2 introduces TutorMoments, a replay-based evaluation framework and dataset for measuring whether LLMs can balance when to help and when to hold back in one-on-one math tutoring. Preliminary results show models tend to over-help, and prompt engineering only partially closes the gap to human tutors.
This paper introduces Contrastive Anchor Probing (CAP) to study and detect preference-induced stance reversal sycophancy (PSRS) in LLMs, analyzing 290,460 labeled responses across 17 models and showing detection is possible from response text alone.
Daniel van Strien uploaded a dataset of 1,080,814 public domain images from 49,455 digitised books (c.1510–1900) from the British Library to Hugging Face Hub, organised into four configs by image type.
IslamicTurathBench (ISTB) is a new multi-task, multi-discipline benchmark for evaluating large language models on classical Islamic scholarship, containing 3,465 expert-reviewed questions across 35 works and seven fields.
This paper presents MMLongBench-Doc-V2, a corrected and semantics-aware revision of the MMLongBench-Doc long-document QA benchmark, fixing annotation errors and replacing string matching with an LLM judge, along with a decision procedure for empty-set keys.
The paper presents a methodology for building a large-scale Russian dataset from social media texts to detect presuicidal and anti-suicidal signals, including annotation guidelines and baseline classification experiments.
Introduces LegalPincite, a large-scale legal information retrieval dataset built from CJEU judgments, featuring masked queries, full corpora, and paragraph-level citation annotations to enable multi-level retrieval evaluation.
Introduces ARB, a matched authorship-rewriting benchmark for evaluating AI-text detectors, showing that detector performance drops significantly when human text is rewritten by an LLM despite high recall on direct LLM-generated text.
Ego2Robot is a scalable pipeline that converts egocentric human manipulation videos into robot training data via action retargeting and visual synthesis, producing 18,561 hours of data across 15 robot morphologies. Experiments show that joint pretraining on this synthesized data improves out-of-distribution generalization for vision-language-action models, including on real-robot deployment.
Introduces GEOID-Flood, a large-scale multi-modal benchmark dataset for flood segmentation with over 14,000 tiles from 219 events across 65 countries, evaluating foundation models against conventional encoders across single-image, multi-temporal, and multi-modal protocols.
Introduces DENSEWORLD, a 1,000-hour dataset of crowded Global South urban scenes, and FactorJEPA, a JEPA variant that factorizes future prediction into layout, agents, and interactions, improving accuracy and robustness under occlusion and heterogeneity.
Introduces a new problem domain for fine-grained analysis of children's gait behaviors from standard RGB video, along with a new dataset of over 1,100 high-frame-rate sequences and a unified framework, demonstrating that current SOTA methods and MLLMs fail on this clinical task.
This paper introduces ICLE++, a new corpus of persuasive student essays annotated with both holistic and trait-specific scores, aimed at improving generalization and multi-trait scoring in automated essay scoring (AES) research.
This paper introduces AHA-Memes, the first large-scale Arabic hateful meme benchmark with fine-grained multi-label annotations, covering 5K manually annotated and ~66K silver-labeled memes, and benchmarks various multimodal models for culturally grounded hate detection.
CG-World is a large-scale world-state dataset and protocol derived from industrial computer graphics pipelines, explicitly recording multimodal world states, interventions, and counterfactual branches to support world model research. It demonstrates improvements in geometry-conditioned video generation, action prediction, and closed-loop transfer of vision-language-action policies.
The author proposes a pipeline for generating diverse synthetic reasoning training data for LLMs using formal solvers, and asks for existing work and advice on avoiding repetitive templates.
This Data Descriptor presents a large-scale corpus of transcribed religious radio broadcasts captured from live webstreams over one month in July 2025, comprising over 700,000 recordings and 60 million transcript lines, annotated using LLMs for program format and topic. It enables descriptive study of religious broadcasting and analysis of social/political issues in religious media.
Introduces ACE-Data-0, a large-scale embodied AI dataset with 150 hours of synchronized multimodal human demonstrations across 200 task categories, captured by the Ambient Capture Engine (ACE) in real home environments. Includes a hierarchical benchmark exposing gaps in current methods under contact, occlusion, and long horizons.
Indic DiarBench is a multilingual joint diarization and ASR benchmark covering all 22 scheduled languages of India with 108 hours of human-corrected multi-speaker audio, capturing conversational nuances like code-mixing and speaker overlap.
This paper introduces GEMCo, the first publicly available German corpus of multi-turn e-mail counselling conversations, validated against real counselling data to serve as an ethically releasable proxy for inaccessible sensitive data.