Tag
Sesame releases TurnBench, a benchmark for evaluating real-time conversational turn-taking, with a leaderboard and labeled dataset to assess model performance in detecting speech events.
This paper introduces MTDiag, a multi-turn diagnostic dialogue dataset for evaluating Large Language Models in realistic clinical diagnostic scenarios, addressing limitations of static QA benchmarks.
This paper introduces Strategic 16K, a leakage-controlled dataset for document sensitivity classification, and benchmarks classical and transformer-based models, with BERT achieving top performance while addressing label leakage issues.
SupraLabs releases reasoning-corpus-4K-5M-v1, a 5M-sample reasoning dataset for training small language models (SLMs), featuring chain-of-thought traces and ChatML format, hosted on Hugging Face.
Researchers introduce a new multimodal benchmark derived from Japan's National Assessment of Academic Ability, featuring 900K aggregated student responses to evaluate MLLM performance in authentic K-12 educational contexts.
This paper introduces Checkup2Action, a multimodal dataset and benchmark for generating patient-oriented action cards from clinical check-up reports, addressing the interpretability gap for laypersons.
PianoCoRe is a large-scale piano MIDI dataset unifying and refining open-source corpora with 250,046 performances of 5,625 pieces by 483 composers, featuring note-level alignments for music information retrieval and including a MIDI quality classifier and alignment refinement pipeline.