@kyutai_labs: KairosQA contains over 7000 temporally grounded questions created using Wikidata. We look for questions whose answers c…
Summary
Kyutai Labs introduces KairosQA, a dataset of over 7,000 temporally grounded questions from Wikidata, focusing on popular Wikipedia pages where answers changed multiple times between 2018 and 2025.
View Cached Full Text
Cached at: 05/27/26, 07:03 AM
KairosQA contains over 7000 temporally grounded questions created using Wikidata. We look for questions whose answers changed at least twice between 2018 and 2025, and filter to only popular Wikipedia pages. https://t.co/kFiJKM3MRT
Similar Articles
@kyutai_labs: We show that LLM accuracy drops off on time-sensitive questions like “Who won the Champions League in year X” - even wi…
Kyutai Labs demonstrates that LLM accuracy degrades on time-sensitive questions within the knowledge cutoff and introduces the KairosQA dataset. They train models that avoid this issue by temporally sorting the training data.
Introducing SimpleQA
OpenAI introduces SimpleQA, a new factuality benchmark dataset with 4,326 short fact-seeking questions designed to evaluate frontier language models on their ability to provide accurate answers without hallucination. The dataset achieves high quality through dual independent annotation, rigorous criteria, and achieves only ~3% estimated error rate, with GPT-4o scoring less than 40%.
ResearchQA: Benchmarking Citation-Grounded Question-Answering on Scientific Papers
ResearchQA is a new benchmark of 6,211 single-paper question-answer pairs from 494 open-access papers across eight domains, designed to evaluate citation-grounded question-answering by requiring verifiable citations and supporting grounded refusal when evidence is insufficient.
LakeQuest: A Three-Domain Benchmark for Grounded Question Answering across Data Lakes
Introduces LakeQuest, a human-validated benchmark of 9,846 QA pairs across three domains for evaluating end-to-end retrieve-and-synthesize pipelines over heterogeneous data lakes, revealing critical failure modes in modern QA systems.
LakeQA: An Exploratory QA Benchmark over a Million-Scale Data Lake
LakeQA is a new benchmark for exploratory question answering over a million-scale data lake, evaluating multi-hop reasoning and compositionality across text, tables, and knowledge graphs.