Tag
This article discusses the limitations of AI models in maintaining context over long conversations, highlighting recency bias and the distinction between context window size and actual comprehension. It suggests practical workarounds like restating constraints and using running context documents.
This paper studies temporal failure modes in LLM-based statutory question answering, including post-cutoff staleness and recency bias. It introduces a benchmark of 312 expert-validated German statutory QA pairs and evaluates LLMs under various inference settings.
This article highlights a common problem in local LLMs where they incorrectly classify real-time information beyond their knowledge cutoff as fictional or satirical, even when provided with tools, often due to excessive RLHF training.