reasoning-benchmark

Tag

Cards List
#reasoning-benchmark

LLMs Can See the Smoke but not the Fire: Evaluating Abductive Reasoning with Elenchos

arXiv cs.AI · 2026-07-15 Cached

The paper introduces Elenchos, a generative evaluation framework for abductive reasoning in LLMs, where models must infer hidden rule changes from behavioral differences under black-box access. It finds a detection-attribution dissociation: models detect alterations but struggle to identify the specific mutations, especially under interacting mutations.

0 favorites 0 likes
#reasoning-benchmark

Positional Failures in Long-Context LLMs: A Blind Spot in Reasoning Benchmarks

arXiv cs.CL · 2026-05-25 Cached

This paper identifies a blind spot in long-context LLM reasoning benchmarks: they fail to control task position within the context, allowing positional failures to go undetected. The authors propose Context Rot Evaluation (CRE) to systematically vary task position, filler content, and context length, revealing severe accuracy drops for some models when reasoning tasks are placed in the middle of long contexts.

0 favorites 0 likes
← Back to home

Submit Feedback