Tag
The paper introduces Elenchos, a generative evaluation framework for abductive reasoning in LLMs, where models must infer hidden rule changes from behavioral differences under black-box access. It finds a detection-attribution dissociation: models detect alterations but struggle to identify the specific mutations, especially under interacting mutations.
This paper identifies a blind spot in long-context LLM reasoning benchmarks: they fail to control task position within the context, allowing positional failures to go undetected. The authors propose Context Rot Evaluation (CRE) to systematically vary task position, filler content, and context length, revealing severe accuracy drops for some models when reasoning tasks are placed in the middle of long contexts.