Tag
This empirical study tests when chain-of-thought prompting helps or hurts LLM reasoning, finding that CoT provides large gains on deep serial tasks like GSM8K and MATH but is redundant on shallow tasks like MMLU and ARC-Challenge, consistent with a serial-depth bottleneck framework.
A pre-registered study proposes a first-token hallucination detection method using internal model signals across 10 models, finding no universal detector but a universal above-chance floor, with public code and verification scripts.
This paper presents a pre-registered audit of whether LLM-judged helpfulness can reliably distinguish answer-giving from pedagogical guidance in AI tutors. The authors find that general-purpose helpfulness is not a dependable pedagogy signal, recommending pedagogy-targeted rubrics and deterministic process measures instead.
This preregistered study tests whether holonomy (a geometric measure) concentrates on active SAE feature planes in the Gemma 2 2B language model. Contrary to the semantic-concentration prediction, active-feature planes carried less holonomy than matched mixed-feature controls, resulting in a narrow operational reversal with the underlying cause remaining open.
This paper reports a pre-registered experiment on small economies of frontier LLM agents (Claude Opus 4.8), testing predictions about information-theoretic capacity regions for wealth growth and mean-field residual-scaling laws for population misalignment. Results confirm a quantitative information law connecting agent knowledge to earnings but reject the mean-field assumption, revealing discrete attractor dynamics instead.