pre-registered

Tag

Cards List
#pre-registered

When Chain-of-Thought Helps and When It Hurts: An Empirical Investigation of the Serial-Depth Bottleneck in LLM Reasoning

arXiv cs.CL · 2d ago Cached

This empirical study tests when chain-of-thought prompting helps or hurts LLM reasoning, finding that CoT provides large gains on deep serial tasks like GSM8K and MATH but is redundant on shallow tasks like MMLU and ARC-Challenge, consistent with a serial-depth bottleneck framework.

0 favorites 0 likes
#pre-registered

No universal hallucination detector, but a universal floor — pre-registered, 10 models. Come break it. [R]

Reddit r/MachineLearning · 2026-08-03

A pre-registered study proposes a first-token hallucination detection method using internal model signals across 10 models, finding no universal detector but a universal above-chance floor, with public code and verification scripts.

0 favorites 0 likes
#pre-registered

Rethinking LLM-Judged Helpfulness as a Pedagogy Signal: A Pre-Registered Audit Across Tutor Models

arXiv cs.CL · 2026-07-31 Cached

This paper presents a pre-registered audit of whether LLM-judged helpfulness can reliably distinguish answer-giving from pedagogical guidance in AI tutors. The authors find that general-purpose helpfulness is not a dependable pedagogy signal, recommending pedagogy-targeted rubrics and deterministic process measures instead.

0 favorites 0 likes
#pre-registered

Do Active SAE Feature Planes Carry More Holonomy? A Preregistered Reversal in Gemma

arXiv cs.LG · 2026-07-24 Cached

This preregistered study tests whether holonomy (a geometric measure) concentrates on active SAE feature planes in the Gemma 2 2B language model. Contrary to the semantic-concentration prediction, active-feature planes carried less holonomy than matched mixed-feature controls, resulting in a narrow operational reversal with the underlying cause remaining open.

0 favorites 0 likes
#pre-registered

Information Limits and Attractor Dynamics in Economies of Frontier LLM Agents: A Pre-Registered Test

arXiv cs.AI · 2026-07-08 Cached

This paper reports a pre-registered experiment on small economies of frontier LLM agents (Claude Opus 4.8), testing predictions about information-theoretic capacity regions for wealth growth and mean-field residual-scaling laws for population misalignment. Results confirm a quantitative information law connecting agent knowledge to earnings but reject the mean-field assumption, revealing discrete attractor dynamics instead.

0 favorites 0 likes
← Back to home

Submit Feedback