causal-ablation

Tag

Cards List
#causal-ablation

Dissecting Hierarchical Reasoning Models: A Mechanistic Study

arXiv cs.LG ↗ · 2026-09-22 Cached

This paper provides a mechanistic analysis of Hierarchical Reasoning Models (HRM) to understand their internal reasoning processes in latent space, using techniques like causal interventions and sparse autoencoders on tasks such as Sudoku and ARC-AGI-2.

0 favorites 0 likes
#causal-ablation

Pattern Selectivity is Not Task-Causal Structure: A Cross-Architecture Mechanistic Study of Composed-Task Circuits in 1B-Class Language Models

arXiv cs.LG ↗ · 2026-06-05 Cached

This paper tests whether the standard recipe for identifying attention-head circuits by task-pattern selectivity and causal ablation yields consistent mechanistic claims across different 1B-class language model families (Pythia, OLMo, OLMoE). It finds no two (task, model) cells share the same primary causal screen, and introduces a five-category taxonomy of screen outcomes, with the MoE model showing a distinct prev-token positional substrate.

0 favorites 0 likes
#causal-ablation

Convergence Without Understanding: When Language Models Agree on Representations but Disagree on Reasoning

arXiv cs.CL ↗ · 2026-05-25 Cached

This paper investigates the Platonic Representation Hypothesis by examining 16 language models across 8 families on 800 reasoning problems. It finds that while models converge in internal representations, they diverge in reasoning processes, especially post-decision, and shared representations have minimal causal influence on predictions.

0 favorites 0 likes
← Back to home

Submit Feedback