centered-kernel-alignment

Tag

Cards List
#centered-kernel-alignment

Convergence Without Understanding: When Language Models Agree on Representations but Disagree on Reasoning

arXiv cs.CL · 2026-05-25 Cached

This paper investigates the Platonic Representation Hypothesis by examining 16 language models across 8 families on 800 reasoning problems. It finds that while models converge in internal representations, they diverge in reasoning processes, especially post-decision, and shared representations have minimal causal influence on predictions.

0 favorites 0 likes
← Back to home

Submit Feedback