Tag
This paper demonstrates that training a world model through random physical exploration leads to latent representations that encode spatial semantic structure (direction and position) without any linguistic supervision, highlighting physical geometry as the organizing principle.
This paper demonstrates that transformers trained on Sudoku solving traces build structured world models organized by domain constraints, and identifies a sparse, monosemantic circuit responsible for the naked-single decision rule. The work provides a fully interpretable algorithmic account of transformer reasoning on a combinatorial task.