training-data-granularity

Tag

Cards List
#training-data-granularity

Decodable But Not Detachable: Training Data Granularity Determines Parametric Modularity in Large Language Models

arXiv cs.AI · 2026-08-12 Cached

This paper investigates when large language models develop domain-specific parametric shells (causally necessary neuron populations), finding that modular training data at the token level (e.g., languages, code) produces functional shells, while academic subject domains do not, despite being linearly decodable.

0 favorites 0 likes
← Back to home

Submit Feedback