Tag
This paper studies what transfers between transformer models of different sizes in the same family (Pythia), showing that representations align while weights don't, and that conversion works best via initialization rather than direct weight projection.