Tag
The paper proposes a two-stage mechanism for subliminal trait transfer in AI models, where optimizer state transports source perturbations and later training determines their behavioral value.
This paper empirically shows that the gradient's top-r subspace in low-rank training methods like GaLore is non-identifiable beyond a small reproducible core, with estimator noise dominating apparent rotations. It analyzes the implications for optimizer state transport and introduces LDAdam, which outperforms GaLore in perplexity.