Tag
This preprint reveals that under certain data symmetries, LLM training dynamics can be reduced to a low-dimensional subspace, making analysis more tractable and interpretable, with each coordinate corresponding to a clear mechanism.
This paper studies stochastic linear bandits where the agent only observes a random subset of action coordinates, proving that sublinear regret is possible when actions have low intrinsic dimension, and proposes the TOFU-POV algorithm with theoretical guarantees.