Tag
This paper proposes a method to close the MLP reachability gap in Low-Rank Clone distillation by training the full deployed matrix, resulting in significant improvements in token efficiency and model performance at no additional inference cost.