Tag
This paper explores surgically retrofitting a pretrained language model (e.g., Qwen2.5-0.5B) with recurrent depth, demonstrating that the resulting model can perform deeper latent reasoning, extrapolate past supervised depth, and outperform dense models fine-tuned to reason in tokens, while also revealing catastrophic interference limits.
This paper introduces CHERRY, a set of techniques for compute-efficient language models including selective token supervision, depth compression via recurrent unrolling, and a mixture of compressed experts, achieving significant efficiency gains on a Korean foundation model.