recurrent-depth

Tag

Cards List
#recurrent-depth

Retrofitting Recurrent Depth into a Pretrained Language Model: Installation, Extrapolation, Transfer, and Retention at Two Parameter Budgets

arXiv cs.CL · 2026-08-13 Cached

This paper explores surgically retrofitting a pretrained language model (e.g., Qwen2.5-0.5B) with recurrent depth, demonstrating that the resulting model can perform deeper latent reasoning, extrapolate past supervised depth, and outperform dense models fine-tuned to reason in tokens, while also revealing catastrophic interference limits.

0 favorites 0 likes
#recurrent-depth

CHERRY: Compressed Hierarchical Experts with Recurrent Representational Yield

arXiv cs.CL · 2026-07-01 Cached

This paper introduces CHERRY, a set of techniques for compute-efficient language models including selective token supervision, depth compression via recurrent unrolling, and a mixture of compressed experts, achieving significant efficiency gains on a Korean foundation model.

0 favorites 0 likes
← Back to home

Submit Feedback