looped-architecture

Tag

Cards List
#looped-architecture

SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers

Hugging Face Daily Papers · 2d ago Cached

SMELT is a method that loops middle layers in Mixture-of-Experts Transformers to improve training efficiency and downstream performance while matching compute, parameter, and cache budgets, leading to faster loss reduction and practical gains.

0 favorites 0 likes
#looped-architecture

LoopUS: Recasting Pretrained LLMs into Looped Latent Refinement Models

Hugging Face Daily Papers · 2026-05-10 Cached

LoopUS is a post-training framework that converts pretrained LLMs into looped architectures for improved reasoning performance via latent-refinement and adaptive early exiting. It addresses computational costs and capability preservation issues found in existing looped computation methods.

0 favorites 0 likes
← Back to home

Submit Feedback