cross-layer-routing

Tag

Cards List
#cross-layer-routing

Linear Attention Architectures: Mechanisms, Trade-offs, and Cross-Layer Routing

arXiv cs.LG · 2026-07-10 Cached

This paper compares softmax attention with four linear attention architectures (DeltaNet, Gated DeltaNet, Kimi Delta Attention, Gated DeltaNet-2) and introduces cross-layer routing mechanisms. Experiments at 350M parameters show Kimi Delta Attention with Muon achieves lowest validation loss, while pure Gated DeltaNet with AdamW has highest throughput.

0 favorites 0 likes
#cross-layer-routing

Rethinking Cross-Layer Information Routing in Diffusion Transformers

Hugging Face Daily Papers · 2026-05-20 Cached

This paper proposes Diffusion-Adaptive Routing (DAR), a learnable, timestep-adaptive residual replacement that improves cross-layer information flow in Diffusion Transformers, leading to significant training acceleration and quality improvements.

0 favorites 0 likes
← Back to home

Submit Feedback