token-skipping

Tag

Cards List
#token-skipping

Beyond Single-Dimensional Compression: The Compound Sparsity Frontier of Large Language Models

arXiv cs.LG · 2d ago Cached

This paper introduces a compound sparsity framework for LLMs that combines static parameter pruning with dynamic token-level computation, showing that mixing both mechanisms outperforms single-dimension compression and delays performance degradation.

0 favorites 0 likes
#token-skipping

@FinanceYF5: MoE models may waste about half of expert computations on tokens that don't need experts 1/ Half of experts are working for nothing MoE models already seem efficient, but a paper finds that many tokens don't need expert processing at all. ZEDA teaches the model to "save when possible," skipping up to 50% of expert computations.

X AI KOLs Following · 2026-05-25 Cached

A paper discovers that about 50% of expert computations in MoE models are wasted on tokens that don't need expert processing. The proposed ZEDA method teaches the model to skip these computations, saving up to half of expert calculations.

0 favorites 0 likes
← Back to home

Submit Feedback