token-adaptive

Tag

Cards List
#token-adaptive

Share First, Route What Remains: A Unified Framework for Token-Adaptive MoE Computation

arXiv cs.LG · 2026-08-12 Cached

This paper proposes UniF-MoE, a unified framework for token-adaptive Mixture-of-Experts computation that first shares reusable computation across experts and then routes the remaining residual demand, improving performance while reducing activated computation, latency, and memory on DomainBed and GLUE benchmarks.

0 favorites 0 likes
#token-adaptive

DepthWeave-KV: Token-Adaptive Cross-Layer Residual Factorization for Long-Context KV Cache Compression

arXiv cs.AI · 2026-07-08 Cached

DepthWeave-KV is a token-adaptive cross-layer residual factorization method for compressing KV cache in long-context transformer inference, achieving 8.3x memory reduction and 72.8 tokens/s at 64K context while preserving near-full-cache task quality across benchmarks.

0 favorites 0 likes
← Back to home

Submit Feedback