shared-experts

Tag

Cards List
#shared-experts

Share First, Route What Remains: A Unified Framework for Token-Adaptive MoE Computation

arXiv cs.LG ↗ · 2026-08-12 Cached

This paper proposes UniF-MoE, a unified framework for token-adaptive Mixture-of-Experts computation that first shares reusable computation across experts and then routes the remaining residual demand, improving performance while reducing activated computation, latency, and memory on DomainBed and GLUE benchmarks.

0 favorites 0 likes
← Back to home

Submit Feedback