Tag
This paper presents a systematic sensitivity analysis of Mixture-of-Experts models using magnitude-based expert masking, finding that late layers are more resilient to masking, which provides a practical path for model compression.
A new sub-6B sparse activation AI model built with a fusion architecture combines weights from LFM2.5-2.6B and Qwen3.6-35B-A3B, achieving near-Qwen3.6-35B performance at a fraction of the size.
Proposes Generic TB-Coverage, a coverage-aware expert pruning method for sparse Mixture-of-Experts language models that uses only generic text corpora for calibration and preserves cross-corpus expert coverage, improving accuracy and reducing perplexity degradation.