sparse-activation

Tag

Cards List
#sparse-activation

Depth-Aware Sensitivity Analysis of Mixture-of-Experts Models via Magnitude-Based Expert Masking

arXiv cs.AI · 2026-08-17 Cached

This paper presents a systematic sensitivity analysis of Mixture-of-Experts models using magnitude-based expert masking, finding that late layers are more resilient to masking, which provides a practical path for model compression.

0 favorites 0 likes
#sparse-activation

@maximelabonne: Wow, this gives me flashbacks of early model merging. Complete insanity, I love it!

X AI KOLs Following · 2026-08-06 Cached

A new sub-6B sparse activation AI model built with a fusion architecture combines weights from LFM2.5-2.6B and Qwen3.6-35B-A3B, achieving near-Qwen3.6-35B performance at a fraction of the size.

0 favorites 0 likes
#sparse-activation

Generic Expert Coverage for Pruning SparseMixture-of-Experts Language Models

arXiv cs.AI · 2026-07-03 Cached

Proposes Generic TB-Coverage, a coverage-aware expert pruning method for sparse Mixture-of-Experts language models that uses only generic text corpora for calibration and preserves cross-corpus expert coverage, improving accuracy and reducing perplexity degradation.

0 favorites 0 likes
← Back to home

Submit Feedback