Tag
MoRA proposes a structured MoE expert pruning framework using learnable router biases optimized with LM loss and a routing-diversity regularizer, plus a post-pruning expert approximation mechanism, outperforming state-of-the-art pruning methods on Qwen3-30B-A3B, DeepSeek-V2-Lite, and Moonlight-16B-A3B with 25-50% expert removal.
Two open-weight fine-tunes of Qwen3.8-Flash-Next are released: Victoria, a pruned coding/agent model (44% of experts removed via REAP, retrained in NVFP4, 70% on Terminal-Bench 2.1) and Maple, a Canada-focused fine-tune that dramatically improves citation of official Canadian sources.
The paper proposes OBC-Prune, a calibration method for pruning large reasoning models that identifies causally important reasoning circuits to improve accuracy and reduce inference overhead on benchmarks like MATH500 and LiveCodeBench.
ToMoE proposes a method to convert dense large language models into Mixture-of-Experts models using dynamic structural pruning without weight updates, outperforming existing techniques.
A trimmed English-only GGUF version of Kimi K3 (IQ2-XXS) reduces model size from 711GB to 478GB by removing multi-language components, with early tests suggesting it may match or outperform the standard 2-bit version on coding tasks.
This paper introduces Maestro, a structured pruning framework for Mixture-of-Experts language models that uses Markov chains to model expert activation trajectories, achieving globally aware pruning and outperforming baselines by up to 10.61% under 50% compression.
This paper proposes entropy-regularized probabilistic gates to maintain uncertainty in sparse federated optimization, improving sparsity recovery and test performance under data heterogeneity and scarce data.
A pruned and quantized version of MiniMax-M3 (MiniMax-M3-Medium-JANG_2L) optimized to run on 128GB Macs using vMLX, featuring 32% expert pruning and JANG_2L mixed-precision quantization to fit within ~105 GB.
This paper demonstrates that cosine similarity is a poor proxy for assessing layer importance in LLMs, and proposes using the actual accuracy drop from layer removal as a more robust metric.
This paper introduces a resource-efficient pruning framework that identifies and removes parameters associated with unsafe behaviors in large language models while preserving utility. Using gradient-free attribution and the Lottery Ticket Hypothesis perspective, the method achieves significant reductions in unsafe generations and improved robustness against jailbreak attacks with minimal performance loss.