model-pruning

Tag

Cards List
#model-pruning

MoRA: MoE Pruning via Router Bias Learning and Expert Approximation

arXiv cs.LG ↗ · 2d ago Cached

MoRA proposes a structured MoE expert pruning framework using learnable router biases optimized with LM loss and a routing-diversity regularizer, plus a post-pruning expert approximation mechanism, outperforming state-of-the-art pruning methods on Qwen3-30B-A3B, DeepSeek-V2-Lite, and Moonlight-16B-A3B with 25-50% expert removal.

0 favorites 0 likes
#model-pruning

Two open-weights releases: Victoria (Qwen3.8-Flash-Next with 44% of experts cut, 70% Terminal-Bench 2.1, GGUF included) and Maple (a Canada-first fine-tune)

Reddit r/LocalLLaMA ↗ · 4d ago

Two open-weight fine-tunes of Qwen3.8-Flash-Next are released: Victoria, a pruned coding/agent model (44% of experts removed via REAP, retrained in NVFP4, 70% on Terminal-Bench 2.1) and Maple, a Canada-focused fine-tune that dramatically improves citation of official Canadian sources.

0 favorites 0 likes
#model-pruning

OBC-Prune: Outcome-Based Calibration for Large Reasoning Model Pruning

arXiv cs.AI ↗ · 2026-09-17 Cached

The paper proposes OBC-Prune, a calibration method for pruning large reasoning models that identifies causally important reasoning circuits to improve accuracy and reduce inference overhead on benchmarks like MATH500 and LiveCodeBench.

0 favorites 0 likes
#model-pruning

[Paper] ToMoE: Converting Dense Large Language Models to Mixture-of-Experts through Dynamic Structural Pruning

Reddit r/LocalLLaMA ↗ · 2026-08-24 Cached

ToMoE proposes a method to convert dense large language models into Mixture-of-Experts models using dynamic structural pruning without weight updates, outperforming existing techniques.

0 favorites 0 likes
#model-pruning

Kimi K3 (Unsloth) IQ2-XXS from 711GB down to 478GB!!! Only Multi-language was removed to trim the size

Reddit r/LocalLLaMA ↗ · 2026-08-08

A trimmed English-only GGUF version of Kimi K3 (IQ2-XXS) reduces model size from 711GB to 478GB by removing multi-language components, with early tests suggesting it may match or outperform the standard 2-bit version on coding tasks.

0 favorites 0 likes
#model-pruning

It Takes a MAESTRO To Prune Bad Experts

arXiv cs.CL ↗ · 2026-07-10 Cached

This paper introduces Maestro, a structured pruning framework for Mixture-of-Experts language models that uses Markov chains to model expert activation trajectories, achieving globally aware pruning and outperforming baselines by up to 10.61% under 50% compression.

0 favorites 0 likes
#model-pruning

Entropy-Regularized Probabilistic Gates for Sparse Model Discovery in Scarce-Data Federated Learning

arXiv cs.LG ↗ · 2026-07-02 Cached

This paper proposes entropy-regularized probabilistic gates to maintain uncertainty in sparse federated optimization, improving sparsity recovery and test performance under data heterogeneity and scarce data.

0 favorites 0 likes
#model-pruning

@dealignai: MiniMax m3, made for 128gb Mac’s Thank you to @hornsby_andrew for preparing the pruning calibration dataset and doing e…

X AI KOLs Timeline ↗ · 2026-06-18 Cached

A pruned and quantized version of MiniMax-M3 (MiniMax-M3-Medium-JANG_2L) optimized to run on 128GB Macs using vMLX, featuring 32% expert pruning and JANG_2L mixed-precision quantization to fit within ~105 GB.

0 favorites 0 likes
#model-pruning

Rethinking Layer Relevance in Large Language Models Beyond Cosine Similarity

arXiv cs.LG ↗ · 2026-05-15 Cached

This paper demonstrates that cosine similarity is a poor proxy for assessing layer importance in LLMs, and proposes using the actual accuracy drop from layer removal as a more robust metric.

0 favorites 0 likes
#model-pruning

Pruning Unsafe Tickets: A Resource-Efficient Framework for Safer and More Robust LLMs

arXiv cs.CL ↗ · 2026-04-20 Cached

This paper introduces a resource-efficient pruning framework that identifies and removes parameters associated with unsafe behaviors in large language models while preserving utility. Using gradient-free attribution and the Lottery Ticket Hypothesis perspective, the method achieves significant reductions in unsafe generations and improved robustness against jailbreak attacks with minimal performance loss.

0 favorites 0 likes
← Back to home

Submit Feedback