model-pruning

Tag

Cards List
#model-pruning

Kimi K3 (Unsloth) IQ2-XXS from 711GB down to 478GB!!! Only Multi-language was removed to trim the size

Reddit r/LocalLLaMA · 2026-08-08

A trimmed English-only GGUF version of Kimi K3 (IQ2-XXS) reduces model size from 711GB to 478GB by removing multi-language components, with early tests suggesting it may match or outperform the standard 2-bit version on coding tasks.

0 favorites 0 likes
#model-pruning

It Takes a MAESTRO To Prune Bad Experts

arXiv cs.CL · 2026-07-10 Cached

This paper introduces Maestro, a structured pruning framework for Mixture-of-Experts language models that uses Markov chains to model expert activation trajectories, achieving globally aware pruning and outperforming baselines by up to 10.61% under 50% compression.

0 favorites 0 likes
#model-pruning

Entropy-Regularized Probabilistic Gates for Sparse Model Discovery in Scarce-Data Federated Learning

arXiv cs.LG · 2026-07-02 Cached

This paper proposes entropy-regularized probabilistic gates to maintain uncertainty in sparse federated optimization, improving sparsity recovery and test performance under data heterogeneity and scarce data.

0 favorites 0 likes
#model-pruning

@dealignai: MiniMax m3, made for 128gb Mac’s Thank you to @hornsby_andrew for preparing the pruning calibration dataset and doing e…

X AI KOLs Timeline · 2026-06-18 Cached

A pruned and quantized version of MiniMax-M3 (MiniMax-M3-Medium-JANG_2L) optimized to run on 128GB Macs using vMLX, featuring 32% expert pruning and JANG_2L mixed-precision quantization to fit within ~105 GB.

0 favorites 0 likes
#model-pruning

Rethinking Layer Relevance in Large Language Models Beyond Cosine Similarity

arXiv cs.LG · 2026-05-15 Cached

This paper demonstrates that cosine similarity is a poor proxy for assessing layer importance in LLMs, and proposes using the actual accuracy drop from layer removal as a more robust metric.

0 favorites 0 likes
#model-pruning

Pruning Unsafe Tickets: A Resource-Efficient Framework for Safer and More Robust LLMs

arXiv cs.CL · 2026-04-20 Cached

This paper introduces a resource-efficient pruning framework that identifies and removes parameters associated with unsafe behaviors in large language models while preserving utility. Using gradient-free attribution and the Lottery Ticket Hypothesis perspective, the method achieves significant reductions in unsafe generations and improved robustness against jailbreak attacks with minimal performance loss.

0 favorites 0 likes
← Back to home

Submit Feedback