model-compression

Tag

Cards List
#model-compression

[Release] GSQ-RCO GGUFs for Qwen3.8-Flash-Next, plus a 50% expert-pruned Coder build at ~1.89 bpw

Reddit r/LocalLLaMA ↗ · 6h ago

Released GSQ-RCO quantized and expert-pruned versions of Qwen3.8-Flash-Next, achieving BF16-level performance at reduced bit-widths and enabling deployment on smaller hardware.

0 favorites 0 likes
#model-compression

What do you think about ToMoE v2 paper, converting dense model to MoE model at near lossless accuracy? I feel Qwen3.8-27B-A16B or something along those lines would be amazing, though there are architectural hurdles, as well as need folr training data.

Reddit r/LocalLLaMA ↗ · 20h ago Cached

ToMoE is a method that converts dense large language models into mixture-of-experts models through dynamic structural pruning, achieving strong performance without weight updates and outperforming state-of-the-art pruning and MoE techniques.

0 favorites 0 likes
#model-compression

Understanding the Role of Prompt Template in Knowledge Distillation for Safety Alignment

arXiv cs.CL ↗ · yesterday Cached

This paper analyzes how prompt template selection during knowledge distillation affects the safety alignment of student large language models, finding that chat templates lead to greater degradation compared to non-chat templates across multiple models and benchmarks.

0 favorites 0 likes
#model-compression

Fewer Tokens, More Self-Teaching: On-Policy Self-Distillation for Extreme Visual Token Reduction

Hugging Face Daily Papers ↗ · 3d ago Cached

The paper proposes LT-OPD, an on-policy self-distillation framework for extreme visual token reduction in multimodal large language models, significantly improving performance under low token budgets while reducing KV-cache usage and FLOPs.

0 favorites 0 likes
#model-compression

Temporal Taxation Compounds Under Post-Training Compression of Whisper Models

arXiv cs.CL ↗ · 4d ago Cached

This paper investigates how post-training weight compression of Whisper ASR models widens demographic disparities in transcription errors, leading to increased correction time for marginalized speakers.

0 favorites 0 likes
#model-compression

Quantization-Robust Unlearning through the Lens of Retain-Forget Loss Landscapes Interaction

arXiv cs.LG ↗ · 5d ago Cached

This paper proposes a quantization-robust unlearning framework for large language models, using loss landscape analysis to ensure effective forgetting while maintaining model utility after compression.

0 favorites 0 likes
#model-compression

Magnitude Profile Pruning: Calibration-Free Structured Attention Head Removal for Transformer Compression

arXiv cs.CL ↗ · 6d ago Cached

The paper introduces Magnitude Profile (MP) scoring, a calibration-free method for pruning attention heads in transformers, achieving better perplexity on models like OPT-6.7B and RoBERTa-large compared to existing methods, with zero forward passes or calibration data.

0 favorites 0 likes
#model-compression

Six Layers Less: Encoder Pruning for Whisper with Label-Free Recovery

Hugging Face Daily Papers ↗ · 6d ago Cached

The paper introduces a method to prune encoder layers in the Whisper ASR model, reducing encoder size by 18.5% and recovering performance through unlabeled data distillation, with code and a pre-trained model released for adoption.

0 favorites 0 likes
#model-compression

efficient fine tune storage

Reddit r/LocalLLaMA ↗ · 2026-09-22

DeltaTensors compresses fine-tuned model checkpoints by storing weight differences from the base model, reducing storage requirements significantly while allowing reconstruction with minimal quality loss.

0 favorites 0 likes
#model-compression

Correlation-Aware Structured Pruning for Large Language Models

arXiv cs.CL ↗ · 2026-09-22 Cached

The paper proposes a correlation-aware structured pruning method for large language models, explicitly modeling cross-unit dependencies to improve pruning decisions and achieve competitive accuracy-efficiency trade-offs.

0 favorites 0 likes
#model-compression

GeoPair: Geometry-Preserving Cross-Layer Factorization for Training-Free Transformer Compression

Hugging Face Daily Papers ↗ · 2026-09-22 Cached

The paper introduces GeoPair, a training-free framework for transformer compression that optimizes cross-layer factorizations while preserving activation geometries, achieving state-of-the-art results across diverse architectures.

0 favorites 0 likes
#model-compression

@Kay2289123: I highly recommend that everyone bookmark this Stanford course from this fall: MS&E 319: Efficient Generative Language …

X AI KOLs Timeline ↗ · 2026-09-21 Cached

This tweet recommends a Stanford course on efficient generative language models, covering techniques from pre-training to inference to balance performance and cost with limited compute.

0 favorites 0 likes
#model-compression

Pruning LLMs Like a Physicist: Block Removal as an Ising Optimization Problem

Hugging Face Blog ↗ · 2026-09-21 Cached

The paper presents a physics-inspired approach to pruning LLM blocks by modeling block removal as a constrained binary optimization problem mapped to an Ising glass, achieving significant compression gains without benchmarking each configuration.

0 favorites 0 likes
#model-compression

Spanish Multiverse Computing used a quantum computer (156-qubit IBM Heron) to create their latest Quasar 1.1 438B model

Reddit r/singularity ↗ · 2026-09-19 Cached

Multiverse Computing released Quasar 1.1 438B, an AI coding model built using quantum-generated data, with improved efficiency and reduced political censorship.

0 favorites 0 likes
#model-compression

@HuggingModels: Meet Ternary-Bonsai-2-27B: a 27B parameter model squeezed into 2-bit ternary format. Runs on llama.cpp with CUDA and Me…

X AI KOLs Timeline ↗ · 2026-09-18 Cached

Ternary-Bonsai-2-27B is a 27B parameter AI model compressed into 2-bit ternary format, running on llama.cpp with CUDA and Metal support, facilitating lightweight on-device AI deployment.

0 favorites 0 likes
#model-compression

@RemiCadene: Who is distilling Astra for robotics?

X AI KOLs Timeline ↗ · 2026-09-18

A tweet posing a question about who is distilling the Astra model for robotics, indicating community interest in AI model optimization for deployment in robotic systems.

0 favorites 0 likes
#model-compression

Layer-wise Curriculum Learning for Efficient LLM Compression

arXiv cs.LG ↗ · 2026-09-18 Cached

The paper introduces a layer-wise curriculum learning method for efficient LLM compression, achieving state-of-the-art performance with significant reductions in GPU memory usage and training time.

0 favorites 0 likes
#model-compression

@BenjaminDEKR: I'm testing this on a 5070 Ti (16gb vram) now and it's really impressive so far

X AI KOLs Following ↗ · 2026-09-18 Cached

A user is testing the newly announced Ternary Bonsai 2 27B AI model, a smaller and quantized version of Qwen3.8 27B, on an NVIDIA 5070 Ti GPU and finds its performance impressive.

0 favorites 0 likes
#model-compression

Ternary Bonsai 2 (27B) just released on Hugging Face. At <6GB in size, it can even run locally in-browser on WebGPU.

Reddit r/LocalLLaMA ↗ · 2026-09-17

Ternary Bonsai 2 is a 27B parameter model derived from Qwen3.8-27B that uses ternary weights to achieve a size under 6GB while retaining 98.2% of its intelligence, enabling it to run in-browser on WebGPU.

0 favorites 0 likes
#model-compression

The Other Half of the Memory Wall: Serving 35B MoEs from SSD with Trained Routing Prediction

arXiv cs.AI ↗ · 2026-09-17 Cached

Edge0 is a streaming MoE inference engine that uses trained routing prediction to serve 35B Mixture-of-Experts models from SSD, achieving near-fp16 performance on consumer hardware with low memory usage.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback