structured-pruning

Tag

Cards List
#structured-pruning

LILA: Calibration-Free Structured Pruning of Large Language Models via Latent Spectral Geometry

arXiv cs.LG · 2026-09-11 Cached

LILA is a calibration-free structured pruning method for large language models that uses latent spectral geometry to score neuron importance, achieving competitive performance without calibration data.

0 favorites 0 likes
#structured-pruning

Damage-Aware Bandit Pruning for Vision and Language Transformers

arXiv cs.AI · 2026-09-10 Cached

This paper proposes a damage-aware multi-armed bandit method for structured post-training pruning of vision and language transformers, showing reduced performance degradation compared to baseline approaches in experiments across various models and datasets.

0 favorites 0 likes
#structured-pruning

COEC: Calibrated Orthogonal-Equivalence Compensation for Structured Pruning of Large Language Models

arXiv cs.LG · 2026-08-24 Cached

The paper proposes COEC, a training-free compensation framework for structured pruning of large language models that applies orthogonal rotations and calibration to reduce output error and improve accuracy after column removal.

0 favorites 0 likes
#structured-pruning

Unifying Depth and Width Pruning for LLMs via Binary Knapsack Optimization

arXiv cs.CL · 2026-08-14 Cached

This paper introduces Sniper, a two-stage structured pruning framework for LLMs that uses binary knapsack optimization to unify depth and width pruning, achieving near-exact compression ratio adherence and improved performance retention across multiple architectures.

0 favorites 0 likes
#structured-pruning

TriSP: Tri-Signal Structured Pruning for Large Language Models

arXiv cs.AI · 2026-07-28 Cached

TriSP introduces a tri-signal importance metric combining weight magnitude, activation norm, and gradient sensitivity for structured pruning of LLMs, achieving lowest perplexity and high throughput improvements on LLaMA-7B.

0 favorites 0 likes
#structured-pruning

Multi-Objective Structured Pruning of LLMs for Latency and Model Size Optimization

arXiv cs.AI · 2026-07-28 Cached

Proposes a two-stage structured pruning framework for LLMs that jointly optimizes latency and model size using multi-objective depth pruning and parallel Bayesian optimization, achieving favorable trade-offs for edge deployment.

0 favorites 0 likes
#structured-pruning

ShortOPD: Recovering Pruned LLMs with Short-to-Long On-Policy Distillation

arXiv cs.LG · 2026-07-16 Cached

ShortOPD proposes a short-to-long on-policy distillation schedule that recovers pruned LLMs for free-form generation by focusing training on effective prefixes, achieving up to 9x improvement over unrecovered models and matching long-horizon distillation with a quarter of the training time.

0 favorites 0 likes
#structured-pruning

It Takes a MAESTRO To Prune Bad Experts

arXiv cs.CL · 2026-07-10 Cached

This paper introduces Maestro, a structured pruning framework for Mixture-of-Experts language models that uses Markov chains to model expert activation trajectories, achieving globally aware pruning and outperforming baselines by up to 10.61% under 50% compression.

0 favorites 0 likes
#structured-pruning

Structured Pruning of Large Language Models via Power Transformation and Sign-Preserving Score Aggregation with Adaptive Feature Retention

arXiv cs.CL · 2026-07-10 Cached

This paper proposes a structured pruning method for LLMs that addresses distribution mismatch, sign-information loss, and outlier influence when adapting unstructured pruning techniques, achieving comparable accuracy with 1.56-1.57x speedup on models like Llama-3-8B and Vicuna-v1.5-13B.

0 favorites 0 likes
#structured-pruning

Cascaded Multi-Granularity Pruning for On-Device LLM Inference in Industrial IoT

arXiv cs.CL · 2026-06-26 Cached

This paper presents a cascaded multi-granularity pruning framework for deploying LLMs on Industrial IoT edge devices, achieving up to 13.8x compression with minimal accuracy loss on MHA+GELU architectures while exposing a collapse on GQA+SwiGLU designs.

0 favorites 0 likes
#structured-pruning

Structured Neuron Pruning in Deep Neural Networks Using Multi-Armed Bandits

arXiv cs.LG · 2026-06-09 Cached

This paper proposes a novel structured neuron pruning framework for deep neural networks using multi-armed bandit algorithms, demonstrating effectiveness on various tasks.

0 favorites 0 likes
#structured-pruning

Knowledge Offloading: Decomposing LLMs into Sparse Backbones and Memory Modules

arXiv cs.LG · 2026-05-29 Cached

Proposes KOFF, a framework that decomposes pretrained LLMs into a sparse shared backbone and domain-specific external memories using structured pruning and LoRA adapters, achieving 12% sparsity without significant performance loss.

0 favorites 0 likes
← Back to home

Submit Feedback