post-training-pruning

Tag

Cards List
#post-training-pruning

PALS: Percentile-Aware Layerwise Sparsity for LLM Pruning

arXiv cs.CL · 2026-07-09 Cached

PALS adjusts per-layer sparsity ratios for LLM pruning based on the 99th percentile of activation magnitudes, achieving significant perplexity improvements on LLaMA-2-7B compared to uniform sparsity, with negligible added cost.

0 favorites 0 likes
← Back to home

Submit Feedback