activation-aware

Tag

Cards List
#activation-aware

Batch-wise Adaptive Pruning: Periodic Neuron Activation-Aware Weight Pruning for Language Reasoning Model

arXiv cs.CL · 2026-08-17 Cached

This paper proposes a training-free adaptive pruning method for large reasoning models during batched inference, using periodic top-k selection and activation memory to improve accuracy and computational efficiency.

0 favorites 0 likes
#activation-aware

PALS: Percentile-Aware Layerwise Sparsity for LLM Pruning

arXiv cs.CL · 2026-07-09 Cached

PALS adjusts per-layer sparsity ratios for LLM pruning based on the 99th percentile of activation magnitudes, achieving significant perplexity improvements on LLaMA-2-7B compared to uniform sparsity, with negligible added cost.

0 favorites 0 likes
#activation-aware

SigmaScale: LLM Compression with SVD-based Low-Rank Decomposition and Learned Scaling Matrices

arXiv cs.CL · 2026-06-08 Cached

Introduces SigmaScale, a method that learns auxiliary scaling matrices for SVD-based LLM compression, showing competitive performance on Llama 3.1 8B and Qwen3-8B benchmarks.

0 favorites 0 likes
#activation-aware

QAM-W: Joint 2D Codebook Quantization for LLM Weights via Hadamard Rotation and Activation-Aware Scaling

arXiv cs.LG · 2026-05-27 Cached

Introduces QAM-W, a joint 2D codebook quantization method for LLM weights using Hadamard rotation and activation-aware scaling, achieving near BF16 perplexity at 5–6 bits per weight and matching SmoothQuant W8A8 quality with 32% fewer weight bits.

0 favorites 0 likes
← Back to home

Submit Feedback