Tag
This paper proposes a training-free adaptive pruning method for large reasoning models during batched inference, using periodic top-k selection and activation memory to improve accuracy and computational efficiency.
This study reveals a 'Smart Pruning Paradox' where activation-aware pruning methods like Wanda preserve perplexity but significantly amplify bias in Large Language Models deployed on edge devices.