Tag
This paper proposes SATS (Sensitivity-Aware Thresholding for Sparsity) and a token routing framework to improve inference efficiency in LLMs by dynamically sparsifying MLP activations. The methods achieve better quality-throughput trade-offs compared to percentile-based baselines.