threshold-calibration

Tag

Cards List
#threshold-calibration

Sensitivity-Aware Thresholding and Token Routing for Activation Sparsification in Large Language Models

arXiv cs.LG · 2026-07-13 Cached

This paper proposes SATS (Sensitivity-Aware Thresholding for Sparsity) and a token routing framework to improve inference efficiency in LLMs by dynamically sparsifying MLP activations. The methods achieve better quality-throughput trade-offs compared to percentile-based baselines.

0 favorites 0 likes
← Back to home

Submit Feedback