Tag
This paper proposes SATS (Sensitivity-Aware Thresholding for Sparsity) and a token routing framework to improve inference efficiency in LLMs by dynamically sparsifying MLP activations. The methods achieve better quality-throughput trade-offs compared to percentile-based baselines.
This paper proposes DPVR-LF, a modality-asymmetric routing framework for MLLMs that routes vision tokens at their saturation point into a lightweight side branch and performs late fusion, reducing visual computation while maintaining competitive performance.
This paper presents Token-Selective Attention (TSA), a differentiable token routing mechanism that learns to skip unnecessary computations per token in transformer layers, reducing token-layer operations by 14–23% with minimal quality loss on language modeling tasks.