Tag
A paper introduces an inference-time optimization for sparse MoE models by adjusting expert selection in late transformer layers, reducing reasoning tokens by 8.5% and latency by 10.9% without retraining, while maintaining accuracy.