router-tuning

Tag

Cards List
#router-tuning

Increasing active parameters per token in MOE (Qwen 35B A4B+) reduce reasoning token by 8.5% - and you don't need to train or finetune!

Reddit r/LocalLLaMA · yesterday

A paper introduces an inference-time optimization for sparse MoE models by adjusting expert selection in late transformer layers, reducing reasoning tokens by 8.5% and latency by 10.9% without retraining, while maintaining accuracy.

0 favorites 0 likes
← Back to home

Submit Feedback