Tag
This paper investigates the predictive value of earlier expert selections in sparse mixture-of-experts models beyond the most recent layer, finding significant gains in routing prediction across layers.