Tag
This paper challenges the common assumption that domain experts for model merging should be trained to their optimal validation loss, showing that the optimal training duration depends strongly on the merging method. Simple averaging degrades with overfitting while sparsification-based methods benefit from training past the optimum, suggesting that training duration and merging method should be chosen jointly.