Tag
This paper proposes Domain-Normalized Multi-Teacher On-Policy Distillation (DN-MOPD) to address unbalanced feedback when merging specialist language models, showing performance improvements on benchmarks like mathematics and instruction-following.