mixed-policy-distillation

Tag

Cards List
#mixed-policy-distillation

Reasoning Compression with Mixed-Policy Distillation

arXiv cs.AI · 2026-05-12 Cached

This paper proposes Mixed-Policy Distillation (MPD), a framework that transfers concise reasoning behaviors from large teacher models to smaller student models, reducing token usage by up to 27.1% while improving performance.

0 favorites 0 likes
← Back to home

Submit Feedback