Tag
The paper analyzes on-policy distillation, revealing it primarily suppresses low-probability tokens rather than relying on teacher guidance, and introduces OPSA, a supervision-free method that significantly enhances reasoning performance.
ByteOtter replicates tensor-level allocation on Qwen 3.5 4B, achieving a 16.67% relative improvement in reasoning performance with only a 0.412% increase in model size, marking the first cross-family application outside Gemma.