reasoning-performance

Tag

Cards List
#reasoning-performance

Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement

Hugging Face Daily Papers · 2026-08-31 Cached

The paper analyzes on-policy distillation, revealing it primarily suppresses low-probability tokens rather than relying on teacher guidance, and introduces OPSA, a supervision-free method that significantly enhances reasoning performance.

0 favorites 0 likes
#reasoning-performance

Qwen 3.5 4B IQ2_XS: +16.67% Reasoning Performance From Tensor-Level Allocation

Reddit r/LocalLLaMA · 2026-08-22

ByteOtter replicates tensor-level allocation on Qwen 3.5 4B, achieving a 16.67% relative improvement in reasoning performance with only a 0.412% increase in model size, marking the first cross-family application outside Gemma.

0 favorites 0 likes
← Back to home

Submit Feedback