reasoning-distillation

Tag

Cards List
#reasoning-distillation

@charles_irl: kino

X AI KOLs Timeline · 2026-07-04 Cached

Ali distilled Claude Fable 5 reasoning traces into Qwen3-4B, claiming 100% self-consistency and zero hallucination variance, and open-sourced the result.

0 favorites 0 likes
#reasoning-distillation

Weight-Space Geometry of Offline Reasoning Training

arXiv cs.LG · 2026-06-24 Cached

This paper investigates whether different offline reinforcement learning losses (RFT, RIFT, DFT, Offline GRPO, DPO) for reasoning distillation produce mechanistically distinct weight updates in a small language model. Using identical math rollouts and a controlled setup with Qwen3-4B and attention-only LoRA, they find that SFT, RFT, and RIFT yield nearly colinear weight deltas, while DPO sits in a near-orthogonal subspace and achieves the highest accuracy.

0 favorites 0 likes
#reasoning-distillation

MARD: Mirror-Augmented Reasoning Distillation for Mechanism-Level Drug-Drug Interaction Prediction

arXiv cs.CL · 2026-06-12 Cached

Introduces MARD, a 7B-parameter model for mechanism-level drug-drug interaction prediction using mirror-augmented reasoning distillation, achieving state-of-the-art accuracy at ~1% of frontier API cost and demonstrating genuine pharmacological reasoning over memorization.

0 favorites 0 likes
#reasoning-distillation

LARK: Learnability-Grounded Trajectory Selection for Efficient Reasoning Distillation

arXiv cs.LG · 2026-06-01 Cached

LARK proposes a learnability-grounded method for selecting reasoning trajectories in LLM distillation, employing a learnability factor and χ²-regularized selection policy that balances efficiency and generalization, consistently outperforming baselines across models and tasks.

0 favorites 0 likes
#reasoning-distillation

Structured Prompt Optimization Meets Reinforcement Learning for Global and Local Interpretability over Complex Text

arXiv cs.CL · 2026-05-29 Cached

Introduces eXTC, a text classifier with three progressive stages: structured prompt optimization to learn a natural-language rulebook, reasoning distillation into a compact LM, and reinforcement learning to expand reasoning, achieving strong performance and interpretability.

0 favorites 0 likes
#reasoning-distillation

Tailoring the Curriculum: Student-Centered Reasoning Distillation via Dynamic Data-Model Compatibility

arXiv cs.AI · 2026-05-29 Cached

Introduces the Data-Model Compatibility (DMC) metric to evaluate how well a reasoning dataset aligns with a student model during distillation. Experiments show DMC strongly correlates with distillation performance and that dynamically selecting datasets based on DMC further improves reasoning capabilities.

0 favorites 0 likes
#reasoning-distillation

Backtracking When It Strays: Mitigating Dual Exposure Biases in LLM Reasoning Distillation

arXiv cs.CL · 2026-05-20 Cached

This paper introduces Motab, a new pipeline for LLM reasoning distillation that mitigates both off-policy and on-policy exposure biases by dynamically monitoring student generation and backtracking to safe states with teacher intervention, achieving ~3% average improvement.

0 favorites 0 likes
#reasoning-distillation

Distribution Corrected Offline Data Distillation for Large Language Models

arXiv cs.CL · 2026-05-15 Cached

This paper proposes a principled offline reasoning distillation framework that corrects teacher-student distribution drift, improving reasoning accuracy on math benchmarks without requiring online rollouts.

0 favorites 0 likes
#reasoning-distillation

hesamation/Qwen3.6-35B-A3B-Claude-4.6-Opus-Reasoning-Distilled-GGUF

Hugging Face Models Trending · 2026-04-18 Cached

A 35B-parameter Qwen3.6 model fine-tuned with Claude-Opus-style chain-of-thought distillation data and released in GGUF quantized formats for efficient local inference.

0 favorites 0 likes
← Back to home

Submit Feedback