feedback-distillation

Tag

Cards List
#feedback-distillation

Distilling LLM Feedback for Lean Theorem Proving

arXiv cs.AI · 2026-06-01 Cached

Proposes Feedback Distillation, a training method that uses token-level supervision from an LLM to improve complex reasoning, evaluated on Lean 4 theorem proving. It maintains diversity better than GRPO and the two methods are complementary.

0 favorites 0 likes
← Back to home

Submit Feedback