data-minimal

Tag

Cards List
#data-minimal

Rethinking On-Policy Distillation of Large Language Models II: One Training Example

Hugging Face Daily Papers · 2d ago Cached

The paper investigates on-policy distillation of large language models, demonstrating that a single training query can achieve substantial state coverage and alignment, suggesting the method is algorithm-starved rather than data-starved.

0 favorites 0 likes
← Back to home

Submit Feedback