Tag
The paper investigates on-policy distillation of large language models, demonstrating that a single training query can achieve substantial state coverage and alignment, suggesting the method is algorithm-starved rather than data-starved.