Tag
Introduces AdaKP, an online adaptive knowledge-point selector that dynamically re-chooses which atomic hints to inject during RL training to mitigate reward sparsity in reasoning tasks, achieving improvements on competition-level math benchmarks with negligible overhead.