zero-reward

Tag

Cards List
#zero-reward

TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition

arXiv cs.LG · 2026-06-10 Cached

Proposes TD-Grokking, a training-time decomposition framework that recursively breaks down intractable zero-reward problems into verifiable subproblems, enabling LLMs to learn from failed trajectories. Outperforms vanilla GRPO and baselines on mathematical and medical reasoning tasks.

0 favorites 0 likes
← Back to home

Submit Feedback