policy-entropy

Tag

Cards List
#policy-entropy

Demystifying Reinforcement Learning Post-Training of Language Models

arXiv cs.LG · 3d ago Cached

This paper deconstructs the reinforcement learning post-training algorithm for large language models, examining how base model distribution, reward signal granularity, and prompt diversity affect post-training outcomes.

0 favorites 0 likes
← Back to home

Submit Feedback