Tag
This paper deconstructs the reinforcement learning post-training algorithm for large language models, examining how base model distribution, reward signal granularity, and prompt diversity affect post-training outcomes.
This paper presents an architecture that uses formally verified law as a reward signal for training legal AI, adaptively autoformalizing legal rules into a formal calculus and employing a verifier to ensure provable correctness, demonstrated on German and US law examples.
Proposes Cross-Model Entropy (CME) as a label-free reward signal for reinforcement learning post-training of large language models, enabling open-ended instruction following without ground-truth verifiers or human preference labels.