Tag
This paper proposes entropy-regularized cross-entropy (ER-CE) to address the misalignment between cross-entropy loss and verifier objectives in verifiable domains, showing consistent improvements in mathematical reasoning and code generation benchmarks.
Foundational empirical study demonstrating power-law scaling relationships between language model performance and model size, dataset size, and compute budget, with implications for optimal training allocation and sample efficiency.