Tag
This paper presents a specialization pipeline for post-training language models to achieve gold-medal performance in coding competitions, demonstrating top scores on IOI benchmarks using techniques like supervised fine-tuning and reinforcement learning.