Tag
PaperGym is a framework that converts scientific papers into training environments by separating research questions from evaluation rubrics, enabling reinforcement learning to improve research planning. It demonstrates improved performance over existing methods, with trained models outperforming larger ones on benchmarks.
Cornell researchers propose POP, a self-play framework that lets an LLM generate its own rubrics and training pairs for open-ended tasks, boosting Qwen-2.5-7B on healthcare QA, creative writing and instruction following without human labels.