Tag
Introduces ReliableTableQA, a framework for training LLMs to annotate statistical reliability of tabular QA results, showing that a small SFT set is sufficient and GRPO only helps when SFT is under-trained.
This paper introduces Retention-aware Policy Optimization (RaPO) to mitigate catastrophic forgetting in visual continual learning using reinforcement fine-tuning. RaPO uses trajectory-level reward shaping and cross-task advantage normalization to close the gap between reinforcement and supervised fine-tuning in class- and domain-incremental learning.