Tag
The paper proposes OneModel, a unified AI paradigm that consolidates complex business workflows into a single model's parameters, achieving over 50% latency reduction and improved accuracy in financial services through continual pre-training and logic compilation.
This paper introduces ICRL, a framework that jointly trains a solver and critic with reinforcement learning to internalize critique guidance, enabling the solver to improve without external critique. It uses distribution calibration and role-wise group advantage estimation, achieving 6-7 point gains over GRPO on agentic and mathematical reasoning tasks.