Tag
The paper introduces CRATE, a two-stage framework using step-level consequence reasoning to evaluate mobile agents, achieving high F1-scores on benchmarks like AndroidWorld and MobileRisk.
This paper introduces SFS-DPO, a reinforcement learning two-stage framework for step-level self-verification and self-correction in LLMs, with a teacher-assisted variant SFS-DPO-R. It demonstrates improvements in self-correction effectiveness across multiple LLMs with less training data than prior approaches.