Tag
The paper introduces RAVEL, a retrieval-aware online reinforcement learning framework for interactive person re-identification that optimizes question selection based on retrieval feedback to improve performance across multiple interaction rounds.
RICE-PO is a critic-free policy optimization framework that turns retrieval interactions into localized credit signals for training reasoning agents, outperforming prompt-based and group-based RL baselines on BRIGHT and BEIR benchmarks.