UniDoc-RL: Coarse-to-Fine Visual RAG with Hierarchical Actions and Dense Rewards
Summary
UniDoc-RL presents a reinforcement learning framework for Large Vision-Language Models that optimizes retrieval, reranking, and visual reasoning through hierarchical decision-making and dense multi-reward supervision, achieving up to 17.7% improvements over prior RL-based methods on visual RAG tasks.
View Cached Full Text
Cached at: 04/20/26, 08:28 AM
Paper page - UniDoc-RL: Coarse-to-Fine Visual RAG with Hierarchical Actions and Dense Rewards
Source: https://huggingface.co/papers/2604.14967
Abstract
UniDoc-RL introduces a reinforcement learning framework for LVLMs that jointly optimizes retrieval, reranking, visual perception, and reasoning through hierarchical decision-making and dense multi-reward supervision.
Retrieval-Augmented Generation (https://huggingface.co/papers?q=Retrieval-Augmented%20Generation)(RAG) extends Large Vision-Language Models (https://huggingface.co/papers?q=Large%20Vision-Language%20Models)(LVLMs) with external visual knowledge. However, existing visual RAG systems typically rely on generic retrieval signals that overlook the fine-grained visual semantics (https://huggingface.co/papers?q=fine-grained%20visual%20semantics)essential for complex reasoning. To address this limitation, we propose UniDoc-RL, a unified reinforcement learning (https://huggingface.co/papers?q=reinforcement%20learning)framework in which an LVLM agent jointly performs retrieval, reranking, active visual perception (https://huggingface.co/papers?q=active%20visual%20perception), and reasoning. UniDoc-RL formulates visual information acquisition (https://huggingface.co/papers?q=visual%20information%20acquisition)as a sequential decision-making (https://huggingface.co/papers?q=sequential%20decision-making)problem with a hierarchical action space (https://huggingface.co/papers?q=hierarchical%20action%20space). Specifically, it progressively refines visual evidence from coarse-grained document retrieval to fine-grained image selection and active region cropping, allowing the model to suppress irrelevant content and attend to information-dense regions. For effective end-to-end training, we introduce a dense multi-reward scheme (https://huggingface.co/papers?q=dense%20multi-reward%20scheme)that provides task-aware supervision for each action. Based on Group Relative Policy Optimization (https://huggingface.co/papers?q=Group%20Relative%20Policy%20Optimization)(GRPO), UniDoc-RL aligns agent behavior with multiple objectives without relying on a separate value network. To support this training paradigm, we curate a comprehensive dataset of high-quality reasoning trajectories with fine-grained action annotations. Experiments on three benchmarks demonstrate that UniDoc-RL consistently surpasses state-of-the-art baselines, yielding up to 17.7% gains over prior RL-based methods.
View arXiv page (https://arxiv.org/abs/2604.14967)View PDF (https://arxiv.org/pdf/2604.14967)GitHub8 (https://github.com/deepglint/UniDoc-RL)Add to collection (https://huggingface.co/login?next=%2Fpapers%2F2604.14967)
Get this paper in your agent:
hf papers read 2604.14967
Don’t have the latest CLI?curl -LsSf https://hf.co/cli/install.sh | bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2604.14967 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2604.14967 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2604.14967 in a Space README.md to link it from this page.
Collections including this paper2
Similar Articles
UniCAR-RL: Seeing Better before Thinking Deeper in Visual Mathematics
UniCAR-RL introduces a reinforcement learning framework that decouples perception and reasoning to improve multimodal large language models' visual mathematical reasoning without annotation.
UniVR: Thinking in Visual Space for Unified Visual Reasoning
UniVR introduces VR-GRPO, a reinforcement learning paradigm for unified visual reasoning, learning complex reasoning and physical dynamics from pure visual demonstrations, achieving up to 25% improvement on the VR-X benchmark.
EasyVideoR1: Easier RL for Video Understanding
EasyVideoR1 is an efficient reinforcement learning framework for training large vision-language models on video understanding tasks, featuring offline preprocessing with tensor caching for 1.47x throughput improvement, a task-aware reward system covering 11 problem types, and evaluation across 22 video benchmarks. It also supports joint image-video training and a mixed offline-online data training paradigm.
VLD-RAG: Agentic Vision-Language Retrieval-Augmented Generation for Long, Visually-Rich Multi-Page Documents
This paper presents VLD-RAG, an agentic multimodal retrieval-augmented generation framework for question answering over long, visually-rich documents. It uses a page-preserving index and a verifier-guided agent workflow to improve cross-page evidence retrieval and reasoning, outperforming prior vision-based baselines on benchmarks like LongDocURL and MMLongBench-Doc.
Region-Level Policy Optimization for Fine-grained MLLM Perception
Vision-RL² is a method for improving fine-grained perception in multimodal large language models by using region-level reinforcement learning to compress visual tokens and enhance performance across multiple benchmarks without full model fine-tuning.