InSight-doc: Agentic Visual Perception for Long-Document Understanding
Summary
InSight-doc is an agentic visual perception framework for long-document understanding that adaptively allocates visual resolution during reasoning, reducing hallucination and inference latency while improving accuracy on document VQA benchmarks. The paper releases an 8B model, datasets, and code.
View Cached Full Text
Cached at: 08/12/26, 12:19 PM
Paper page - InSight-doc: Agentic Visual Perception for Long-Document Understanding
Source: https://huggingface.co/papers/2608.10628
Abstract
InSight-doc adaptively allocates visual resolution during reasoning to improve long-document understanding while reducing latency and hallucinations.
Long-document understanding often requires reasoning over many visually rich pages, making inference costly and prone to context rot. In this work, we propose InSight-doc, anagentic visual perceptionframework that treats visual resolution as anadaptive reasoning-time resource. InSight-doc starts from low resolution and selectively zooms into high-resolution regions for finer evidence, without relying on any external retriever. To train such an agent, we construct anactive-perception corpusof 17.9K high-qualitySFTexamples withregion-level zoom-intrajectories, accompanied by 19.2K hardRLexamples. ThroughSFT+RL, InSight-doc-8B improves the baseline by 4.3--16.4 accuracy points overdocument VQAbenchmarks. On long documents, it reduceshallucinationby more than 40% andinference latencyby 41%--68% while maintaining an accuracy lead. Our code, datasets, and model are released at https://github.com/m-Just/InSight-doc .
View arXiv pageView PDFProject pageGitHub4Add to collection
Community
Paper submitter
Upload images, audio, and videos by dragging in the text input, pasting, orclicking here.
Tap or paste here to upload images
Get this paper in your agent:
hf papers read 2608\.10628
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper1
#### InSight-doc/InSight-doc-8B Image-Text-to-Text• 9B• Updatedabout 3 hours ago • 90 • 2
Datasets citing this paper2
#### m-Just/InSight-doc-RL-19k Updatedabout 3 hours ago • 2 • 2 #### m-Just/InSight-doc-SFT-18k Viewer• Updatedabout 3 hours ago • 17.9k • 1 • 2
Spaces citing this paper1
Collections including this paper1
Similar Articles
VLD-RAG: Agentic Vision-Language Retrieval-Augmented Generation for Long, Visually-Rich Multi-Page Documents
This paper presents VLD-RAG, an agentic multimodal retrieval-augmented generation framework for question answering over long, visually-rich documents. It uses a page-preserving index and a verifier-guided agent workflow to improve cross-page evidence retrieval and reasoning, outperforming prior vision-based baselines on benchmarks like LongDocURL and MMLongBench-Doc.
InsightEmb: Learning Action-Intent Embeddings for Agentic Insight Retrieval
InsightEmb is a contrastive embedding framework for agentic insight retrieval that learns progress-oriented retrieval geometry from mathematical reasoning data alone, improving retrieval for LLM agents without environment-specific training.
InSight: Self-Guided Skill Acquisition via Steerable VLAs
InSight presents a framework for autonomous skill acquisition in vision-language-action (VLA) models by enabling steerability at the primitive-action level and using a VLM-guided data flywheel to generate demonstrations, achieving manipulation tasks like block flipping and pouring without human demonstrations.
DeepVoyager-VL: Incentivizing Vision-in-the-Loop Search for Long-Horizon Multimodal Agents
DeepVoyager-VL proposes a long-horizon multimodal deep-search framework that integrates visual evidence into intermediate reasoning, using a multimodal event graph for data synthesis and fine-tuning without reinforcement learning, achieving strong performance across ten benchmarks.
VisualThink-VLA: Visual Intermediate Reasoning for Effective and Low-Latency Vision-Language-Action Policies
VisualThink-VLA introduces a visual intermediate reasoning framework for vision-language-action policies that preserves spatial precision and dramatically reduces latency compared to text-based reasoning, achieving sub-second inference and state-of-the-art success rates on robot manipulation benchmarks.