InSight-doc: Agentic Visual Perception for Long-Document Understanding

Hugging Face Daily Papers Papers

Summary

InSight-doc is an agentic visual perception framework for long-document understanding that adaptively allocates visual resolution during reasoning, reducing hallucination and inference latency while improving accuracy on document VQA benchmarks. The paper releases an 8B model, datasets, and code.

Long-document understanding often requires reasoning over many visually rich pages, making inference costly and prone to context rot. In this work, we propose InSight-doc, an agentic visual perception framework that treats visual resolution as an adaptive reasoning-time resource. InSight-doc starts from low resolution and selectively zooms into high-resolution regions for finer evidence, without relying on any external retriever. To train such an agent, we construct an active-perception corpus of 17.9K high-quality SFT examples with region-level zoom-in trajectories, accompanied by 19.2K hard RL examples. Through SFT+RL, InSight-doc-8B improves the baseline by 4.3--16.4 accuracy points over document VQA benchmarks. On long documents, it reduces hallucination by more than 40% and inference latency by 41%--68% while maintaining an accuracy lead. Our code, datasets, and model are released at https://github.com/m-Just/InSight-doc .
Original Article
View Cached Full Text

Cached at: 08/12/26, 12:19 PM

Paper page - InSight-doc: Agentic Visual Perception for Long-Document Understanding

Source: https://huggingface.co/papers/2608.10628

Abstract

InSight-doc adaptively allocates visual resolution during reasoning to improve long-document understanding while reducing latency and hallucinations.

Long-document understanding often requires reasoning over many visually rich pages, making inference costly and prone to context rot. In this work, we propose InSight-doc, anagentic visual perceptionframework that treats visual resolution as anadaptive reasoning-time resource. InSight-doc starts from low resolution and selectively zooms into high-resolution regions for finer evidence, without relying on any external retriever. To train such an agent, we construct anactive-perception corpusof 17.9K high-qualitySFTexamples withregion-level zoom-intrajectories, accompanied by 19.2K hardRLexamples. ThroughSFT+RL, InSight-doc-8B improves the baseline by 4.3--16.4 accuracy points overdocument VQAbenchmarks. On long documents, it reduceshallucinationby more than 40% andinference latencyby 41%--68% while maintaining an accuracy lead. Our code, datasets, and model are released at https://github.com/m-Just/InSight-doc .

View arXiv pageView PDFProject pageGitHub4Add to collection

Community

Paper submitter

about 3 hours ago

Upload images, audio, and videos by dragging in the text input, pasting, orclicking here.

Tap or paste here to upload images

Get this paper in your agent:

hf papers read 2608\.10628

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper1

#### InSight-doc/InSight-doc-8B Image-Text-to-Text• 9B• Updatedabout 3 hours ago • 90 • 2

Datasets citing this paper2

#### m-Just/InSight-doc-RL-19k Updatedabout 3 hours ago • 2 • 2 #### m-Just/InSight-doc-SFT-18k Viewer• Updatedabout 3 hours ago • 17.9k • 1 • 2

Spaces citing this paper1

Collections including this paper1

Similar Articles

InSight: Self-Guided Skill Acquisition via Steerable VLAs

Hugging Face Daily Papers

InSight presents a framework for autonomous skill acquisition in vision-language-action (VLA) models by enabling steerability at the primitive-action level and using a VLM-guided data flywheel to generate demonstrations, achieving manipulation tasks like block flipping and pouring without human demonstrations.