document-vqa

Tag

Cards List
#document-vqa

Do MLLMs Really Understand Low-Resource Khmer Documents? A Pilot Study on Khmer Document VQA

arXiv cs.CL · 5d ago Cached

This paper evaluates the performance of open multimodal large language models on Khmer document VQA, finding that while models handle English and numeric content well, understanding native Khmer script remains challenging. The study uses a diagnostic subset from the KH-FUNSD collection and compares different model configurations.

0 favorites 0 likes
#document-vqa

InSight-doc: Agentic Visual Perception for Long-Document Understanding

Hugging Face Daily Papers · 2026-08-11 Cached

InSight-doc is an agentic visual perception framework for long-document understanding that adaptively allocates visual resolution during reasoning, reducing hallucination and inference latency while improving accuracy on document VQA benchmarks. The paper releases an 8B model, datasets, and code.

0 favorites 0 likes
#document-vqa

DocTrace: Towards Traceable Long Document VQA via Hierarchical Evidence Graph Reasoning

arXiv cs.AI · 2026-08-05 Cached

The paper introduces DocTrace, a hierarchical framework for long document visual question answering that casts the task as explicit evidence graph reasoning. It achieves state-of-the-art results on three benchmarks while enabling traceable evidence provenance, outperforming Qwen3-VL-8B-Instruct by 11-14 points.

0 favorites 0 likes
← Back to home

Submit Feedback