Tag
This paper introduces TestHallVQA, a multi-image VQA benchmark for evaluating Large Vision-Language Models' document-level reasoning under redundant contexts, and proposes a new metric F1-R2 to quantify computational reasoning and robustness.