Tag
Introduces a language interface for visual evidence attribution in document understanding, using verbatim quotes instead of coordinate-based bounding boxes, achieving significantly higher evidence recall and lower attribution hallucination on CiteVQA.
CiteVQA is a benchmark for document vision-language models that evaluates both answer correctness and citation of supporting evidence, revealing widespread attribution hallucinations where models provide correct answers but cite wrong regions.