Tag
This paper empirically characterizes uncertainty in thinking-mode visual language models, demonstrating that the thinking chain entropy is a more reliable signal for hallucination detection than conventional answer token distribution, which collapses in these models.
This paper audits knowledge-based VQA benchmarks, revealing systematic violations of assumptions that make accuracy a misleading metric. It introduces a repair protocol and multi-entity augmentation to restore answer derivability and question clarity, showing that corrected settings yield markedly different model rankings.
This paper proposes a large-scale multi-modal dataset (MMIO) for zero-shot industrial defect detection and introduces the Refined Text-Visual Prompt (RTVP) method, achieving state-of-the-art results on the benchmark.
This paper introduces a zero-shot strategy for chart summarization using Program-of-Thoughts prompting, where lightweight visual language models (VLMs) generate Python programs to compute statistics, improving factual accuracy over existing methods.