visual-language-models

Tag

Cards List
#visual-language-models

When Thinking Hurts: Epistemic Signals in the Reasoning Chains of Visual Language Models

arXiv cs.LG ↗ · 2026-07-10 Cached

This paper empirically characterizes uncertainty in thinking-mode visual language models, demonstrating that the thinking chain entropy is a more reliable signal for hallucination detection than conventional answer token distribution, which collapses in these models.

0 favorites 0 likes
#visual-language-models

Identifying and Resolving Pitfalls of Knowledge-Based VQA Benchmarks: Auditing, Repairing, and Augmenting

arXiv cs.CL ↗ · 2026-07-02 Cached

This paper audits knowledge-based VQA benchmarks, revealing systematic violations of assumptions that make accuracy a misleading metric. It introduces a repair protocol and multi-entity augmentation to restore answer derivability and question clarity, showing that corrected settings yield markedly different model rankings.

0 favorites 0 likes
#visual-language-models

Zero-Shot Learning in Industrial Scenarios: New Large-Scale Benchmark, Challenges and Baseline

arXiv cs.AI ↗ · 2026-06-09 Cached

This paper proposes a large-scale multi-modal dataset (MMIO) for zero-shot industrial defect detection and introduces the Refined Text-Visual Prompt (RTVP) method, achieving state-of-the-art results on the benchmark.

0 favorites 0 likes
#visual-language-models

From Data to Insights: Exploring Program-of-Thoughts Prompting for Chart Summarization

arXiv cs.CL ↗ · 2026-05-29 Cached

This paper introduces a zero-shot strategy for chart summarization using Program-of-Thoughts prompting, where lightweight visual language models (VLMs) generate Python programs to compute statistics, improving factual accuracy over existing methods.

0 favorites 0 likes
← Back to home

Submit Feedback