lvlm

Tag

Cards List
#lvlm

SpatialBlock: Enhancing Spatial Intelligence in LVLMs via Synthetic Block-Stacking Problem

Hugging Face Daily Papers ↗ · 2026-09-07 Cached

This paper introduces SpatialBlock-15k, a synthetic dataset for block-stacking problems, to enhance 3D spatial reasoning in large vision-language models, demonstrating improved performance and generalization to real-world tasks.

0 favorites 0 likes
#lvlm

When Seeing Is Not Enough: Benchmarking Interactive Visual Grounding in LVLMs

arXiv cs.AI ↗ · 2026-08-26 Cached

This paper introduces a controlled evaluation framework for interactive visual grounding in large vision-language models (LVLMs), showing that current LVLMs perform below human baselines and struggle with proactive question-driven grounding.

0 favorites 0 likes
#lvlm

Learning from Failures: Retrieval-Centric CoT via Hard Negatives for Unified Multimodal Retrieval

Hugging Face Daily Papers ↗ · 2026-08-06 Cached

This paper introduces UniME-R1, an embedder-adviser framework for unified multimodal retrieval that generates Retrieval-Centric Chain-of-Thought (RC-CoT) conditioned on retrieval feedback, improving retrieval performance by learning from hard negatives.

0 favorites 0 likes
#lvlm

[R] CausalVLBench: Benchmarking Visual Causal Reasoning in Large VLMs.

Reddit r/MachineLearning ↗ · 2026-08-02 Cached

This arXiv paper introduces CausalVLBench, a benchmark for evaluating visual causal reasoning in large vision-language models across three tasks: causal structure inference, intervention target prediction, and counterfactual prediction. It evaluates open-source LVLMs on three causal representation learning datasets, revealing strengths and weaknesses.

0 favorites 0 likes
#lvlm

Can LVLMs Uncover the Truth Behind Visual Illusions? An Analysis of Perceptual and Reasoning Capabilities

arXiv cs.CL ↗ · 2026-07-31 Cached

This paper introduces IllusionReasoning, a benchmark using real-world visual illusions to jointly evaluate the perception and reasoning capabilities of Large Vision Language Models (LVLMs), finding that current models' reasoning abilities are not as advanced as claimed.

0 favorites 0 likes
#lvlm

Implicit vs. Explicit Prompting Strategies for LVLMs in Referential Communication

arXiv cs.CL ↗ · 2026-06-17 Cached

This paper investigates seemingly contradictory findings on whether large vision-language models (LVLMs) can coordinate efficient referring expressions. The authors show that models can achieve efficiency when explicitly prompted, but fail to infer the need for efficiency from implicit prompts, revealing key differences between human and AI communication.

0 favorites 0 likes
#lvlm

UniDoc-RL: Coarse-to-Fine Visual RAG with Hierarchical Actions and Dense Rewards

Hugging Face Daily Papers ↗ · 2026-04-16 Cached

UniDoc-RL presents a reinforcement learning framework for Large Vision-Language Models that optimizes retrieval, reranking, and visual reasoning through hierarchical decision-making and dense multi-reward supervision, achieving up to 17.7% improvements over prior RL-based methods on visual RAG tasks.

0 favorites 0 likes
← Back to home

Submit Feedback