vision-language-models

Tag

Cards List
#vision-language-models

What Converges in the Platonic Representation Hypothesis? Structure over Geometry

arXiv cs.LG ↗ · 15h ago Cached

This paper challenges the interpretation of the Platonic Representation Hypothesis by distinguishing between relational structure and metric geometry, showing that relational convergence is robust while metric geometry convergence is weaker in various models after calibration.

0 favorites 0 likes
#vision-language-models

PRISM-VLM: A Multi-Axis Discriminative Benchmark for Compact Vision-Language Models

arXiv cs.CL ↗ · 15h ago Cached

PRISM-VLM is a multi-axis discriminative benchmark that evaluates compact vision-language models along seven axes, including task quality and behavioral robustness, to provide more reliable separation and insights compared to single-axis benchmarks. It aims to release an open pipeline for the community to improve AI evaluation methods.

0 favorites 0 likes
#vision-language-models

SpecialEduBench: Benchmarking Vision-Language Models on Knowledge, Skill, and Attitude in Language Intervention for Autistic Children

arXiv cs.CL ↗ · yesterday Cached

SpecialEduBench is a benchmark introduced to evaluate vision-language models on their ability to perform language intervention for autistic children, assessing knowledge, skill, and attitude dimensions.

0 favorites 0 likes
#vision-language-models

Spatial-Interactor: Learning Spatial Reasoning through Interaction with the Observable Physical World

arXiv cs.AI ↗ · yesterday Cached

Spatial-Interactor is a framework that trains vision-language models to enhance spatial reasoning through interaction with the physical world, employing a three-level curriculum and two-stage training strategy to improve state transition modeling and long-horizon integration.

0 favorites 0 likes
#vision-language-models

When Should a VLM Look? Paying Only for Visual Calls That Were Needed and Used

arXiv cs.AI ↗ · yesterday Cached

The paper introduces CounterCredit, a training method for vision-language agents that ensures visual calls are both needed and used, leading to higher performance and fewer spurious calls on benchmarks.

0 favorites 0 likes
#vision-language-models

ECG Mirage: Revealing and Mitigating the Underutilisation of ECGs in Vision-Language Models for Clinical Prediction

arXiv cs.AI ↗ · 3d ago Cached

The study identifies 'ECG Mirage' in vision-language models for clinical prediction, where models underutilize ECG data despite apparent multimodal capability, and proposes visual prompt tuning as an efficient mitigation strategy.

0 favorites 0 likes
#vision-language-models

Transferring the Intelligence of VLMs to Robotic Control

Hugging Face Daily Papers ↗ · 5d ago Cached

This paper presents RoboDawn, a method to transfer Vision-Language Model intelligence to robotic control, achieving state-of-the-art results on benchmarks with zero-shot and one-shot learning and successful real-world applications.

0 favorites 0 likes
#vision-language-models

Uni-LaDiR: Latent Diffusion Unifies Multimodal Reasoning

arXiv cs.LG ↗ · 6d ago Cached

Uni-LaDiR introduces a unified latent diffusion framework for multimodal reasoning, mapping modality-specific thoughts into a shared latent space and using diffusion to generate reasoning steps, achieving improved performance on vision-language benchmarks.

0 favorites 0 likes
#vision-language-models

MintAct: A Unified Visual Agent for Digital Environments

Hugging Face Daily Papers ↗ · 6d ago Cached

MintAct is a family of vision-language models that unifies UI grounding, multi-step navigation across mobile, desktop, and web, and visual tool use, achieving state-of-the-art performance on various benchmarks.

0 favorites 0 likes
#vision-language-models

Anchoring What Matters: A Dual-Level Learning Framework for Visually-Grounded Multimodal Reasoning

arXiv cs.AI ↗ · 2026-09-17 Cached

This paper proposes PIVOT, a dual-level learning framework that enhances visually-grounded reasoning in large vision-language models by using self-calibrated experience replay and vision-guided advantage allocation to optimize reinforcement learning.

0 favorites 0 likes
#vision-language-models

Collaborative Memory for Multi-Agent VLM Systems

arXiv cs.AI ↗ · 2026-09-17 Cached

This paper proposes a collaborative memory framework for multi-agent vision-language model systems to address distributed perception and improve shared visual context and reasoning consistency.

0 favorites 0 likes
#vision-language-models

CapMem: A Benchmark for Caption-Based Episodic Memory in Egocentric Video

arXiv cs.AI ↗ · 2026-09-17 Cached

CapMem is a human-annotated benchmark for episodic memory in egocentric video using captions, showing that caption-based QA outperforms direct video QA on long videos.

0 favorites 0 likes
#vision-language-models

Disentangling Algorithmic Bias from Archival Artifacts: A Controlled Audit of Vision-Language Model Valuation in Metropolitan Museum Archives

arXiv cs.LG ↗ · 2026-09-17 Cached

This study audits CLIP models for gender bias in Metropolitan Museum artwork metadata, finding no statistically significant bias but emphasizing the need for multivariate confound control in AI fairness assessments.

0 favorites 0 likes
#vision-language-models

Bridging Learned Visual Perception and Symbolic Belief-Space Planning

arXiv cs.AI ↗ · 2026-09-16 Cached

This paper introduces a novel paradigm, VLM-as-probabilistic-grounder, which models uncertainty in vision-language model groundings as probability distributions for symbolic belief-space planning, enhancing robustness in partially observable settings.

0 favorites 0 likes
#vision-language-models

ReDraft, Don't Just Distill: Reference-Driven Revision for Continual VLLM Post-Training

arXiv cs.AI ↗ · 2026-09-16 Cached

ReDraft is a reference-driven revision method for continual post-training of large vision-language models that balances learning new tasks and preserving old ones, achieving higher accuracy and less forgetting than standard approaches like SFT.

0 favorites 0 likes
#vision-language-models

Efficient Reasoning Distillation: Small Video-Language Models via Synthetic CoT and Difficulty-Aware Fine-Tuning

arXiv cs.LG ↗ · 2026-09-16 Cached

The paper presents a method to distill reasoning into compact video-language models using synthetic chain-of-thought rationales and difficulty-aware fine-tuning, enabling smaller models to outperform larger ones with minimal compute.

0 favorites 0 likes
#vision-language-models

Learning to Refer from Estimated Listener Gaze

arXiv cs.CL ↗ · 2026-09-15 Cached

This paper proposes a method to enhance referring expression generation in vision-language models by leveraging listener gaze data for training, leading to more efficient and successful communication.

0 favorites 0 likes
#vision-language-models

Inside VLM Chart Reading: Tracing Value Reading from Vertical Bar Charts Across Space and Depth

arXiv cs.CL ↗ · 2026-09-15 Cached

This paper investigates how vision-language models read exact values from vertical bar charts using controlled counterfactual activation patching, revealing insights into internal computations and differences between models like Qwen and InternVL.

0 favorites 0 likes
#vision-language-models

@LLMJunky: This is one of the coolest applications of AI (VLM) that I've ever seen.

X AI KOLs Following ↗ · 2026-09-14 Cached

A tweet by @LLMJunky highlights an impressive application of Vision-Language Models (VLM) in AI, sharing enthusiasm about its cool capabilities.

0 favorites 0 likes
#vision-language-models

PhysBrain 1.5: From Vision-Language Models to Physical Foundation Models

Hugging Face Daily Papers ↗ · 2026-09-14 Cached

PhysBrain 1.5 is a unified model that integrates physical environment understanding, action generation, and future state prediction via autoregressive training, achieving state-of-the-art open-source performance on 28 embodied benchmarks.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback