Tag
This paper challenges the interpretation of the Platonic Representation Hypothesis by distinguishing between relational structure and metric geometry, showing that relational convergence is robust while metric geometry convergence is weaker in various models after calibration.
PRISM-VLM is a multi-axis discriminative benchmark that evaluates compact vision-language models along seven axes, including task quality and behavioral robustness, to provide more reliable separation and insights compared to single-axis benchmarks. It aims to release an open pipeline for the community to improve AI evaluation methods.
SpecialEduBench is a benchmark introduced to evaluate vision-language models on their ability to perform language intervention for autistic children, assessing knowledge, skill, and attitude dimensions.
Spatial-Interactor is a framework that trains vision-language models to enhance spatial reasoning through interaction with the physical world, employing a three-level curriculum and two-stage training strategy to improve state transition modeling and long-horizon integration.
The paper introduces CounterCredit, a training method for vision-language agents that ensures visual calls are both needed and used, leading to higher performance and fewer spurious calls on benchmarks.
The study identifies 'ECG Mirage' in vision-language models for clinical prediction, where models underutilize ECG data despite apparent multimodal capability, and proposes visual prompt tuning as an efficient mitigation strategy.
This paper presents RoboDawn, a method to transfer Vision-Language Model intelligence to robotic control, achieving state-of-the-art results on benchmarks with zero-shot and one-shot learning and successful real-world applications.
Uni-LaDiR introduces a unified latent diffusion framework for multimodal reasoning, mapping modality-specific thoughts into a shared latent space and using diffusion to generate reasoning steps, achieving improved performance on vision-language benchmarks.
MintAct is a family of vision-language models that unifies UI grounding, multi-step navigation across mobile, desktop, and web, and visual tool use, achieving state-of-the-art performance on various benchmarks.
This paper proposes PIVOT, a dual-level learning framework that enhances visually-grounded reasoning in large vision-language models by using self-calibrated experience replay and vision-guided advantage allocation to optimize reinforcement learning.
This paper proposes a collaborative memory framework for multi-agent vision-language model systems to address distributed perception and improve shared visual context and reasoning consistency.
CapMem is a human-annotated benchmark for episodic memory in egocentric video using captions, showing that caption-based QA outperforms direct video QA on long videos.
This study audits CLIP models for gender bias in Metropolitan Museum artwork metadata, finding no statistically significant bias but emphasizing the need for multivariate confound control in AI fairness assessments.
This paper introduces a novel paradigm, VLM-as-probabilistic-grounder, which models uncertainty in vision-language model groundings as probability distributions for symbolic belief-space planning, enhancing robustness in partially observable settings.
ReDraft is a reference-driven revision method for continual post-training of large vision-language models that balances learning new tasks and preserving old ones, achieving higher accuracy and less forgetting than standard approaches like SFT.
The paper presents a method to distill reasoning into compact video-language models using synthetic chain-of-thought rationales and difficulty-aware fine-tuning, enabling smaller models to outperform larger ones with minimal compute.
This paper proposes a method to enhance referring expression generation in vision-language models by leveraging listener gaze data for training, leading to more efficient and successful communication.
This paper investigates how vision-language models read exact values from vertical bar charts using controlled counterfactual activation patching, revealing insights into internal computations and differences between models like Qwen and InternVL.
A tweet by @LLMJunky highlights an impressive application of Vision-Language Models (VLM) in AI, sharing enthusiasm about its cool capabilities.
PhysBrain 1.5 is a unified model that integrates physical environment understanding, action generation, and future state prediction via autoregressive training, achieving state-of-the-art open-source performance on 28 embodied benchmarks.