modality-gap

Tag

Cards List
#modality-gap

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs

arXiv cs.CL · 2026-08-03 Cached

This paper introduces TokenSwap, a method to convert text-only benchmarks into image-interleaved counterparts, and TokenSwap-Bench to measure the modality gap across 42 multimodal LLMs. It finds reasoning models have smaller gaps and shows that TokenSwap-based training can reduce the gap.

0 favorites 0 likes
#modality-gap

Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM

arXiv cs.CL · 2026-05-08 Cached

Proposes TextPro-SLM, a speech large language model that minimizes the modality gap by processing spoken input to resemble prosody-aware text input, achieving strong paralinguistic understanding with low training data.

0 favorites 0 likes
#modality-gap

Anisotropic Modality Align

Hugging Face Daily Papers · 2026-05-08 Cached

This paper proposes AnisoAlign, a framework that addresses the modality gap in multimodal models by applying anisotropic geometric correction to enable effective unpaired modality alignment.

0 favorites 0 likes
#modality-gap

Do Vision-Language Models Truly Perform Vision Reasoning? A Rigorous Study of the Modality Gap

arXiv cs.CL · 2026-04-20 Cached

This paper introduces CrossMath, a controlled multimodal reasoning benchmark that reveals a critical limitation in current vision-language models: they perform reasoning primarily in textual space rather than genuine vision-grounded reasoning, with visual input often degrading performance compared to text-only baselines. The authors propose fine-tuning approaches to mitigate this modality gap and improve multimodal reasoning capabilities.

0 favorites 0 likes
← Back to home

Submit Feedback