Tag
MME-Safety is a rigorously verified benchmark for evaluating the safety of Multimodal Large Language Models, featuring a four-dimensional annotation schema and a hierarchical framework to assess risk scenarios, harm severity, and modality-specific stealth levels.
This paper proposes a multimodal anomaly detection framework for fault detection in mechanical systems using self-supervised cross-modal reconstruction and adaptive thresholding to improve robustness under distribution shifts.
Omni-Streaming Thinking improves streaming omni-modal reasoning by deferring claims until cross-modal verification, reducing premature commitment and auditory hallucinations.
OmniHallu introduces a unified hallucination detection framework for multimodal large language models, covering comprehension and generation tasks across image, video, and audio modalities, with a benchmark and multi-agent architecture.
The paper introduces Tri-PvP, a benchmark that exposes visual bias and asymmetric evidence-form preferences in omni-modal large language models, revealing deep-seated modality biases that are linearly decodable from early layers and resistant to surface mitigation.
This paper introduces HalluPrism, a behavioral diagnostic method for multimodal large language models that uses visual perturbation probes to identify hallucination failure modes, improving failure-family classification over confidence-only methods.
VA-Judger is the first reward model for joint video-audio generation that uses human preference feedback to evaluate holistic quality, including a dataset and benchmark, and demonstrates significant improvements in human preference rates when applied to models like LTX-2.
This paper introduces The Unwritten Benchmark, a new challenge to evaluate abstract perceptual reasoning in multimodal AI models, revealing a significant performance gap between humans and current models like GPT-4o and Gemini 2.5-Pro.
This paper introduces HC-RAG, a hierarchical cross-modal retrieval-augmented generation framework for evidence-centric financial question answering over 10-K filings, along with a new benchmark Multi-Doc-2025. It outperforms RAPTOR and GraphRAG on financial QA benchmarks, especially for long-document and table-related queries.
This paper introduces a method to predict middle-layer attention in multimodal LLMs to prune visual tokens efficiently, using question-contrastive teacher selection and cross-modal attention distillation.
This paper introduces ConCor-1, a grounding model that treats vision-language grounding as bidirectional concept correspondence, jointly recovering text spans, image segments, and cross-modal matches without prespecified phrases. It unifies phrase grounding, referring expression grounding, and open-vocabulary detection, achieving significant F1 improvements on long-caption and zero-shot LVIS benchmarks.
Introduces C3PO, a benchmark of 3,404 samples for evaluating cross-modal composition and counterfactual reasoning in multimodal LLMs. It finds modality dominance causes most failures, with even the best model (Gemini-3.1-Pro) far below human accuracy.
This paper introduces Kontrast, a framework for automatically detecting knowledge inconsistencies across Wikipedia text, tables, and Wikidata knowledge graphs using Text-to-SPARQL and LLM reasoning.
Proposes a unified post-hoc detection framework for copyright infringement in AI models, using conditional sensitivity and differential privacy to measure memorization across modalities.
OmniVAE is a jointly trained audio-video VAE that uses segment-level contrastive learning and feature distillation to align latent spaces, improving joint generation quality and synchronization in text-to-audio-video generation.
This paper proposes UMMT, a token-level cross-modal transformer with contrastive multi-task learning for breast cancer subtype classification and survival prediction, achieving state-of-the-art results on METABRIC and TCGA-BRCA datasets.
Proposes PEACE, a knowledge-guided framework for transferring adult ECG interpretation to pediatric populations using label-conditioned contrastive alignment, achieving significant improvements under limited supervision.
Introduces LakeQuest, a human-validated benchmark of 9,846 QA pairs across three domains for evaluating end-to-end retrieve-and-synthesize pipelines over heterogeneous data lakes, revealing critical failure modes in modern QA systems.
VTaMo introduces explicit multi-granularity video-text alignment for sign language translation using optimal transport and contrastive learning, achieving state-of-the-art performance on four benchmarks.
This paper proposes a cross-modal generative framework that synthesizes fetal Doppler ultrasound waveforms from fetal-maternal electrocardiograms, using cross-modal attention and dilated convolutions, achieving improved synthesis quality and quantifying the influence of maternal-fetal coupling.