cross-modal

Tag

Cards List
#cross-modal

MME-Safety: A Fine-grained Benchmark for Safety Evaluation of MLLMs

arXiv cs.CL ↗ · 2026-09-21 Cached

MME-Safety is a rigorously verified benchmark for evaluating the safety of Multimodal Large Language Models, featuring a four-dimensional annotation schema and a hierarchical framework to assess risk scenarios, harm severity, and modality-specific stealth levels.

0 favorites 0 likes
#cross-modal

Robust Fault Detection in Mechanical Multimodal Time Series via Self-Supervised Cross-Modal Reconstruction

arXiv cs.LG ↗ · 2026-09-16 Cached

This paper proposes a multimodal anomaly detection framework for fault detection in mechanical systems using self-supervised cross-modal reconstruction and adaptive thresholding to improve robustness under distribution shifts.

0 favorites 0 likes
#cross-modal

Omni-Streaming Thinking

Hugging Face Daily Papers ↗ · 2026-09-14 Cached

Omni-Streaming Thinking improves streaming omni-modal reasoning by deferring claims until cross-modal verification, reducing premature commitment and auditory hallucinations.

0 favorites 0 likes
#cross-modal

OmniHallu: Unified Hallucination Detection for Cross-Modal Comprehension and Generation in Multimodal Large Language Models

arXiv cs.CL ↗ · 2026-09-11 Cached

OmniHallu introduces a unified hallucination detection framework for multimodal large language models, covering comprehension and generation tasks across image, video, and audio modalities, with a benchmark and multi-agent architecture.

0 favorites 0 likes
#cross-modal

Tri-PvP: Exposing Modality Bias in Omni-Modal Large Language Models through Perceptual-Propositional Evidence Conflicts

Hugging Face Daily Papers ↗ · 2026-09-05 Cached

The paper introduces Tri-PvP, a benchmark that exposes visual bias and asymmetric evidence-form preferences in omni-modal large language models, revealing deep-seated modality biases that are linearly decodable from early layers and resistant to surface mitigation.

0 favorites 0 likes
#cross-modal

HalluPrism: When Multimodal Uncertainty Should Diagnose, Not Decide

arXiv cs.LG ↗ · 2026-09-01 Cached

This paper introduces HalluPrism, a behavioral diagnostic method for multimodal large language models that uses visual perturbation probes to identify hallucination failure modes, improving failure-family classification over confidence-only methods.

0 favorites 0 likes
#cross-modal

VA-Judger: Reward Modeling from Human Preference Feedback for Joint Video-Audio Generation

Hugging Face Daily Papers ↗ · 2026-08-19 Cached

VA-Judger is the first reward model for joint video-audio generation that uses human preference feedback to evaluate holistic quality, including a dataset and benchmark, and demonstrates significant improvements in human preference rates when applied to models like LTX-2.

0 favorites 0 likes
#cross-modal

The Unwritten Benchmark: A New Challenge for Multimodal Machine Learning in Abstract Perceptual Reasoning

arXiv cs.AI ↗ · 2026-08-18 Cached

This paper introduces The Unwritten Benchmark, a new challenge to evaluate abstract perceptual reasoning in multimodal AI models, revealing a significant performance gap between humans and current models like GPT-4o and Gemini 2.5-Pro.

0 favorites 0 likes
#cross-modal

HC-RAG: Evidence-Centric Retrieval-Augmented Generation over Heterogeneous Financial Filings

arXiv cs.CL ↗ · 2026-08-14 Cached

This paper introduces HC-RAG, a hierarchical cross-modal retrieval-augmented generation framework for evidence-centric financial question answering over 10-K filings, along with a new benchmark Multi-Doc-2025. It outperforms RAPTOR and GraphRAG on financial QA benchmarks, especially for long-document and table-related queries.

0 favorites 0 likes
#cross-modal

Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin

arXiv cs.AI ↗ · 2026-08-10 Cached

This paper introduces a method to predict middle-layer attention in multimodal LLMs to prune visual tokens efficiently, using question-contrastive teacher selection and cross-modal attention distillation.

0 favorites 0 likes
#cross-modal

Vision-Language Grounding as Bidirectional Concept Correspondence

Hugging Face Daily Papers ↗ · 2026-08-08 Cached

This paper introduces ConCor-1, a grounding model that treats vision-language grounding as bidirectional concept correspondence, jointly recovering text spans, image segments, and cross-modal matches without prespecified phrases. It unifies phrase grounding, referring expression grounding, and open-vocabulary detection, achieving significant F1 improvements on long-caption and zero-shot LVIS benchmarks.

0 favorites 0 likes
#cross-modal

C$^3$PO: Evaluating Cross-Modal Composition and Counterfactual Performance in Omnimodal Models

arXiv cs.AI ↗ · 2026-08-07 Cached

Introduces C3PO, a benchmark of 3,404 samples for evaluating cross-modal composition and counterfactual reasoning in multimodal LLMs. It finds modality dominance causes most failures, with even the best model (Gemini-3.1-Pro) far below human accuracy.

0 favorites 0 likes
#cross-modal

Detecting Knowledge Inconsistencies Across Text, Tables, and Knowledge Graphs

arXiv cs.CL ↗ · 2026-07-29 Cached

This paper introduces Kontrast, a framework for automatically detecting knowledge inconsistencies across Wikipedia text, tables, and Wikidata knowledge graphs using Text-to-SPARQL and LLM reasoning.

0 favorites 0 likes
#cross-modal

DCS: A Unified Conditional Sensitivity Framework for Cross-Modal Copyright Infringement Detection

arXiv cs.LG ↗ · 2026-07-27 Cached

Proposes a unified post-hoc detection framework for copyright infringement in AI models, using conditional sensitivity and differential privacy to measure memorization across modalities.

0 favorites 0 likes
#cross-modal

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation

Hugging Face Daily Papers ↗ · 2026-07-26 Cached

OmniVAE is a jointly trained audio-video VAE that uses segment-level contrastive learning and feature distillation to align latent spaces, improving joint generation quality and synchronization in text-to-audio-video generation.

0 favorites 0 likes
#cross-modal

Token-Level Cross-Modal Transformer with Contrastive Multi-Task Learning for Breast Cancer Subtype Classification and Survival Prediction

arXiv cs.LG ↗ · 2026-07-21 Cached

This paper proposes UMMT, a token-level cross-modal transformer with contrastive multi-task learning for breast cancer subtype classification and survival prediction, achieving state-of-the-art results on METABRIC and TCGA-BRCA datasets.

0 favorites 0 likes
#cross-modal

Knowledge-Guided Cross-Modal Fusion for Adult-to-Pediatric ECG Transfer via Label-Conditioned Contrastive Alignment

arXiv cs.LG ↗ · 2026-07-20 Cached

Proposes PEACE, a knowledge-guided framework for transferring adult ECG interpretation to pediatric populations using label-conditioned contrastive alignment, achieving significant improvements under limited supervision.

0 favorites 0 likes
#cross-modal

LakeQuest: A Three-Domain Benchmark for Grounded Question Answering across Data Lakes

arXiv cs.CL ↗ · 2026-07-15 Cached

Introduces LakeQuest, a human-validated benchmark of 9,846 QA pairs across three domains for evaluating end-to-end retrieve-and-synthesize pipelines over heterogeneous data lakes, revealing critical failure modes in modern QA systems.

0 favorites 0 likes
#cross-modal

VTaMo: Video-Text Alignment Model for Sign Language Translation

arXiv cs.CL ↗ · 2026-07-13 Cached

VTaMo introduces explicit multi-granularity video-text alignment for sign language translation using optimal transport and contrastive learning, achieving state-of-the-art performance on four benchmarks.

0 favorites 0 likes
#cross-modal

Cross-Modal Generative Framework for Signal Translation from Fetal-Maternal Electrocardiograms to Fetal Doppler Waveforms

arXiv cs.LG ↗ · 2026-07-10 Cached

This paper proposes a cross-modal generative framework that synthesizes fetal Doppler ultrasound waveforms from fetal-maternal electrocardiograms, using cross-modal attention and dilated convolutions, achieving improved synthesis quality and quantifying the influence of maternal-fetal coupling.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback