mllm

Tag

Cards List
#mllm

Where MLLMs Fail and Why: Causal Task Decomposition for Capability Failure Diagnosis

arXiv cs.CL ↗ · 4d ago Cached

The paper proposes a causal decomposition framework that isolates whether MLLM failures on compositional tasks stem from intrinsic capability deficits or cascading errors from upstream prerequisites, and introduces CADET, a benchmark of 46 unit tasks with over 33,000 annotated questions, revealing that supplying critical prerequisites can eliminate much of the apparent reasoning deficit.

0 favorites 0 likes
#mllm

Can Multimodal Large Language Models Generate and Detect Multimodal Social Media Fake News?

arXiv cs.CL ↗ · 5d ago Cached

An EMNLP Findings 2026 paper introduces a multi-agent framework (story, image, and critic agents) that generates over 9,000 multimodal fake news posts and benchmarks 16 open- and closed-source MLLMs, finding they fall short of human-level detection accuracy, especially at judging image authenticity.

0 favorites 0 likes
#mllm

ThinkV2V: Unleashing the Reasoning Capability of MLLMs for Instruction-Guided Video Editing

Hugging Face Daily Papers ↗ · 6d ago Cached

ThinkV2V introduces a reasoning-driven framework that activates MLLM thinking before visual generation for instruction-guided video editing, using an MLLM-to-DiT architecture with progressive curriculum training and inference-time thinking scaling. The authors also release the ThinkV2V-150K dataset and ThinkV2V-Bench, showing their 5B DiT model outperforms larger 10B baselines.

0 favorites 0 likes
#mllm

MaLiang-Harness: A Programmable Path to Image and Video Generation

Hugging Face Daily Papers ↗ · 2026-09-28 Cached

MaLiang-Harness introduces a unified framework to address the Program-to-Visual gap in executable programs for image and video generation, evaluating multiple MLLMs on benchmarks and revealing insights into visual generation performance.

0 favorites 0 likes
#mllm

DEEPO: Dual-Entropy Enhanced Policy Optimization for Hallucination in MLLMs

arXiv cs.AI ↗ · 2026-09-25 Cached

The paper proposes DEEPO, a dual-stage reinforcement learning optimization method to reduce hallucination in multimodal large language models by addressing weaknesses in the correction chain from reward to parameter update.

0 favorites 0 likes
#mllm

One Domain, Many Tongues: Composing Domain and Language LoRAs for Cross-Lingual Remote-Sensing MLLMs without Paired Data

arXiv cs.CL ↗ · 2026-09-23 Cached

This paper introduces Modl, a technique for creating cross-lingual remote-sensing multimodal large language models by composing domain and language LoRAs with mutual orthogonality, achieving superior performance without paired multilingual data.

0 favorites 0 likes
#mllm

VideoMM: Adaptive Macro-Micro Inference for Efficient Video MLLMs

arXiv cs.AI ↗ · 2026-09-16 Cached

VideoMM introduces an adaptive macro-micro inference framework that reduces visual token overhead in video MLLMs, achieving a 6.13× speedup and 7.4% accuracy gain for efficient long-form video understanding.

0 favorites 0 likes
#mllm

OmniHarness: Harnessing Generalizable Visual Generation via Symbolic Policy Learning

arXiv cs.LG ↗ · 2026-09-16 Cached

OmniHarness introduces a framework for generalizable visual generation using symbolic policy learning, addressing limitations in multimodal large language models and multi-agent systems, and achieving strong performance on benchmarks like ComfyBench.

0 favorites 0 likes
#mllm

ReactVAU: A Slow-Fast Decoupled Framework for Streaming Video Anomaly Understanding

Hugging Face Daily Papers ↗ · 2026-09-07 Cached

ReactVAU introduces a slow-fast decoupled framework for real-time streaming video anomaly understanding, leveraging a fast detection module, persistent anomaly-aware memory, and on-demand slow reasoning to enhance efficiency and performance.

0 favorites 0 likes
#mllm

CORE: Improving Compositional Reasoning in MLLM Embedding via Reranker Distillation

Hugging Face Daily Papers ↗ · 2026-09-03 Cached

CORE introduces a distillation method that transfers compositional ranking judgments from a reranker to an embedding model using a Rank-KL objective, enhancing compositional retrieval performance across benchmarks without compromising standard tasks.

0 favorites 0 likes
#mllm

MetaSpace: Metamorphic Testing for Spatial Cognition in Embodied Agents

arXiv cs.AI ↗ · 2026-08-11 Cached

MetaSpace is a framework that applies metamorphic testing to evaluate spatial cognition in embodied agents, generating test cases from execution trajectories and encoding logical/physical constraints as Prolog rules. Benchmarking shows state-of-the-art MLLM-driven agents score far below human levels on spatial cognition, with directional tasks being especially weak.

0 favorites 0 likes
#mllm

MoCA: Implicit Social Context Analysis

arXiv cs.CL ↗ · 2026-08-07 Cached

Introduces MoCA (Implicit Social Context Analysis), a new task and benchmark for modeling implicit social scenarios across affection, intent, and stance, along with a Conflict-Driven Abductive Reasoning (CoDAR) framework. Experiments show state-of-the-art multimodal LLMs struggle on this task, while CoDAR improves performance but still lags behind human reasoning.

0 favorites 0 likes
#mllm

PRISM: Priority-aware Rubric Internalization via Structured Multimodal Data Synthesis

arXiv cs.LG ↗ · 2026-08-07 Cached

This paper introduces PRISM, a four-stage data synthesis framework for training multimodal LLMs to follow prioritized rubrics, and PRISM-Eval, a judge-free evaluation suite. With only 10K samples, PRISM lifts Qwen3-VL-4B from 9.5% to 30.1% Strict accuracy on PRISM-Eval while preserving general benchmark performance, and gains transfer to other open-source MLLMs.

0 favorites 0 likes
#mllm

Multi-Agent Forensic Reasoning for Generalizable Deepfake Video Detection

Hugging Face Daily Papers ↗ · 2026-08-07 Cached

This paper introduces FaceVid-Forensics-100K, a large-scale deepfake video dataset with fine-grained annotations, and proposes a multi-agent forensic reasoning framework (ARGUS) that uses four specialized expert agents and a judge agent to outperform closed-source models on deepfake detection.

0 favorites 0 likes
#mllm

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning

arXiv cs.AI ↗ · 2026-08-06 Cached

This paper introduces Visualized Task Semantics (VTS), a controlled intervention to study when multimodal LLMs are asked questions in the image rather than text. It reveals a consistent accuracy drop across all tested models and tasks, and proposes prompt-region grounding to recover the clean task representation, improving VTS accuracy from 58.0 to 66.3.

0 favorites 0 likes
#mllm

Can MLLMs Decode the Creative Leap? Introducing C4 for Cross-Concept Understanding

Hugging Face Daily Papers ↗ · 2026-08-06 Cached

This paper introduces C4, a cognition-inspired evaluation framework for cross-concept understanding using Chinese idioms (Chengyu), and finds that current multimodal LLMs struggle with creatively encoded meaning.

0 favorites 0 likes
#mllm

PaDoc: Layout-Grounded Parallel Decoding for Document Parsing

Hugging Face Daily Papers ↗ · 2026-08-06 Cached

PaDoc introduces a layout-grounded parallel decoding method for end-to-end document parsers, decoupling layout and content decoding to reduce decoding depth and improve throughput. It achieves state-of-the-art results on OmniDocBenchFull and significantly speeds up inference compared to sequential baselines.

0 favorites 0 likes
#mllm

ArtECulture: Benchmarking Culture-Conditioned Visual Emotion Understanding in Multimodal Large Language Models

arXiv cs.CL ↗ · 2026-08-05 Cached

This paper introduces ArtECulture, a benchmark for culture-conditioned visual emotion understanding in multimodal large language models, covering English, Chinese, and Arabic cultures with balanced Western and non-Western artwork. Evaluations reveal the task remains challenging, and the authors propose a retrieval-augmented framework to inject cultural knowledge into MLLMs.

0 favorites 0 likes
#mllm

Better, Stronger, Faster, and Broader: Structured All-Mask Prediction for MLLM-Based Segmentation

Hugging Face Daily Papers ↗ · 2026-08-03 Cached

This paper introduces Structured All-Mask Prediction and STAMPlus, a method for MLLM-based segmentation that jointly predicts all target masks in one non-autoregressive pass, resolving the trilemma of high segmentation performance, preserved dialogue ability, and fast inference.

0 favorites 0 likes
#mllm

Evaluation-Verification Reward for Consistent Multi-Reference Image Editing

Hugging Face Daily Papers ↗ · 2026-07-31 Cached

This paper introduces a Multi-dimensional Evaluation-Verification Reward (EVR) for reinforcement learning fine-tuning of multi-reference image editing models, improving visual consistency and harmony.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback