Tag
The paper proposes a causal decomposition framework that isolates whether MLLM failures on compositional tasks stem from intrinsic capability deficits or cascading errors from upstream prerequisites, and introduces CADET, a benchmark of 46 unit tasks with over 33,000 annotated questions, revealing that supplying critical prerequisites can eliminate much of the apparent reasoning deficit.
An EMNLP Findings 2026 paper introduces a multi-agent framework (story, image, and critic agents) that generates over 9,000 multimodal fake news posts and benchmarks 16 open- and closed-source MLLMs, finding they fall short of human-level detection accuracy, especially at judging image authenticity.
ThinkV2V introduces a reasoning-driven framework that activates MLLM thinking before visual generation for instruction-guided video editing, using an MLLM-to-DiT architecture with progressive curriculum training and inference-time thinking scaling. The authors also release the ThinkV2V-150K dataset and ThinkV2V-Bench, showing their 5B DiT model outperforms larger 10B baselines.
MaLiang-Harness introduces a unified framework to address the Program-to-Visual gap in executable programs for image and video generation, evaluating multiple MLLMs on benchmarks and revealing insights into visual generation performance.
The paper proposes DEEPO, a dual-stage reinforcement learning optimization method to reduce hallucination in multimodal large language models by addressing weaknesses in the correction chain from reward to parameter update.
This paper introduces Modl, a technique for creating cross-lingual remote-sensing multimodal large language models by composing domain and language LoRAs with mutual orthogonality, achieving superior performance without paired multilingual data.
VideoMM introduces an adaptive macro-micro inference framework that reduces visual token overhead in video MLLMs, achieving a 6.13× speedup and 7.4% accuracy gain for efficient long-form video understanding.
OmniHarness introduces a framework for generalizable visual generation using symbolic policy learning, addressing limitations in multimodal large language models and multi-agent systems, and achieving strong performance on benchmarks like ComfyBench.
ReactVAU introduces a slow-fast decoupled framework for real-time streaming video anomaly understanding, leveraging a fast detection module, persistent anomaly-aware memory, and on-demand slow reasoning to enhance efficiency and performance.
CORE introduces a distillation method that transfers compositional ranking judgments from a reranker to an embedding model using a Rank-KL objective, enhancing compositional retrieval performance across benchmarks without compromising standard tasks.
MetaSpace is a framework that applies metamorphic testing to evaluate spatial cognition in embodied agents, generating test cases from execution trajectories and encoding logical/physical constraints as Prolog rules. Benchmarking shows state-of-the-art MLLM-driven agents score far below human levels on spatial cognition, with directional tasks being especially weak.
Introduces MoCA (Implicit Social Context Analysis), a new task and benchmark for modeling implicit social scenarios across affection, intent, and stance, along with a Conflict-Driven Abductive Reasoning (CoDAR) framework. Experiments show state-of-the-art multimodal LLMs struggle on this task, while CoDAR improves performance but still lags behind human reasoning.
This paper introduces PRISM, a four-stage data synthesis framework for training multimodal LLMs to follow prioritized rubrics, and PRISM-Eval, a judge-free evaluation suite. With only 10K samples, PRISM lifts Qwen3-VL-4B from 9.5% to 30.1% Strict accuracy on PRISM-Eval while preserving general benchmark performance, and gains transfer to other open-source MLLMs.
This paper introduces FaceVid-Forensics-100K, a large-scale deepfake video dataset with fine-grained annotations, and proposes a multi-agent forensic reasoning framework (ARGUS) that uses four specialized expert agents and a judge agent to outperform closed-source models on deepfake detection.
This paper introduces Visualized Task Semantics (VTS), a controlled intervention to study when multimodal LLMs are asked questions in the image rather than text. It reveals a consistent accuracy drop across all tested models and tasks, and proposes prompt-region grounding to recover the clean task representation, improving VTS accuracy from 58.0 to 66.3.
This paper introduces C4, a cognition-inspired evaluation framework for cross-concept understanding using Chinese idioms (Chengyu), and finds that current multimodal LLMs struggle with creatively encoded meaning.
PaDoc introduces a layout-grounded parallel decoding method for end-to-end document parsers, decoupling layout and content decoding to reduce decoding depth and improve throughput. It achieves state-of-the-art results on OmniDocBenchFull and significantly speeds up inference compared to sequential baselines.
This paper introduces ArtECulture, a benchmark for culture-conditioned visual emotion understanding in multimodal large language models, covering English, Chinese, and Arabic cultures with balanced Western and non-Western artwork. Evaluations reveal the task remains challenging, and the authors propose a retrieval-augmented framework to inject cultural knowledge into MLLMs.
This paper introduces Structured All-Mask Prediction and STAMPlus, a method for MLLM-based segmentation that jointly predicts all target masks in one non-autoregressive pass, resolving the trilemma of high segmentation performance, preserved dialogue ability, and fast inference.
This paper introduces a Multi-dimensional Evaluation-Verification Reward (EVR) for reinforcement learning fine-tuning of multi-reference image editing models, improving visual consistency and harmony.