mllm

Tag

Cards List
#mllm

DEEPO: Dual-Entropy Enhanced Policy Optimization for Hallucination in MLLMs

arXiv cs.AI ↗ · 23h ago Cached

The paper proposes DEEPO, a dual-stage reinforcement learning optimization method to reduce hallucination in multimodal large language models by addressing weaknesses in the correction chain from reward to parameter update.

0 favorites 0 likes
#mllm

One Domain, Many Tongues: Composing Domain and Language LoRAs for Cross-Lingual Remote-Sensing MLLMs without Paired Data

arXiv cs.CL ↗ · 2d ago Cached

This paper introduces Modl, a technique for creating cross-lingual remote-sensing multimodal large language models by composing domain and language LoRAs with mutual orthogonality, achieving superior performance without paired multilingual data.

0 favorites 0 likes
#mllm

VideoMM: Adaptive Macro-Micro Inference for Efficient Video MLLMs

arXiv cs.AI ↗ · 2026-09-16 Cached

VideoMM introduces an adaptive macro-micro inference framework that reduces visual token overhead in video MLLMs, achieving a 6.13× speedup and 7.4% accuracy gain for efficient long-form video understanding.

0 favorites 0 likes
#mllm

OmniHarness: Harnessing Generalizable Visual Generation via Symbolic Policy Learning

arXiv cs.LG ↗ · 2026-09-16 Cached

OmniHarness introduces a framework for generalizable visual generation using symbolic policy learning, addressing limitations in multimodal large language models and multi-agent systems, and achieving strong performance on benchmarks like ComfyBench.

0 favorites 0 likes
#mllm

ReactVAU: A Slow-Fast Decoupled Framework for Streaming Video Anomaly Understanding

Hugging Face Daily Papers ↗ · 2026-09-07 Cached

ReactVAU introduces a slow-fast decoupled framework for real-time streaming video anomaly understanding, leveraging a fast detection module, persistent anomaly-aware memory, and on-demand slow reasoning to enhance efficiency and performance.

0 favorites 0 likes
#mllm

CORE: Improving Compositional Reasoning in MLLM Embedding via Reranker Distillation

Hugging Face Daily Papers ↗ · 2026-09-03 Cached

CORE introduces a distillation method that transfers compositional ranking judgments from a reranker to an embedding model using a Rank-KL objective, enhancing compositional retrieval performance across benchmarks without compromising standard tasks.

0 favorites 0 likes
#mllm

MetaSpace: Metamorphic Testing for Spatial Cognition in Embodied Agents

arXiv cs.AI ↗ · 2026-08-11 Cached

MetaSpace is a framework that applies metamorphic testing to evaluate spatial cognition in embodied agents, generating test cases from execution trajectories and encoding logical/physical constraints as Prolog rules. Benchmarking shows state-of-the-art MLLM-driven agents score far below human levels on spatial cognition, with directional tasks being especially weak.

0 favorites 0 likes
#mllm

MoCA: Implicit Social Context Analysis

arXiv cs.CL ↗ · 2026-08-07 Cached

Introduces MoCA (Implicit Social Context Analysis), a new task and benchmark for modeling implicit social scenarios across affection, intent, and stance, along with a Conflict-Driven Abductive Reasoning (CoDAR) framework. Experiments show state-of-the-art multimodal LLMs struggle on this task, while CoDAR improves performance but still lags behind human reasoning.

0 favorites 0 likes
#mllm

PRISM: Priority-aware Rubric Internalization via Structured Multimodal Data Synthesis

arXiv cs.LG ↗ · 2026-08-07 Cached

This paper introduces PRISM, a four-stage data synthesis framework for training multimodal LLMs to follow prioritized rubrics, and PRISM-Eval, a judge-free evaluation suite. With only 10K samples, PRISM lifts Qwen3-VL-4B from 9.5% to 30.1% Strict accuracy on PRISM-Eval while preserving general benchmark performance, and gains transfer to other open-source MLLMs.

0 favorites 0 likes
#mllm

Multi-Agent Forensic Reasoning for Generalizable Deepfake Video Detection

Hugging Face Daily Papers ↗ · 2026-08-07 Cached

This paper introduces FaceVid-Forensics-100K, a large-scale deepfake video dataset with fine-grained annotations, and proposes a multi-agent forensic reasoning framework (ARGUS) that uses four specialized expert agents and a judge agent to outperform closed-source models on deepfake detection.

0 favorites 0 likes
#mllm

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning

arXiv cs.AI ↗ · 2026-08-06 Cached

This paper introduces Visualized Task Semantics (VTS), a controlled intervention to study when multimodal LLMs are asked questions in the image rather than text. It reveals a consistent accuracy drop across all tested models and tasks, and proposes prompt-region grounding to recover the clean task representation, improving VTS accuracy from 58.0 to 66.3.

0 favorites 0 likes
#mllm

Can MLLMs Decode the Creative Leap? Introducing C4 for Cross-Concept Understanding

Hugging Face Daily Papers ↗ · 2026-08-06 Cached

This paper introduces C4, a cognition-inspired evaluation framework for cross-concept understanding using Chinese idioms (Chengyu), and finds that current multimodal LLMs struggle with creatively encoded meaning.

0 favorites 0 likes
#mllm

PaDoc: Layout-Grounded Parallel Decoding for Document Parsing

Hugging Face Daily Papers ↗ · 2026-08-06 Cached

PaDoc introduces a layout-grounded parallel decoding method for end-to-end document parsers, decoupling layout and content decoding to reduce decoding depth and improve throughput. It achieves state-of-the-art results on OmniDocBenchFull and significantly speeds up inference compared to sequential baselines.

0 favorites 0 likes
#mllm

ArtECulture: Benchmarking Culture-Conditioned Visual Emotion Understanding in Multimodal Large Language Models

arXiv cs.CL ↗ · 2026-08-05 Cached

This paper introduces ArtECulture, a benchmark for culture-conditioned visual emotion understanding in multimodal large language models, covering English, Chinese, and Arabic cultures with balanced Western and non-Western artwork. Evaluations reveal the task remains challenging, and the authors propose a retrieval-augmented framework to inject cultural knowledge into MLLMs.

0 favorites 0 likes
#mllm

Better, Stronger, Faster, and Broader: Structured All-Mask Prediction for MLLM-Based Segmentation

Hugging Face Daily Papers ↗ · 2026-08-03 Cached

This paper introduces Structured All-Mask Prediction and STAMPlus, a method for MLLM-based segmentation that jointly predicts all target masks in one non-autoregressive pass, resolving the trilemma of high segmentation performance, preserved dialogue ability, and fast inference.

0 favorites 0 likes
#mllm

Evaluation-Verification Reward for Consistent Multi-Reference Image Editing

Hugging Face Daily Papers ↗ · 2026-07-31 Cached

This paper introduces a Multi-dimensional Evaluation-Verification Reward (EVR) for reinforcement learning fine-tuning of multi-reference image editing models, improving visual consistency and harmony.

0 favorites 0 likes
#mllm

RoCo-ACE: Rollout-Conditioned Online Distillation for Retention-Aware Knowledge Injection

arXiv cs.AI ↗ · 2026-07-29 Cached

This paper introduces RoCo-ACE, a rollout-conditioned online distillation objective for knowledge injection into multimodal large language models. It improves injected knowledge accuracy while limiting drift in non-updated behaviors.

0 favorites 0 likes
#mllm

VlogReward: Learning Multi-Dimensional Evaluation for Vlog Editing

arXiv cs.AI ↗ · 2026-07-28 Cached

Introduces VlogReward, a reward model for evaluating vlog editing plans across six dimensions, along with a large-scale dataset and benchmark, achieving state-of-the-art results against GPT-5 and Gemini-3-Pro.

0 favorites 0 likes
#mllm

DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth

arXiv cs.LG ↗ · 2026-07-21 Cached

DocOCR-Eval proposes an annotation-free framework that uses a correction and ranking strategy to evaluate and select OCR tools without ground truth labels, showing that aggregating multiple multimodal large language models improves alignment with human rankings.

0 favorites 0 likes
#mllm

HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enchancement

Hugging Face Daily Papers ↗ · 2026-07-20 Cached

HOMIE is a framework for human-object centric video personalization that integrates MLLM features to improve subject fidelity and interaction patterns, achieving state-of-the-art performance.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback