multimodal-large-language-models

Tag

Cards List
#multimodal-large-language-models

Reasoning with Memory: A Temporal Granularity-Adaptive Framework for Training-Free Long Video Understanding

arXiv cs.AI · 2d ago Cached

ReMem introduces a dual-level memory-augmented keyframe selection framework for training-free long video understanding, achieving state-of-the-art zero-shot performance on multiple benchmarks.

0 favorites 0 likes
#multimodal-large-language-models

ConsiSpace: Learning Geometric Consistency Matters for Video Spatial Reasoning

Hugging Face Daily Papers · 2026-07-20 Cached

ConsiSpace is a geometry-consistency-aware framework for video spatial reasoning that improves performance on spatial reasoning benchmarks by using a geometry-consistent memory and unified consistency self-supervised reinforcement learning.

0 favorites 0 likes
#multimodal-large-language-models

SpaR3D-MoE: Adaptive 3D Spatial Reasoning from Sparse Views Meets Geometry-Inductive Mixture-of-Experts

arXiv cs.AI · 2026-07-09 Cached

Introduces SpaR3D-MoE, an end-to-end framework for adaptive 3D spatial reasoning from sparse RGB views, using manifold sampling and geometry-inductive mixture-of-experts to achieve state-of-the-art performance on VSI-Bench, ScanQA, and SQA3D.

0 favorites 0 likes
#multimodal-large-language-models

BaFCo: A Document Understanding Benchmark for Complex Bangla Form Comprehension

arXiv cs.CL · 2026-07-08 Cached

Introduces BaFCo, a benchmark dataset for Bangla form comprehension focusing on Document Layout Analysis (DLA) and Key Information Extraction (KIE). It includes 200 multi-page complex Bangladeshi government forms with fine-grained annotations across 26 entity types and evaluates multiple MLLMs, revealing limitations in understanding complex Bangla forms.

0 favorites 0 likes
#multimodal-large-language-models

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models

arXiv cs.CL · 2026-06-26 Cached

This survey paper systematically reviews the paradigm evolution of unified vision-language perception in multimodal large language models (MLLMs), proposing a five-stage taxonomy and identifying open challenges toward general multimodal intelligence.

0 favorites 0 likes
#multimodal-large-language-models

Adapting Reinforcement Learning with Chain-of-Thought Supervision for Explainable Detection of Hateful and Propagandistic Memes

arXiv cs.CL · 2026-06-16 Cached

Proposes a reinforcement learning-based post-training method using Group Relative Policy Optimization (GRPO) and chain-of-thought supervision to improve classification and explanation quality for hateful and propagandistic meme detection in thinking-based multimodal large language models, achieving improvements on the Hateful Memes and ArMeme benchmarks.

0 favorites 0 likes
#multimodal-large-language-models

P3D-Bench: Benchmarking MLLMs for Parametric 3D Generation and Structural Reasoning

Hugging Face Daily Papers · 2026-06-09 Cached

This paper introduces P3D-Bench, a benchmark for evaluating multimodal large language models on parametric 3D generation tasks, including text-to-3D, image-to-3D, and assembly-3D, with metrics for geometric precision, semantic alignment, and part-level structure.

0 favorites 0 likes
#multimodal-large-language-models

Optical Reasoning: Rethinking Images as an Expressive Reasoning Medium Beyond Text

Hugging Face Daily Papers · 2026-06-08 Cached

Proposes optical reasoning, using images as a standalone reasoning medium for language and multimodal tasks, achieving higher token efficiency than traditional text-based approaches.

0 favorites 0 likes
#multimodal-large-language-models

Robust-U1: Can MLLMs Self-Recover Corrupted Visual Content for Robust Understanding?

Hugging Face Daily Papers · 2026-06-06 Cached

Robust-U1 is a framework that enables multimodal large language models (MLLMs) to self-recover corrupted visual content using supervised fine-tuning, reinforcement learning with dual rewards, and joint multimodal reasoning, achieving state-of-the-art robustness on corruption benchmarks.

0 favorites 0 likes
#multimodal-large-language-models

Learning Geometric Representations from Videos for Spatial Intelligent Multimodal Large Language Models

Hugging Face Daily Papers · 2026-06-04 Cached

GeoVR enhances multimodal large language models with 3D awareness by restructuring their semantic latent space through geometric knowledge distillation from 3D foundation models using multiple geometric targets.

0 favorites 0 likes
#multimodal-large-language-models

Benchmarking Visual State Tracking in Multimodal Video Understanding

Hugging Face Daily Papers · 2026-06-02 Cached

Introduces VSTAT, a benchmark for evaluating visual state tracking in multimodal large language models (MLLMs) using 834 clips and 1,500 questions. Current MLLMs perform poorly compared to humans, failing at visual perception rather than reasoning.

0 favorites 0 likes
#multimodal-large-language-models

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding

Hugging Face Daily Papers · 2026-05-18 Cached

SWIM is a novel training strategy that aligns vision and language representations for fine-grained object understanding using only textual prompts, leveraging mask supervision during training to improve cross-modal attention. It introduces the NL-Refer dataset and achieves superior performance over visual-prompt-based methods.

0 favorites 0 likes
#multimodal-large-language-models

Video2GUI: Synthesizing Large-Scale Interaction Trajectories for Generalized GUI Agent Pretraining

arXiv cs.CL · 2026-05-15 Cached

Proposes Video2GUI, a framework to automatically extract GUI interaction trajectories from unlabeled instructional videos, building WildGUI dataset with 12M trajectories across 1500+ apps. Pre-training on this data yields 5-20% improvements on GUI grounding and action benchmarks.

0 favorites 0 likes
#multimodal-large-language-models

CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence

Hugging Face Daily Papers · 2026-05-13 Cached

CiteVQA is a benchmark for document vision-language models that evaluates both answer correctness and citation of supporting evidence, revealing widespread attribution hallucinations where models provide correct answers but cite wrong regions.

0 favorites 0 likes
#multimodal-large-language-models

EDU-CIRCUIT-HW: Evaluating Multimodal Large Language Models on Real-World University-Level STEM Student Handwritten Solutions

Hugging Face Daily Papers · 2026-04-30 Cached

This paper introduces the EDU-CIRCUIT-HW dataset for evaluating multimodal large language models on real-world university-level STEM handwritten solutions, revealing significant recognition limitations and proposing a hybrid approach that combines automated recognition with minimal human oversight to enhance grading robustness.

0 favorites 0 likes
← Back to home

Submit Feedback