Tag
BenchBench is a benchmark that challenges multimodal LLMs to compete in assembling Ikea furniture, testing their practical AI capabilities.
This paper introduces SCOPE, a joint training method for computer-use agents that improves both task completion and safety, using a synthesized dataset and achieving strong performance on benchmarks OSWorld and OS-BLIND.
A systematic study benchmarking training-free uncertainty quantification strategies for multimodal Large Language Models, categorizing methods into token-level, verbalized, and semantic approaches and finding optimal strategies depend on response length.
PolyBridgeBench is a new executable benchmark for evaluating multimodal LLMs on physics-grounded bridge design tasks, revealing gaps between deterministic validity and dynamic success in structure synthesis and repair.
This paper presents AREA, a training-free inference-time method that adaptively allocates evidence highlighting in multimodal large language models, improving performance on knowledge-based visual question answering and standard multimodal benchmarks.
OmniHallu introduces a unified hallucination detection framework for multimodal large language models, covering comprehension and generation tasks across image, video, and audio modalities, with a benchmark and multi-agent architecture.
ReactHuman is a physics-grounded benchmark that evaluates multimodal large language models for human-like reactive decision-making in simulated humanoid robots facing household hazards, revealing persistent safety failures and offering a scalable diagnostic tool.
The paper proposes HEAL, a method to mitigate hallucinations in Multimodal Large Language Models by analyzing and calibrating information distribution in synergy heads through causal interventions and counterfactual analysis.
The paper introduces TA-SPA, a black-box jailbreak attack framework for multimodal large language models that uses text-anchored semantic perturbations to achieve effective and transferable attacks against safety alignments.
Video-IFBench introduces a comprehensive benchmark for evaluating how well multimodal large language models follow diverse instructions in video understanding tasks, covering various constraints and task types.
The paper proposes Length-Aware Contrastive Learning for GUI Agents (LACL-GUI), a contrastive reinforcement learning framework that incorporates trajectory-level quality signals to improve agent performance by addressing reward-gradient misalignment.
OraRL improves reinforcement learning post-training for video multimodal language models by integrating annotations as oracle rollouts, achieving higher sample efficiency and performance without chain-of-thought generation.
The tweet points out that multimodal large language models capable of processing smells do not yet exist, highlighting a current limitation in AI technology.
MobileForge is a benchmark for project-level multi-screen mobile app generation, evaluating multimodal LLMs on build success, cross-page navigation, visual fidelity, maintainability, and efficiency. Experiments on six frontier multimodal LLMs show current models can compile and reach pages but still struggle with interactive navigation and visual quality.
This paper introduces TokenSwap, a method to convert text-only benchmarks into image-interleaved counterparts, and TokenSwap-Bench to measure the modality gap across 42 multimodal LLMs. It finds reasoning models have smaller gaps and shows that TokenSwap-based training can reduce the gap.
This paper introduces Symphony-Bias, a multimodal dataset for evaluating gender bias in LLMs' associations with musical instruments across text, vision, and audio. It finds that 92% of instrument-level outcomes align with prior social-science findings, with gender biases strongest in text and weakest in audio.
This paper investigates why multimodal LLMs fail on vision-centric tasks when visual evidence conflicts with language priors, introducing the WhatIfVis benchmark and showing that models often encode visual evidence but cannot reliably control their reliance on it.
A comprehensive survey on multimodal humor understanding using large language models, covering methods, datasets, evaluation protocols, and challenges in interpreting humor in memes, cartoons, and comics.
This paper introduces ActiveVision, a benchmark to evaluate active observation in multimodal large language models. Frontier models like GPT-5.5 and Claude Fable 5 perform poorly, solving only 10.6% and 3.5% of tasks respectively, compared to human 96.1%, highlighting a lack of iterative visual perception.
This paper evaluates the zero-shot capability of multimodal large language models (MLLMs) for localized concept naming in images, proposing a reproducible evaluation protocol that achieves 62-88% object-level accuracy without training.