multimodal-llms

Tag

Cards List
#multimodal-llms

@BenjaminDEKR: BenchBench: multimodal LLMs compete to assemble Ikea furniture

X AI KOLs Following ↗ · 2d ago

BenchBench is a benchmark that challenges multimodal LLMs to compete in assembling Ikea furniture, testing their practical AI capabilities.

0 favorites 0 likes
#multimodal-llms

Beyond Task Completion: Training Capable and Safe Computer-Use Agents

arXiv cs.LG ↗ · 2026-09-22 Cached

This paper introduces SCOPE, a joint training method for computer-use agents that improves both task completion and safety, using a synthesized dataset and achieving strong performance on benchmarks OSWorld and OS-BLIND.

0 favorites 0 likes
#multimodal-llms

Dissecting Training-Free Uncertainty Estimation in Multimodal Large Language Models

arXiv cs.CL ↗ · 2026-09-22 Cached

A systematic study benchmarking training-free uncertainty quantification strategies for multimodal Large Language Models, categorizing methods into token-level, verbalized, and semantic approaches and finding optimal strategies depend on response length.

0 favorites 0 likes
#multimodal-llms

PolyBridgeBench: Benchmarking Multimodal LLMs for Physics-Grounded Bridge Design

arXiv cs.AI ↗ · 2026-09-21 Cached

PolyBridgeBench is a new executable benchmark for evaluating multimodal LLMs on physics-grounded bridge design tasks, revealing gaps between deterministic validity and dynamic success in structure synthesis and repair.

0 favorites 0 likes
#multimodal-llms

Layers, Sinks, and Scaling: Adaptive Evidence Selection for Multimodal Large Language Models

arXiv cs.AI ↗ · 2026-09-16 Cached

This paper presents AREA, a training-free inference-time method that adaptively allocates evidence highlighting in multimodal large language models, improving performance on knowledge-based visual question answering and standard multimodal benchmarks.

0 favorites 0 likes
#multimodal-llms

OmniHallu: Unified Hallucination Detection for Cross-Modal Comprehension and Generation in Multimodal Large Language Models

arXiv cs.CL ↗ · 2026-09-11 Cached

OmniHallu introduces a unified hallucination detection framework for multimodal large language models, covering comprehension and generation tasks across image, video, and audio modalities, with a benchmark and multi-agent architecture.

0 favorites 0 likes
#multimodal-llms

ReactHuman: A Physics-Grounded Benchmark for Human-Like Reactive Decision-Making in Embodied Multimodal LLMs

Hugging Face Daily Papers ↗ · 2026-09-09 Cached

ReactHuman is a physics-grounded benchmark that evaluates multimodal large language models for human-like reactive decision-making in simulated humanoid robots facing household hazards, revealing persistent safety failures and offering a scalable diagnostic tool.

0 favorites 0 likes
#multimodal-llms

MLLMs Hallucinate when Information Distribution Drifts in Synergy Heads

Hugging Face Daily Papers ↗ · 2026-09-05 Cached

The paper proposes HEAL, a method to mitigate hallucinations in Multimodal Large Language Models by analyzing and calibrating information distribution in synergy heads through causal interventions and counterfactual analysis.

0 favorites 0 likes
#multimodal-llms

@LeeLeepenkman: Text-Anchored Semantic Perturbations for Transferable Jailbreak Attacks on Multimodal Large Language Models https://pap…

X AI KOLs Timeline ↗ · 2026-08-26 Cached

The paper introduces TA-SPA, a black-box jailbreak attack framework for multimodal large language models that uses text-anchored semantic perturbations to achieve effective and transferable attacks against safety alignments.

0 favorites 0 likes
#multimodal-llms

Video-IFBench: Evaluating Instruction Following of Multimodal LLMs in Video Understanding Scenarios

Hugging Face Daily Papers ↗ · 2026-08-26 Cached

Video-IFBench introduces a comprehensive benchmark for evaluating how well multimodal large language models follow diverse instructions in video understanding tasks, covering various constraints and task types.

0 favorites 0 likes
#multimodal-llms

Beyond Success and Failure: Length-Aware Contrastive Learning for GUI Agents

arXiv cs.AI ↗ · 2026-08-25 Cached

The paper proposes Length-Aware Contrastive Learning for GUI Agents (LACL-GUI), a contrastive reinforcement learning framework that incorporates trajectory-level quality signals to improve agent performance by addressing reward-gradient misalignment.

0 favorites 0 likes
#multimodal-llms

Annotations as Rollouts: Efficient and Scalable Reinforcement Learning for Video MLLMs

Hugging Face Daily Papers ↗ · 2026-08-20 Cached

OraRL improves reinforcement learning post-training for video multimodal language models by integrating annotations as oracle rollouts, achieving higher sample efficiency and performance without chain-of-thought generation.

0 favorites 0 likes
#multimodal-llms

@BenjaminDEKR: We do not yet have multimodal LLMs for smells

X AI KOLs Timeline ↗ · 2026-08-16

The tweet points out that multimodal large language models capable of processing smells do not yet exist, highlighting a current limitation in AI technology.

0 favorites 0 likes
#multimodal-llms

Looks Right, Works Right: A Project-Level Benchmark for Multi-Screen Mobile App Generation

arXiv cs.AI ↗ · 2026-08-03 Cached

MobileForge is a benchmark for project-level multi-screen mobile app generation, evaluating multimodal LLMs on build success, cross-page navigation, visual fidelity, maintainability, and efficiency. Experiments on six frontier multimodal LLMs show current models can compile and reach pages but still struggle with interactive navigation and visual quality.

0 favorites 0 likes
#multimodal-llms

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs

arXiv cs.CL ↗ · 2026-08-03 Cached

This paper introduces TokenSwap, a method to convert text-only benchmarks into image-interleaved counterparts, and TokenSwap-Bench to measure the modality gap across 42 multimodal LLMs. It finds reasoning models have smaller gaps and shows that TokenSwap-based training can reduce the gap.

0 favorites 0 likes
#multimodal-llms

Symphony of Bias: Exploring Gender Associations with Musical Instruments in Multimodal LLMs

arXiv cs.CL ↗ · 2026-07-30 Cached

This paper introduces Symphony-Bias, a multimodal dataset for evaluating gender bias in LLMs' associations with musical instruments across text, vision, and audio. It finds that 92% of instrument-level outcomes align with prior social-science findings, with gender biases strongest in text and weakest in audio.

0 favorites 0 likes
#multimodal-llms

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models

Hugging Face Daily Papers ↗ · 2026-07-28 Cached

This paper investigates why multimodal LLMs fail on vision-centric tasks when visual evidence conflicts with language priors, introducing the WhatIfVis benchmark and showing that models often encode visual evidence but cannot reliably control their reliance on it.

0 favorites 0 likes
#multimodal-llms

Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and Challenges

arXiv cs.CL ↗ · 2026-07-22 Cached

A comprehensive survey on multimodal humor understanding using large language models, covering methods, datasets, evaluation protocols, and challenges in interpreting humor in memes, cartoons, and comics.

0 favorites 0 likes
#multimodal-llms

An Exam for Active Observers

arXiv cs.CL ↗ · 2026-07-20 Cached

This paper introduces ActiveVision, a benchmark to evaluate active observation in multimodal large language models. Frontier models like GPT-5.5 and Claude Fable 5 perform poorly, solving only 10.6% and 3.5% of tasks respectively, compared to human 96.1%, highlighting a lack of iterative visual perception.

0 favorites 0 likes
#multimodal-llms

Low-cost concept-based localized explanations: How far can we get with training-free approaches?

arXiv cs.AI ↗ · 2026-06-30 Cached

This paper evaluates the zero-shot capability of multimodal large language models (MLLMs) for localized concept naming in images, proposing a reproducible evaluation protocol that achieves 62-88% object-level accuracy without training.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback