Tag
UniCAR-RL introduces a reinforcement learning framework that decouples perception and reasoning to improve multimodal large language models' visual mathematical reasoning without annotation.
The paper introduces multimodal contextual sycophancy in large language models, where external text overrides visual evidence, and proposes a diagnostic method using System-2 Visual Arbitration to improve performance.
CoVA-SFT is a large-scale dataset and benchmark designed to train multimodal language models in chain of visual abstractions, improving performance over baselines in visual reasoning tasks.
TempCloze is a video cloze benchmark for evaluating visual temporal reasoning in Video-LLMs, requiring models to identify missing video segments from distractors to reduce linguistic shortcuts in semantics, alignment, and progression.
The author questions traditional mobile agent approaches relying on API access and proposes using screen understanding via hardware like aiden-firmware to enable more general-purpose agents that interact like humans.
This paper proposes Evidence-First Reflection (EFR) to improve reflection in desktop GUI agents by decoupling action-induced visual difference extraction from outcome verification, yielding accuracy gains of 7.11% on benchmarks.
VBVR-Pro introduces a closed-loop testbed for scalable and verifiable native visual reasoning through generation, featuring task scaling, verifiable rewards, and mechanism studies across diverse visual substrates.
VGI-Bench evaluates visual reasoning in video generation models through 27 tasks, revealing limited reliability and self-correction in current systems.
The paper introduces Internalized Visual Thinking (IVT), a post-training framework that trains multimodal models to predict future frame embeddings, enabling direct answer generation without synthesizing intermediate images, thereby reducing latency over 5x for proactive video reasoning.
This paper introduces Counterfactual Evidence Disentanglement (CED), a training-time method that makes vision-language models rely on concrete image evidence rather than language priors or shortcuts, improving visual reasoning grounding across benchmarks.
This paper proposes a test-time alignment approach for large vision-language models using trajectory-guided structured sampling and iterative MCMC refinement, improving visual reasoning accuracy without heavy post-training.
Introduces OPD-V, a visual on-policy self-distillation paradigm for multimodal large language models that leverages positive and negative teachers to exploit modality balance as privileged information, improving reasoning performance across benchmarks while reducing training cost.
This paper introduces Beacon, an agentic visual reasoning model that improves multimodal LLMs' ability to decide when to use tools and benefit from tool use, using Necessity-Aware Adaptive Reward and Hint-Guided Capability Expansion in reinforcement learning.
Visual prompt engineering (VIPE) automatically modifies task images to improve video model reasoning performance, often more effective than text-based prompting or test-time scaling.
Proposes NOPD, a self-distillation method that improves vision-language models without external supervision by leveraging prediction discrepancies between clean and corrupted inputs. Achieves significant gains on visual reasoning tasks, matching or exceeding RL and distillation from external models.
This paper introduces Visual Prompt Engineering (VIPE), a method that automatically modifies task images to improve video model performance, showing it can be more effective than text-based prompt engineering or test-time scaling.
TRACE is a taxonomy-guided environment with 1,000 visual reasoning tasks across 11 domains. Training Qwen2.5-VL-3B and Qwen2.5-VL-7B on 64,000 TRACE instances improves their macro-average performance across 24 external benchmarks by 3.51 and 4.06 percentage points respectively.
HDR is a unified framework integrating hierarchical latents into causal video generation for multi-step visual reasoning, achieving better reasoning consistency, lower latency, and strong data efficiency compared to baselines.
UniVR introduces VR-GRPO, a reinforcement learning paradigm for unified visual reasoning, learning complex reasoning and physical dynamics from pure visual demonstrations, achieving up to 25% improvement on the VR-X benchmark.
Muse Spark 1.1 outperforms GPT-5.6 Sol and Gemini 3.1 on Radiology's Last Exam 2.0, a new visual reasoning benchmark for autonomous AI diagnosis in healthcare, though it still lags behind Fable and human radiologists.