visual-reasoning

Tag

Cards List
#visual-reasoning

Visual Reasoning through Tool-supervised Reinforcement Learning

Hugging Face Daily Papers ↗ · 2026-04-21 Cached

Introduces ToolsRL, a two-stage reinforcement learning framework that teaches multimodal LLMs to use simple visual tools for complex visual reasoning tasks.

0 favorites 0 likes
#visual-reasoning

Chain-of-Thought Degrades Visual Spatial Reasoning Capabilities of Multimodal LLMs

Hugging Face Daily Papers ↗ · 2026-04-17 Cached

Research shows Chain-of-Thought prompting harms visual-spatial reasoning in multimodal LLMs due to shortcut learning and hallucinating visual details from text alone.

0 favorites 0 likes
#visual-reasoning

Learning Adaptive Reasoning Paths for Efficient Visual Reasoning

Hugging Face Daily Papers ↗ · 2026-04-16 Cached

AVR is an adaptive visual reasoning framework that dynamically selects optimal reasoning formats to reduce token usage by 50-90% while maintaining accuracy in visual reasoning tasks. The method addresses reasoning path redundancy by decomposing visual reasoning into three cognitive functions and using FS-GRPO training to encourage efficient format selection.

0 favorites 0 likes
#visual-reasoning

Boosting Visual Instruction Tuning with Self-Supervised Guidance

Hugging Face Daily Papers ↗ · 2026-04-14 Cached

This paper proposes augmenting visual instruction tuning in multimodal language models with self-supervised tasks expressed as natural language instructions, improving vision-centric reasoning without additional architecture or annotations. By reformulating classical self-supervised pretext tasks as image-instruction-response triplets, the method achieves consistent performance improvements across multiple benchmarks by injecting only 3-10% visually grounded instructions into the training data.

0 favorites 0 likes
#visual-reasoning

A better method for planning complex visual tasks

MIT News — Artificial Intelligence ↗ · 2026-03-11 Cached

MIT researchers developed VLMFP, a two-stage generative AI approach combining vision-language models with formal planning software to achieve 70% success rate on complex visual planning tasks like robot navigation, nearly 2.3x better than existing baselines. The method automatically translates visual scenarios into planning files that classical solvers can process, enabling effective long-horizon planning in novel environments.

0 favorites 0 likes
#visual-reasoning

Thinking with images

OpenAI Blog ↗ · 2025-04-16 Cached

OpenAI releases o3 and o4-mini models that can reason with images in their chain-of-thought process, enabling visual understanding through native image manipulation tools like cropping and zooming without separate specialized models. These models achieve state-of-the-art performance on multimodal benchmarks including STEM questions, chart reading, and visual search tasks.

0 favorites 0 likes
← Previous
← Back to home

Submit Feedback