visual-grounding

Tag

Cards List
#visual-grounding

Evidence-Backed Video Question Answering

Hugging Face Daily Papers · 2026-07-13 Cached

This paper introduces Evidence-Backed Video Question Answering (E-VQA), a new task requiring models to output both semantic answers and precise spatio-temporal evidence like tracked object segmentation masklets. The authors create a human-verified benchmark and a scalable training dataset, showing significant improvements over baselines.

0 favorites 0 likes
#visual-grounding

GUICrafter: Weakly-Supervised GUI Agent Leveraging Massive Unannotated Screenshots

Hugging Face Daily Papers · 2026-06-29 Cached

GUICrafter introduces a weakly-supervised GUI agent that leverages massive unannotated screenshots and a two-stage curriculum learning framework to reduce reliance on expensive human annotations, achieving competitive performance with advanced systems like UI-TARS using only 0.1% of its data.

0 favorites 0 likes
#visual-grounding

@VincentLogic: NVIDIA open-sourced a visual grounding model: LocateAnything-3B. Dozens of minions densely piled together — it detects every single one without missing any, all boxed. The technological shift behind this is worth more than just saying 'more accurate'.

X AI KOLs Timeline · 2026-06-26 Cached

NVIDIA has open-sourced the visual grounding model LocateAnything-3B, which can accurately detect and bound all target objects in dense scenes.

0 favorites 0 likes
#visual-grounding

@DataChaz: @NVIDIA just dropped LocateAnything, making object detection ~10x faster by fixing one core bottleneck: How the model w…

X AI KOLs Following · 2026-06-17 Cached

NVIDIA released LocateAnything, an open-source model that achieves ~10x faster object detection by predicting all coordinates simultaneously instead of sequentially, reaching 12.7 FPS on a single H100 and outperforming 32B parameter models.

0 favorites 0 likes
#visual-grounding

iVGR: Internalizing Visually Grounded Reasoning for MLLMs with Reinforcement Learning

Hugging Face Daily Papers · 2026-05-29 Cached

Introduces iVGR, a reinforcement learning framework that internalizes visual localization into textual reasoning for multimodal language models, eliminating the need for explicit visual grounding during inference while improving fine-grained perception performance.

0 favorites 0 likes
#visual-grounding

@ZhidingYu: Thank you NVIDIA! I will be presenting LocateAnything at #CVPR2026 at the NVIDIA Booth: June 5 4:20 - 4:40 pm MDT (Frid…

X AI KOLs Following · 2026-05-28 Cached

NVIDIA introduces LocateAnything, a unified generative grounding and detection framework that uses Parallel Box Decoding to improve decoding throughput and localization accuracy. This work will be presented at CVPR 2026.

0 favorites 0 likes
#visual-grounding

LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding

Hugging Face Daily Papers · 2026-05-26 Cached

LocateAnything proposes Parallel Box Decoding for unified visual grounding and object detection, decoding geometric elements as atomic units to improve throughput and localization accuracy, supported by a large-scale dataset of 138M samples.

0 favorites 0 likes
#visual-grounding

Transcoders Trace Visual Grounding and Hallucinations in Vision-Language Models

arXiv cs.LG · 2026-05-25 Cached

This paper presents a function-centric framework using Transcoders to trace computational pathways in vision-language models, demonstrating stronger attribution of visual grounding and the ability to predict hallucinations via graph-based features.

0 favorites 0 likes
#visual-grounding

ForMaT: Dataset for Visually-Grounded Multilingual PDF Translation

arXiv cs.CL · 2026-05-18 Cached

This paper introduces ForMaT, a parallel corpus of 3,956 PDFs across 15 language pairs designed for visually-grounded multilingual translation, preserving layout metadata to benchmark layout-aware MT systems.

0 favorites 0 likes
#visual-grounding

MementoGUI: Learning Agentic Multimodal Memory Control for Long-Horizon GUI Agents

Hugging Face Daily Papers · 2026-05-18 Cached

MementoGUI introduces a plug-in agentic memory framework for GUI agents that uses learned controllers for selective memory management and retrieval, improving performance on long-horizon tasks with compressed visual and textual representations.

0 favorites 0 likes
#visual-grounding

Towards Visually-Guided Movie Subtitle Translation for Indic Languages

arXiv cs.CL · 2026-05-13 Cached

This paper presents a case study on visually-guided movie subtitle translation for low-resource Indic languages, demonstrating that selective visual grounding improves translation quality while addressing temporal misalignment challenges.

0 favorites 0 likes
#visual-grounding

SeePhys Pro: Diagnosing Modality Transfer and Blind-Training Effects in Multimodal RLVR for Physics Reasoning

Hugging Face Daily Papers · 2026-05-10 Cached

The paper introduces SeePhys Pro, a benchmark to diagnose modality transfer issues in multimodal RL for physics reasoning, revealing that models struggle with representation-invariant reasoning and often rely on residual textual cues rather than visual evidence.

0 favorites 0 likes
#visual-grounding

HyperEyes: Dual-Grained Efficiency-Aware Reinforcement Learning for Parallel Multimodal Search Agents

Hugging Face Daily Papers · 2026-05-08 Cached

HyperEyes is a parallel multimodal search agent that uses dual-grained reinforcement learning to optimize inference efficiency, achieving higher accuracy with significantly fewer tool-call rounds compared to existing agents.

0 favorites 0 likes
← Back to home

Submit Feedback