visual-rag

Tag

Cards List
#visual-rag

@GYLQ520: PixelRAG is an open-source project from UC Berkeley SkyLab and other teams, ready to play. Its approach is straightforward, no longer relying on HTML or text parsing. Instead, it directly captures web pages and PDFs as screenshots and uses visual indexing for RAG retrieval. Everything that HTML parsing loses, such as tables, charts, and layouts, is preserved…

X AI KOLs Timeline ↗ · 2026-08-21 Cached

PixelRAG is an open-source project from UC Berkeley SkyLab and other teams. By capturing web pages and PDFs as screenshots and using visual indexing for retrieval, it improves RAG accuracy and significantly reduces token costs for AI Agents.

0 favorites 0 likes
#visual-rag

Does More Retrieved Evidence Help Visual Retrieval-Augmented Generation with Diffusion Language Models?

arXiv cs.CL ↗ · 2026-08-10 Cached

The paper investigates whether retrieving more evidence helps visual retrieval-augmented generation with diffusion language models, finding that unconditionally expanding evidence hurts accuracy due to semantic conflict, and proposes a training-free Entropy-Based Candidate Filter (ECF) to selectively admit evidence, improving accuracy across benchmarks.

0 favorites 0 likes
#visual-rag

UniDoc-RL: Coarse-to-Fine Visual RAG with Hierarchical Actions and Dense Rewards

Hugging Face Daily Papers ↗ · 2026-04-16 Cached

UniDoc-RL presents a reinforcement learning framework for Large Vision-Language Models that optimizes retrieval, reranking, and visual reasoning through hierarchical decision-making and dense multi-reward supervision, achieving up to 17.7% improvements over prior RL-based methods on visual RAG tasks.

0 favorites 0 likes
← Back to home

Submit Feedback