@akshay_pachaar: Web scraping will never be the same. (100% open-source visual search at scale) PixelRAG is a retrieval system that skip…
Summary
PixelRAG is an open-source retrieval system that bypasses HTML parsing by screenshotting web pages and using vision-language models to read answers directly from pixels, claiming significant accuracy improvements over text-based RAG.
View Cached Full Text
Cached at: 06/20/26, 02:36 PM
Web scraping will never be the same.
(100% open-source visual search at scale)
PixelRAG is a retrieval system that skips HTML parsing completely.
Instead of scraping a page into text and embedding chunks, it screenshots the page and retrieves the image. A vision-language model reads the answer straight off the pixels.
Why that matters: parsing is where web RAG quietly loses information.
- A single HTML-to-text parser can drop 40%+ of a page.
- Tables, charts, and layout get flattened or thrown out.
- Swapping parsers alone can move accuracy ~10 points on the same docs.
PixelRAG indexes the page a person actually sees. The team built a visual index of all of Wikipedia, 30M+ screenshots, and it still beats the strongest text RAG baseline by 18.1% on text-only QA.
The repo also ships a Claude Code plugin that gives Claude eyes.
It lets Claude screenshot any URL and read the rendered page instead of scraping the DOM. So you can hand it a live page, an arXiv paper, or your local site and ask what it actually looks like.
One setup script. No MCP server, no backend.
How the pipeline works:
- Renders each document (web, PDF, image) to image tiles.
- Embeds them with Qwen3-VL-Embedding, LoRA fine-tuned on screenshots.
- Builds a FAISS index and serves a search API.
A stronger reader model lifts accuracy with no re-indexing, since the index is just pixels.
Everything is open-source under Apache-2.0.
GitHub repo: https://github.com/StarTrail-org/PixelRAG…
Talking about RAG, I recently wrote an article on a new approach that makes retrieval much more efficient by cutting corpus size by 40x, reducing tokens per query by 3x, and improving vector search relevance by 2.3x.
The article is quoted below.
StarTrail-org/PixelRAG
Source: https://github.com/StarTrail-org/PixelRAG
Official codebase for PIXELRAG: Web Screenshots Beat Text for Retrieval-Augmented Generation
Yichuan Wang*,
Zhifei Li*,
Zirui Wang,
Paul Teiletche,
Lesheng Jin
Matei Zaharia†,
Joseph E. Gonzalez†,
Sewon Min†
* Equal contribution † Equal advising
Work done at Berkeley SkyLab & BAIR & Berkeley NLP
Search any document by how it looks, not just the text it contains.
What it is · Give Claude eyes · How it works · Pipelines
pip install pixelrag
The two core operations — render a page to screenshots, search a visual index:
# Render any page or document to screenshot tiles
pixelshot https://en.wikipedia.org/wiki/Python --output ./tiles
# Search a hosted index of 8.28M Wikipedia pages — no setup, runs against the live API
curl -X POST https://api.pixelrag.ai/search \
-H "Content-Type: application/json" \
-d '{"queries": [{"text": "What is the capital of France?"}], "n_docs": 5}'
Live, hosted endpoint —
https://api.pixelrag.aiserves a pre-built index of 8.28M Wikipedia pages. No setup, no API key. It even takes an image as the query (visual search) — see the API reference →.
Or try it in the browser at pixelrag.ai, or run the demo notebook in
Colab — it
renders a page and searches the hosted index, with the images inline.
What it is
PixelRAG renders documents — web pages, PDFs, images — as screenshots and retrieves over the images directly. Visual structure that HTML parsing throws away — tables, charts, layout, infographics — stays intact, so the reader model can actually answer questions about it. Wikipedia’s 8.28M articles ship as a pre-built index; the pipeline itself is general-purpose.
Give Claude eyes
The renderer also ships as a Claude Code plugin — the pixelbrowse skill. Instead of fetching
raw HTML, Claude screenshots a page with pixelshot and reads the image, so it sees
charts, diagrams, tables, and layout the way a person does.
Install it — no clone needed (pixelshot comes from pip install pixelrag):
pip install pixelrag # provides the pixelshot command
claude plugin marketplace add StarTrail-org/PixelRAG
claude plugin install pixelbrowse@pixelrag-plugins
Then just ask Claude to look at a page:
claude -p "screenshot https://news.ycombinator.com and summarize the top stories"
claude -p "screenshot https://arxiv.org/abs/2404.12387 and explain the key findings"
Or use the slash command in an interactive session: /screenshot https://example.com.
No MCP server, no backend: the skill just calls pixelshot (Playwright/CDP) on your machine.
How it works
Text-based RAG parses the page to text chunks and loses the table — the reader can’t find the answer. PixelRAG renders the page to screenshot tiles, retrieves the right tile, and the reader reads the number straight off the image.
Two pieces make this work: (1) rendering documents to images instead of parsing them to text, and
(2) a Qwen3-VL-Embedding model, LoRA-fine-tuned on screenshot data, that embeds page images into
a space where visual content is retrievable.
Pipelines
Capture is the standalone pixelshot command; the rest of the pipeline runs through the
pixelrag umbrella — pixelrag <stage>. Install only the stages you need:
| Command | What it does | Install |
|---|---|---|
pixelshot | Document → image tiles (Playwright CDP, PDF) | pip install pixelrag |
pixelrag chunk · embed · build-index | Tiles → vectors → FAISS index | pip install 'pixelrag[embed]' |
pixelrag index | Orchestrates the full pipeline: source → ingest → embed → index | pip install 'pixelrag[index]' |
pixelrag serve | FAISS search API (FastAPI, CPU or GPU) | pip install 'pixelrag[serve]' |
render ←── index ──→ embed serve (independent) train → serve (HTTP)
train is a separate uv project with its own pinned env (torch==2.9.1+cu129,
transformers==4.57.1, cuDNN 9.20) — install it from inside train/, not from the root.
Search a pre-built index
pip install 'pixelrag[serve]'
# Download a pre-built index from Hugging Face. The dataset repo holds four FAISS indexes
# (base/LoRA Wikipedia pixel, Wikipedia text, news pixel); grab just the base one (~217G) here.
huggingface-cli download StarTrail-org/pixelrag-faiss-indexes \
--repo-type dataset --include "search_index_normed_v2/*" --local-dir ./index
# Serve, then query
pixelrag serve --index-dir ./index/search_index_normed_v2 --port 30001
curl -X POST http://localhost:30001/search \
-H "Content-Type: application/json" \
-d '{"queries": [{"text": "What is the capital of France?"}], "n_docs": 5}'
Build an index from your own documents
pip install 'pixelrag[index]'
# Create pixelrag.yaml
cat > pixelrag.yaml << 'EOF'
source:
type: local
path: ./my_docs
embed:
model: Qwen/Qwen3-VL-Embedding-2B
device: cuda
gpu_ids: [0]
output: ./my_index
EOF
# Build, then serve
pixelrag index build
pixelrag serve --index-dir ./my_index --port 30001
Render a page programmatically
from pixelrag_render import render_url
# render a single page to tiles — e.g. for an agent to read
tiles = render_url("https://en.wikipedia.org/wiki/Python", "./tiles")
Embed tools (standalone)
Each stage runs independently, without the orchestrator:
pip install 'pixelrag[embed]'
pixelrag chunk --tiles-dir ./tiles
pixelrag embed --shard-dir ./tiles --output-dir ./embeddings --gpu-ids 0,1
pixelrag build-index --embeddings-dir ./embeddings --output-dir ./index
Training
Fine-tuning lives in train/ — a separate uv project (wiki-screenshot-training) with its own
pinned env. It LoRA-fine-tunes Qwen/Qwen3-VL-Embedding-2B for webpage retrieval; run it from
inside train/ (cd train && uv sync). See train/README.md for the full recipe.
You don’t need to retrain to use the model — the trained adapters are published at
Chrisyichuan/wiki-screenshot-embedding-lora.
We also release the full training set
(Chrisyichuan/screenshot-training-natural-filtered-v2),
so you can adapt other backbones yourself — a larger Qwen, or any other embedding model.
Data Curation
Visualization of some very early version of the training data: early training data viewer
Reproduce: TBD
Acknowledgments
Thanks to Rulin Shao for support.
Thanks also to Claude Code and OpenAI Codex for supporting open-source contributors with credits and plans, which we earned by working on LEANN.
This work is done by the Berkeley Sky Computing Lab, BAIR, and the Berkeley NLP Group.
License
Apache-2.0
Similar Articles
@RoundtableSpace: Web scraping is dead. PixelRAG skips HTML parsing completely. It screenshots the page and a vision model reads the answ…
PixelRAG is an open-source tool that replaces traditional web scraping by using screenshots and a vision model to extract data from web pages. It includes a plugin for Claude Code.
@GYLQ520: PixelRAG is an open-source project from UC Berkeley SkyLab and other teams, ready to play. Its approach is straightforward, no longer relying on HTML or text parsing. Instead, it directly captures web pages and PDFs as screenshots and uses visual indexing for RAG retrieval. Everything that HTML parsing loses, such as tables, charts, and layouts, is preserved…
PixelRAG is an open-source project from UC Berkeley SkyLab and other teams. By capturing web pages and PDFs as screenshots and using visual indexing for retrieval, it improves RAG accuracy and significantly reduces token costs for AI Agents.
@LTChives: Web scraping is dead. This PixelRAG in the video completely bypasses HTML parsing. It takes a screenshot of the webpage and then lets the vision model read answers from the pixels. Previously, AI reading a webpage meant first parsing the code, extracting text, and splitting paragraphs. Now it just looks at the page. 100% open source, plus it comes with Claude Code…
PixelRAG is a novel open-source tool that bypasses traditional HTML parsing by directly taking screenshots of webpages and using vision models to extract answers from the pixels. It also supports the Claude Code plugin, giving Claude visual capabilities.
@oliviscusAI: The entire RAG industry is about to get cooked. Researchers developed a new RAG approach that bypasses almost everythin…
Researchers introduced PageIndex, a novel RAG approach that eliminates reliance on vector databases, embeddings, and chunking, achieving 98.7% accuracy on FinanceBench and outperforming existing methods while being free and open source.
How we index images for RAG
Kapa.ai describes their approach to indexing images for RAG by using a cheap vision model to generate text descriptions at indexing time, avoiding query-time vision costs, resulting in better answers with minimal per-query overhead.