Twitter/X

PixelRAG, from UC Berkeley researchers, replaces HTML-to-text scraping with page…

Brief

PixelRAG, from UC Berkeley researchers, replaces HTML-to-text scraping with page screenshots that a vision-language model reads, avoiding parser-induced data loss (single-parser drops >40%, parser swaps change accuracy ~10 points). It outperforms text RAG by 18.1%, was scaled to 8.28M Wikipedia articles (30M+ screenshots), and is Apache-2.0 on GitHub with a 'pixelbrowse' Claude Code plugin.

Source evidence

STOP PARSING HTML FOR RAG. JUST SCREENSHOT IT 🔥

Researchers from UC Berkeley just released PixelRAG, an open-source system that skips HTML parsing entirely.

Why is it changing web scraping for good?

Well, instead of scraping a page into text and embedding chunks:

1 it screenshots that page and retrieves the image

2 a vision-language model then reads the answer straight off the pixels

This solves a real problem: parsing is where RAG quietly loses context:
→ A single HTML-to-text parser can drop 40%+ of a page
→ Tables, charts, and complex layouts get flattened or deleted
→ Changing parsers alone can swing accuracy by 10 points

By indexing the page exactly as a human sees it, PixelRAG beats the strongest text RAG baseline by 18.1% on text-only QA.

To prove it scales, the team built a visual index of all 8.28M Wikipedia articles using 30M+ screenshots.

It even includes a Claude Code plugin (pixelbrowse) that gives Claude native "eyes" to read live pages locally, how cool is that?

100% free and open-source (Apache-2.0).

Here's the repo → github.com/StarTrail-org/Pix…

BONUS: If you're building retrieval systems, my friend @akshay_pachaar recently wrote a great guide on slashing your RAG corpus by 40x and tokens by 3x.

Check it out below ↓

Akshay 🚀 (@akshay_pachaar)

— https://nitter.net/akshay_pachaar/status/2052743644411765230#m