PixelRAG
View on GitHubhttps://arxiv.org/abs/2606.28344. The end of web parsing. The beginning of scalable pixel-native search. link: https://pixelrag.ai/
PixelRAG renders web pages and PDFs to screenshot tiles and retrieves over the images with a LoRA-tuned Qwen3-VL embedding model, so tables, charts and layout survive retrieval. Includes a pixelshot CLI, FAISS/Qdrant indexing, a FastAPI search server, a hosted 8.28M-page Wikipedia index, and a Claude Code screenshot sk
Use Cases
Search documents by visual appearance instead of parsed textRetrieve table/chart content that HTML-to-text parsing destroysBuild a FAISS or Qdrant index over screenshots of your own PDFs and pagesQuery an 8.28M-page Wikipedia screenshot index via hosted APIGive Claude Code eyes via the pixelbrowse screenshot skillRender web pages or PDFs to image tiles with the pixelshot CLIServe a local visual search API with pixelrag serveFine-tune a VLM embedding model on screenshot dataEvaluate visual retrieval on LiveVQA / Monaco benchmarksVisual search where the query itself is an image
Built With
- Language
- Python
- Frameworks
- PyTorch · Hugging Face Transformers · FAISS · FastAPI · Qdrant · Playwright · Chromium CDP · PyMuPDF · Uvicorn · Pydantic · NumPy · uv · pillow
Tags
visual-rag · pixel-native-search · screenshot-retrieval · multimodal · embeddings · faiss · vector-search · qwen3-vl · vlm · document-rendering · pdf-search · web-screenshots · agent-memory · claude-code-plugin · qdrant · searchengine