pdf-inspector
★ 19KRust library that classifies PDFs (text-based, scanned, image, mixed) in milliseconds and extracts position-aware text, tables, and clean Markdown, with optional per-page OCR routing. Ships Python, Node, WASM bindings and CLI tools, aimed at fast local document ingestion for LLM/RAG pipelines.
RAG | Rust · pdf-parsing · text-extraction
View Project →