Vibe Coding Discover

RAG

pdf-inspector

View on GitHub

Fast Rust library for PDF inspection, classification, and text extraction. Intelligently detects scanned vs text-based PDFs to enable smart routing decisions.

★ 19K1,311 forksRustMITfirecrawl

Rust library that classifies PDFs (text-based, scanned, image, mixed) in milliseconds and extracts position-aware text, tables, and clean Markdown, with optional per-page OCR routing. Ships Python, Node, WASM bindings and CLI tools, aimed at fast local document ingestion for LLM/RAG pipelines.

Use Cases

RAG document ingestion pipelinesPDF to Markdown conversion for LLM contextRoute scanned vs text-based PDFs to avoid OCR costPer-page selective OCR for mixed documentsExtract tables from financial and legal PDFsMulti-column and RTL reading-order reconstructionBrowser/Web Worker client-side PDF parsing via WASMBatch benchmark PDF parsers on a corpusInvoice, report, and research paper text extractionCLI batch conversion of PDF corpora to JSON/Markdown

Built With

Language
Rust
Frameworks
PyO3 · maturin · napi-rs · wasm-bindgen · ONNX Runtime · PDFium

Tags

pdf-parsing · text-extraction · markdown-conversion · ocr-routing · document-classification · table-detection · reading-order · rust · python-bindings · node-bindings · wasm · cli · document-ingestion · layout-analysis · cid-fonts