pdf-inspector
View on GitHubFast Rust library for PDF inspection, classification, and text extraction. Intelligently detects scanned vs text-based PDFs to enable smart routing decisions.
Rust library that classifies PDFs (text-based, scanned, image, mixed) in milliseconds and extracts position-aware text, tables, and clean Markdown, with optional per-page OCR routing. Ships Python, Node, WASM bindings and CLI tools, aimed at fast local document ingestion for LLM/RAG pipelines.
Use Cases
RAG document ingestion pipelinesPDF to Markdown conversion for LLM contextRoute scanned vs text-based PDFs to avoid OCR costPer-page selective OCR for mixed documentsExtract tables from financial and legal PDFsMulti-column and RTL reading-order reconstructionBrowser/Web Worker client-side PDF parsing via WASMBatch benchmark PDF parsers on a corpusInvoice, report, and research paper text extractionCLI batch conversion of PDF corpora to JSON/Markdown
Built With
- Language
- Rust
- Frameworks
- PyO3 · maturin · napi-rs · wasm-bindgen · ONNX Runtime · PDFium
Tags
pdf-parsing · text-extraction · markdown-conversion · ocr-routing · document-classification · table-detection · reading-order · rust · python-bindings · node-bindings · wasm · cli · document-ingestion · layout-analysis · cid-fonts