opendataloader-pdf
View on GitHubPDF Parser for AI-ready data. Automate PDF accessibility. Open-source.
Java-based PDF parser with Python and Node.js SDKs. Extracts Markdown, JSON with bounding boxes, and HTML; hybrid mode supports OCR and complex layouts, and it can auto-tag PDFs for accessibility.
Use Cases
Prepare PDF content for RAG pipelinesExtract structured text and tables from PDFsProcess scanned PDFs with OCRConvert PDFs to Markdown, JSON, or HTMLAuto-tag PDFs for accessibility
Built With
- Language
- Java
- Frameworks
- LangChain
Tags
PDF parsing · document extraction · RAG · OCR · accessibility · Markdown · JSON · HTML · bounding boxes · table extraction · layout analysis