marker
View on GitHubConvert PDF to markdown + JSON quickly with high accuracy
Marker is a Python document-intelligence tool that converts PDFs, images, PPTX, DOCX, XLSX, HTML and EPUB into markdown, JSON, HTML or chunks. It runs local Surya VLM inference (vLLM/llama.cpp) for layout, OCR and tables, with optional LLM boosting for higher accuracy.
Use Cases
Convert PDFs/images/Office docs to markdown, JSON, HTML or chunksPrepare clean text corpora for RAG pipelinesOCR scanned documents in many languagesExtract and format tables, forms and equationsConvert inline math to LaTeXBatch-convert large document sets on GPU or CPUImprove extraction accuracy with optional LLM post-processingExtract embedded images and strip headers/footers
Built With
- Language
- Python
- Frameworks
- PyTorch · transformers · surya-ocr · pdftext · FastAPI · Streamlit · vLLM · llama.cpp · Pydantic · click
Tags
pdf · ocr · markdown · document-intelligence · document-conversion · vlm · vision-model · table-extraction · latex · chunking · json-output · batch-processing · cpu-gpu-mps · rag-data-prep