docling
View on GitHubGet your documents ready for gen AI
Docling is a Python library and CLI that parses PDFs, Office files, HTML, EPUB, images and audio into a unified DoclingDocument, exporting to Markdown, HTML or lossless JSON. It is widely used as the document ingestion stage for RAG and agentic AI pipelines.
Use Cases
Converting PDFs, DOCX, PPTX, XLSX, HTML and EPUB into structured Markdown/JSON for LLM pipelinesDocument ingestion and parsing stage for RAG and vector knowledge basesOCR of scanned PDFs and images with layout, reading order and table structureExtracting tables from documents into machine-readable formVLM-based page understanding using GraniteDocling or other visual language modelsAudio and video transcription via ASR for meeting/podcast contentParsing patents, JATS articles and XBRL financial reports into unified representationsLocal/air-gapped document processing for sensitive dataServing document conversion as an API or via an MCP server to agents
Built With
- Language
- Python
- Frameworks
- LangChain · LlamaIndex · CrewAI · Haystack · Pydantic · Hugging Face Transformers · PyTorch · Docling Core · MCP
Tags
document-parsing · pdf · ocr · markdown · table-extraction · document-conversion · docling-document · vlm · asr · docx · pptx · xlsx · html · json-export · chunking · rag-ingestion