unstructured
View on GitHubConvert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for language models. Visit our website to learn more about our enterprise grade Platform product for production grade workflows, partitioning, enrichments, chunking and embedding.
Python library that partitions 60+ document types (PDF, HTML, DOCX, email, images) into structured elements for LLM and RAG pipelines, with chunking, enrichment, and an MCP server for agent workflows.
Use Cases
RAG data ingestion from PDFs and office docsConvert PDF/HTML/DOCX to structured JSON for LLMsChunk and embed documents for vector storesOCR scanned documents and imagesParse emails (.eml/.msg) and attachmentsExtract tables and layout elements from documentsPrepare training/eval corpora for language modelsAgent-driven document processing via MCP serverBatch ETL preprocessing pipelines for ML
Built With
- Language
- HTML
- Frameworks
- LangChain · spaCy · Unstructured MCP · unstructured-client · pdfminer.six · pypdf · python-docx
Tags
document-parsing · etl · rag · ocr · pdf · chunking · embeddings · mcp · data-ingestion · nlp · preprocessing · pdf-to-json · langchain · vector-database · document-conversion · unstructured-data