Vibe Coding Discover

RAG

unstructured

View on GitHub

Convert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for language models. Visit our website to learn more about our enterprise grade Platform product for production grade workflows, partitioning, enrichments, chunking and embedding.

★ 15K1,332 forksHTMLApache-2.0Unstructured-IO

Python library that partitions 60+ document types (PDF, HTML, DOCX, email, images) into structured elements for LLM and RAG pipelines, with chunking, enrichment, and an MCP server for agent workflows.

Use Cases

RAG data ingestion from PDFs and office docsConvert PDF/HTML/DOCX to structured JSON for LLMsChunk and embed documents for vector storesOCR scanned documents and imagesParse emails (.eml/.msg) and attachmentsExtract tables and layout elements from documentsPrepare training/eval corpora for language modelsAgent-driven document processing via MCP serverBatch ETL preprocessing pipelines for ML

Built With

Language
HTML
Frameworks
LangChain · spaCy · Unstructured MCP · unstructured-client · pdfminer.six · pypdf · python-docx

Tags

document-parsing · etl · rag · ocr · pdf · chunking · embeddings · mcp · data-ingestion · nlp · preprocessing · pdf-to-json · langchain · vector-database · document-conversion · unstructured-data