knowhere
View on GitHubKnowhere extracts, parses, and outputs structured chunks ready for AI Agents and RAG.
Knowhere is an open-source document parsing and retrieval backend that converts PDFs, Office files, images, and text into hierarchy-native chunks with citations and cross-document links. It serves this memory to AI agents and RAG pipelines, including via MCP.
Use Cases
Turn dirty PDFs and slide decks into structured chunks for RAGFeed AI agents persistent, navigable document memoryAgentic RAG with hierarchy and cross-document graph navigationParse ultra-long PDFs (300-500+ pages) with page-grounded citationsExtract and link tables and images to source sections via VLM summariesExpose corpus retrieval to Claude Code, Cursor, or Codex via MCPBuild offline/local document knowledge bases for agentsDeterministic top-K vector retrieval over parsed sections
Built With
- Language
- Python
- Frameworks
- LangChain · MCP · Docker · FastAPI
Tags
rag · agentic-rag · document-parsing · document-intelligence · chunking · vision-model · ocr · knowledge-base · retrieval · citations · pdf · pptx · mcp · vector-database · memory · long-context