scribe.js
View on GitHubJavaScript OCR and text extraction for images and PDFs.
JavaScript OCR and text-extraction library for images and PDFs, powered by WebAssembly Tesseract with no build step. Runs in browser or Node, outputs searchable PDFs, and ships a CLI plus an MCP server for agent/tool integration.
Use Cases
Recognize text from imagesExtract text from user-uploaded PDFsAdd searchable invisible text layers to scanned PDFsRun OCR fully client-side in the browserBatch OCR via CLIDocument ingestion for RAG pipelinesExpose OCR as an MCP tool for AI agentsTable and form-field extraction from documentsExtract text from text-native PDFs without OCRFill and sign PDF forms programmatically
Built With
- Language
- JavaScript
- Frameworks
- Node.js · Tesseract · Vite · Vitest · WebdriverIO · ESLint · Express · WebAssembly
Tags
ocr · text-extraction · pdf · tesseract · webassembly · document-ai · searchable-pdf · mcp · cli · browser · nodejs · image-processing · table-extraction · form-fields · javascript-library