Vibe Coding Discover

AI Tools

A very fast document classifier/splitter using Jev

★ 30823 forksPythonApache-2.0jerryjliu

DocJev is a Python library, CLI, and local web app that classifies PDFs/DOCX/PPTX into your own natural-language categories and splits multi-document packets into page ranges. It uses local LiteParse OCR plus hosted Jev (or an OpenAI baseline) decisions, with review flags, PDF export, and a benchmark harness.

Use Cases

Classify a PDF/DOCX/PPTX into one of user-defined natural-language categoriesSplit a multi-document PDF packet into ordered page-range segmentsExport each detected document segment as a separate PDF with hash verificationClassify a whole folder of documents into JSONL records with per-file error handlingOCR hard or scanned documents via LiteParse or LlamaParse tiersCompare LLM decision latency and accuracy against a small-model baselineRun a local web app for live classification/splitting demosFlag low-confidence boundaries, 'other' categories, and blank-page uncertainties for review

Built With

Language
Python
Frameworks
Typer · Pydantic · FastAPI · Uvicorn · pytest · Hatchling · pypdf · pypdfium2 · ReportLab · python-docx · python-pptx · hypothesis

Tags

document-classification · document-splitting · ocr · pdf · llm-inference · cli · python · llamaparse · document-ai · benchmarking · pdf-parsing · docx · pptx · fastapi · pydantic · page-boundaries