ocrmypdf/OCRmyPDFπ₯ active
OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched
Extract structured data from documents quickly and accurately.
$ git clone https://github.com/datalab-to/lift.gitOCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched
PyMuPDF is a high performance Python library for data extraction, analysis, conversion & manipulation of PDF (and other) documents.
The official repo for βDolphin: Document Image Parsing via Heterogeneous Anchor Promptingβ, ACL, 2025.
Transforms PDF, Documents and Images into Enriched Structured Data
A curated collection of practical AI projects implementing OCR systems, RAG, AI agents, and other AI use cases.
Document (PDF, Word, PPTX ...) extraction and parse API using state of the art modern OCRs + Ollama supported models. Anonymize documents. Remove PII. Convert any document or picture to structured JSON or Markdown
Data from GitHub Β· snapshot Sep 24, 2026