OCR & Documents
pymupdf/PyMuPDFπ₯ active
PyMuPDF is a high performance Python library for data extraction, analysis, conversion & manipulation of PDF (and other) documents.
PythonAGPL-3.0updated Sep 24, 2026
Demos, examples and utilities using PyMuPDF
$ git clone https://github.com/pymupdf/PyMuPDF-Utilities.gitPyMuPDF is a high performance Python library for data extraction, analysis, conversion & manipulation of PDF (and other) documents.
OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched
The official repo for βDolphin: Document Image Parsing via Heterogeneous Anchor Promptingβ, ACL, 2025.
Transforms PDF, Documents and Images into Enriched Structured Data
A set of tools for extracting tables from PDF files helping to do data mining on (OCR-processed) scanned documents.
Extract structured data from documents quickly and accurately.
Data from GitHub Β· snapshot Sep 24, 2026