ocrmypdf/OCRmyPDFπ₯ active
OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched
A set of tools for extracting tables from PDF files helping to do data mining on (OCR-processed) scanned documents.
$ git clone https://github.com/WZBSocialScienceCenter/pdftabextract.gitOCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched
Ready-to-use OCR with 80+ supported languages and all popular writing scripts including Latin, Chinese, Arabic, Devanagari, Cyrillic and etc.
PyMuPDF is a high performance Python library for data extraction, analysis, conversion & manipulation of PDF (and other) documents.
The official repo for βDolphin: Document Image Parsing via Heterogeneous Anchor Promptingβ, ACL, 2025.
Transforms PDF, Documents and Images into Enriched Structured Data
Open Source Virtual (Network) Printer for Windows that allows you to create PDFs, OCR text, and print images, with advanced features usually available only in enterprise solutions.
Data from GitHub Β· snapshot Sep 24, 2026