WZBSocialScienceCenter/pdftabextract
A set of tools for extracting tables from PDF files helping to do data mining on (OCR-processed) scanned documents.
OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched
$ git clone https://github.com/ocrmypdf/OCRmyPDF.gitA set of tools for extracting tables from PDF files helping to do data mining on (OCR-processed) scanned documents.
PyMuPDF is a high performance Python library for data extraction, analysis, conversion & manipulation of PDF (and other) documents.
The official repo for βDolphin: Document Image Parsing via Heterogeneous Anchor Promptingβ, ACL, 2025.
Transforms PDF, Documents and Images into Enriched Structured Data
Open Source Virtual (Network) Printer for Windows that allows you to create PDFs, OCR text, and print images, with advanced features usually available only in enterprise solutions.
Data from GitHub Β· snapshot Sep 24, 2026