OCR & Documents
pd3f/pd3f
π PDF text extraction pipeline: self-hosted, local-first, Docker-based
π€ Language Models
HTMLAGPL-3.0updated Oct 13, 2023
Sample applications and demos for Document AI, the end-to-end document processing platform on Google Cloud
$ git clone https://github.com/GoogleCloudPlatform/document-ai-samples.gitπ PDF text extraction pipeline: self-hosted, local-first, Docker-based
OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched
PyMuPDF is a high performance Python library for data extraction, analysis, conversion & manipulation of PDF (and other) documents.
The official repo for βDolphin: Document Image Parsing via Heterogeneous Anchor Promptingβ, ACL, 2025.
Transforms PDF, Documents and Images into Enriched Structured Data
A set of tools for extracting tables from PDF files helping to do data mining on (OCR-processed) scanned documents.
Data from GitHub Β· snapshot Sep 24, 2026