OCR & Documents
eikek/docspell๐ฅ active
Assist in organizing your piles of documents, resulting from scanners, e-mails and other sources with miminal effort.
๐ง Natural Language Processing
ElmAGPL-3.0updated Sep 22, 2026
Transforms PDF, Documents and Images into Enriched Structured Data
$ git clone https://github.com/axa-group/Parsr.gitAssist in organizing your piles of documents, resulting from scanners, e-mails and other sources with miminal effort.
An on-premises, OCR-free unstructured data extraction, markdown conversion and benchmarking toolkit. (https://idp-leaderboard.org/)
๐ญ PDF text extraction pipeline: self-hosted, local-first, Docker-based
Fast and efficient unstructured data extraction. Written in Rust with bindings for many languages.
Collection of open-source libraries and tools for Robotic Process Automation (RPA), designed to be used with both Robot Framework and Python
OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched
Data from GitHub ยท snapshot Sep 24, 2026