OCR & Documents
axa-group/Parsr
Transforms PDF, Documents and Images into Enriched Structured Data
๐ง Natural Language Processing
JavaScriptApache-2.0updated Mar 20, 2026
An on-premises, OCR-free unstructured data extraction, markdown conversion and benchmarking toolkit. (https://idp-leaderboard.org/)
$ git clone https://github.com/NanoNets/docext.gitTransforms PDF, Documents and Images into Enriched Structured Data
Fast and efficient unstructured data extraction. Written in Rust with bindings for many languages.
Assist in organizing your piles of documents, resulting from scanners, e-mails and other sources with miminal effort.
ExtractThinker is a Document Intelligence library for LLMs, offering ORM-style interaction for flexible and powerful document workflows.
An AI-powered Personal Identifiable Information (PII) scanner.
A python based library for NLP in Nepali language
Data from GitHub ยท snapshot Sep 24, 2026