OCR & Documents
NanoNets/docext
An on-premises, OCR-free unstructured data extraction, markdown conversion and benchmarking toolkit. (https://idp-leaderboard.org/)
๐ง Natural Language Processing
PythonApache-2.0updated Mar 17, 2026
An AI-powered Personal Identifiable Information (PII) scanner.
$ git clone https://github.com/redhuntlabs/Octopii.gitAn on-premises, OCR-free unstructured data extraction, markdown conversion and benchmarking toolkit. (https://idp-leaderboard.org/)
Fast and efficient unstructured data extraction. Written in Rust with bindings for many languages.
ExtractThinker is a Document Intelligence library for LLMs, offering ORM-style interaction for flexible and powerful document workflows.
OCR, Archive, Index and Search: Implementation agnostic OCR framework.
A python based library for NLP in Nepali language
Ready-to-use OCR with 80+ supported languages and all popular writing scripts including Latin, Chinese, Arabic, Devanagari, Cyrillic and etc.
Data from GitHub ยท snapshot Sep 24, 2026