OCR & Documents
NanoNets/docext
An on-premises, OCR-free unstructured data extraction, markdown conversion and benchmarking toolkit. (https://idp-leaderboard.org/)
๐ง Natural Language Processing
PythonApache-2.0updated Mar 17, 2026
Fast and efficient unstructured data extraction. Written in Rust with bindings for many languages.
$ git clone https://github.com/yobix-ai/extractous.gitAn on-premises, OCR-free unstructured data extraction, markdown conversion and benchmarking toolkit. (https://idp-leaderboard.org/)
ExtractThinker is a Document Intelligence library for LLMs, offering ORM-style interaction for flexible and powerful document workflows.
Transforms PDF, Documents and Images into Enriched Structured Data
A list of open-source AI projects you can use to generate income easily.
An AI-powered Personal Identifiable Information (PII) scanner.
A python based library for NLP in Nepali language
Data from GitHub ยท snapshot Sep 24, 2026