๐Ÿ† #607 overall#101 of 313 in OCR & Documents

NanoNets /docext

An on-premises, OCR-free unstructured data extraction, markdown conversion and benchmarking toolkit. (https://idp-leaderboard.org/)

$ git clone https://github.com/NanoNets/docext.git
GitHub social preview for NanoNets/docext
Stars
2.1K
2,086
Forks
156
156
Language
Python
License
Apache-2.0
Created
Mar 25, 2025
1.5 years old
Last push
Mar 17, 2026

Categories

GitHub topics

More in OCR & Documents

OCR & Documents

axa-group/Parsr

Transforms PDF, Documents and Images into Enriched Structured Data

๐Ÿง  Natural Language Processing
JavaScriptApache-2.0updated Mar 20, 2026
GitHub โ†—โ˜… 6.2Kโ‘‚ 316
OCR & Documents

yobix-ai/extractous

Fast and efficient unstructured data extraction. Written in Rust with bindings for many languages.

๐Ÿง  Natural Language Processing
RustApache-2.0updated Dec 21, 2024
GitHub โ†—โ˜… 1.8Kโ‘‚ 96
OCR & Documents

eikek/docspell๐Ÿ”ฅ active

Assist in organizing your piles of documents, resulting from scanners, e-mails and other sources with miminal effort.

๐Ÿง  Natural Language Processing
ElmAGPL-3.0updated Sep 22, 2026
GitHub โ†—โ˜… 2.3Kโ‘‚ 184
OCR & Documents

enoch3712/ExtractThinker

ExtractThinker is a Document Intelligence library for LLMs, offering ORM-style interaction for flexible and powerful document workflows.

๐Ÿง  Natural Language Processing
PythonApache-2.0updated Sep 16, 2026
GitHub โ†—โ˜… 1.6Kโ‘‚ 153
OCR & Documents

redhuntlabs/Octopii

An AI-powered Personal Identifiable Information (PII) scanner.

๐Ÿง  Natural Language Processing
PythonOtherupdated Jan 22, 2025
GitHub โ†—โ˜… 745โ‘‚ 64
OCR & Documents

sushil79g/Nepali_nlp

A python based library for NLP in Nepali language

๐Ÿ“ Text Summarization๐Ÿง  Natural Language Processing
Jupyter NotebookMITupdated May 22, 2023
GitHub โ†—โ˜… 171โ‘‚ 55

Data from GitHub ยท snapshot Sep 24, 2026