๐Ÿ† #774 overall#120 of 313 in OCR & Documents

NanoNets /docstrange

Extract and convert data from any document, images, pdfs, word doc, ppt or URL into multiple formats (Markdown, JSON, CSV, HTML) with intelligent structured data extraction and advanced OCR.

$ git clone https://github.com/NanoNets/docstrange.git
GitHub social preview for NanoNets/docstrange
Stars
1.6K
1,571
Forks
139
139
Language
Python
License
MIT
Created
Jul 31, 2025
1.2 years old
Last push
Oct 31, 2025

Categories

GitHub topics

More in OCR & Documents

OCR & Documents

enoch3712/ExtractThinker

ExtractThinker is a Document Intelligence library for LLMs, offering ORM-style interaction for flexible and powerful document workflows.

๐Ÿง  Natural Language Processing
PythonApache-2.0updated Sep 16, 2026
GitHub โ†—โ˜… 1.6Kโ‘‚ 153
OCR & Documents

ibrahimqureshae/mdflux

Turn any document into clean, AI-ready Markdown. Local-first desktop app: reads scanned PDFs, batches folders, runs offline, and uses far fewer tokens than vision models.

PythonMITupdated Sep 7, 2026
GitHub โ†—โ˜… 424โ‘‚ 28
OCR & Documents

PaddlePaddle/PaddleOCR

Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.

PythonApache-2.0updated Sep 16, 2026
GitHub โ†—โ˜… 90.1Kโ‘‚ 11.4K
OCR & Documents

paperless-ngx/paperless-ngx๐Ÿ”ฅ active

A community-supported supercharged document management system: scan, index and archive all your documents

PythonGPL-3.0updated Sep 24, 2026
GitHub โ†—โ˜… 46Kโ‘‚ 3.2K
OCR & Documents

zcaceres/markdownify-mcp๐Ÿ”ฅ active

A Model Context Protocol server for converting almost anything to Markdown

TypeScriptMITupdated Sep 22, 2026
GitHub โ†—โ˜… 3Kโ‘‚ 255

Data from GitHub ยท snapshot Sep 24, 2026