๐Ÿ† #46 overall#10 of 313 in OCR & Documents๐Ÿ”ฅ active this week

opendataloader-project /opendataloader-pdf

PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.

$ git clone https://github.com/opendataloader-project/opendataloader-pdf.git
GitHub social preview for opendataloader-project/opendataloader-pdf
Stars
29.4K
29,363
Forks
2.8K
2,797
Language
Java
License
Apache-2.0
Created
May 13, 2025
1.4 years old
Last push
Sep 23, 2026
๐Ÿ”ฅ this week

Categories

GitHub topics

More in OCR & Documents

OCR & Documents

NanoNets/docstrange

Extract and convert data from any document, images, pdfs, word doc, ppt or URL into multiple formats (Markdown, JSON, CSV, HTML) with intelligent structured data extraction and advanced OCR.

PythonMITupdated Oct 31, 2025
GitHub โ†—โ˜… 1.6Kโ‘‚ 139
OCR & Documents

zcaceres/markdownify-mcp๐Ÿ”ฅ active

A Model Context Protocol server for converting almost anything to Markdown

TypeScriptMITupdated Sep 22, 2026
GitHub โ†—โ˜… 3Kโ‘‚ 255
OCR & Documents

enoch3712/ExtractThinker

ExtractThinker is a Document Intelligence library for LLMs, offering ORM-style interaction for flexible and powerful document workflows.

๐Ÿง  Natural Language Processing
PythonApache-2.0updated Sep 16, 2026
GitHub โ†—โ˜… 1.6Kโ‘‚ 153
OCR & Documents

PaddlePaddle/PaddleOCR

Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.

PythonApache-2.0updated Sep 16, 2026
GitHub โ†—โ˜… 90.1Kโ‘‚ 11.4K
OCR & Documents

paperless-ngx/paperless-ngx๐Ÿ”ฅ active

A community-supported supercharged document management system: scan, index and archive all your documents

PythonGPL-3.0updated Sep 24, 2026
GitHub โ†—โ˜… 46Kโ‘‚ 3.2K
OCR & Documents

datalab-to/chandra

OCR model that handles complex tables, forms, handwriting with full layout.

PythonApache-2.0updated Jun 26, 2026
GitHub โ†—โ˜… 12.3Kโ‘‚ 1.2K

Data from GitHub ยท snapshot Sep 24, 2026