๐Ÿ† #760 overall#118 of 313 in OCR & Documents

enoch3712 /ExtractThinker

ExtractThinker is a Document Intelligence library for LLMs, offering ORM-style interaction for flexible and powerful document workflows.

$ git clone https://github.com/enoch3712/ExtractThinker.git
GitHub social preview for enoch3712/ExtractThinker
Stars
1.6K
1,598
Forks
153
153
Language
Python
License
Apache-2.0
Created
Feb 1, 2024
2.6 years old
Last push
Sep 16, 2026

Categories

GitHub topics

More in OCR & Documents

OCR & Documents

yobix-ai/extractous

Fast and efficient unstructured data extraction. Written in Rust with bindings for many languages.

๐Ÿง  Natural Language Processing
RustApache-2.0updated Dec 21, 2024
GitHub โ†—โ˜… 1.8Kโ‘‚ 96
OCR & Documents

garylab/MakeMoneyWithAI๐Ÿ”ฅ active

A list of open-source AI projects you can use to generate income easily.

๐Ÿง  Natural Language Processing
Pythonno licenseupdated Sep 24, 2026
GitHub โ†—โ˜… 948โ‘‚ 148
OCR & Documents

paperless-ngx/paperless-ngx๐Ÿ”ฅ active

A community-supported supercharged document management system: scan, index and archive all your documents

PythonGPL-3.0updated Sep 24, 2026
GitHub โ†—โ˜… 46Kโ‘‚ 3.2K
OCR & Documents

NanoNets/docstrange

Extract and convert data from any document, images, pdfs, word doc, ppt or URL into multiple formats (Markdown, JSON, CSV, HTML) with intelligent structured data extraction and advanced OCR.

PythonMITupdated Oct 31, 2025
GitHub โ†—โ˜… 1.6Kโ‘‚ 139
OCR & Documents

run-llama/ParseBench๐Ÿ”ฅ active

ParseBench - A Document Parsing Benchmark for AI Agents

PythonApache-2.0updated Sep 23, 2026
GitHub โ†—โ˜… 592โ‘‚ 109
OCR & Documents

Unstructured-IO/unstructured๐Ÿ”ฅ active

Convert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for language models. Visit our website to learn more about our enterprise grade Platform product for production grade workflows, partitioning, enrichments, chunking and embedding.

๐Ÿ” Search & Retrieval๐Ÿง  Natural Language Processing
HTMLApache-2.0updated Sep 24, 2026
GitHub โ†—โ˜… 15.5Kโ‘‚ 1.3K

Data from GitHub ยท snapshot Sep 24, 2026