๐Ÿ† #1,142 overall#168 of 313 in OCR & Documents

th1nhhdk /local_ai_ocr

An local, offline (after initial setup), portable OCR software that can process images and PDF files, using DeepSeek-OCR-2 AI (running directly on your machine).

$ git clone https://github.com/th1nhhdk/local_ai_ocr.git
GitHub social preview for th1nhhdk/local_ai_ocr
Stars
876
876
Forks
207
207
Language
Python
License
Apache-2.0
Created
Nov 21, 2025
0.8 years old
Last push
Jul 5, 2026

Categories

GitHub topics

More in OCR & Documents

OCR & Documents

filyp/autocorrect

Spelling corrector in python

๐Ÿง  Natural Language Processing
PythonLGPL-3.0updated Jul 4, 2025
GitHub โ†—โ˜… 500โ‘‚ 88
OCR & Documents

paperless-ngx/paperless-ngx๐Ÿ”ฅ active

A community-supported supercharged document management system: scan, index and archive all your documents

PythonGPL-3.0updated Sep 24, 2026
GitHub โ†—โ˜… 46Kโ‘‚ 3.2K
OCR & Documents

oomol-lab/pdf-craft๐Ÿ”ฅ active

PDF craft can convert PDF files into various other formats. This project will focus on processing PDF files of scanned books.

PythonMITupdated Sep 23, 2026
GitHub โ†—โ˜… 6.3Kโ‘‚ 458
OCR & Documents

icereed/paperless-gpt๐Ÿ”ฅ active

Use LLMs and LLM Vision (OCR) to handle paperless-ngx - Document Digitalization powered by AI

GoMITupdated Sep 24, 2026
GitHub โ†—โ˜… 2.7Kโ‘‚ 211
OCR & Documents

yobix-ai/extractous

Fast and efficient unstructured data extraction. Written in Rust with bindings for many languages.

๐Ÿง  Natural Language Processing
RustApache-2.0updated Dec 21, 2024
GitHub โ†—โ˜… 1.8Kโ‘‚ 96
OCR & Documents

enoch3712/ExtractThinker

ExtractThinker is a Document Intelligence library for LLMs, offering ORM-style interaction for flexible and powerful document workflows.

๐Ÿง  Natural Language Processing
PythonApache-2.0updated Sep 16, 2026
GitHub โ†—โ˜… 1.6Kโ‘‚ 153

Data from GitHub ยท snapshot Sep 24, 2026