๐Ÿ† #448 overall#72 of 313 in OCR & Documents

Dicklesworthstone /llm_aided_ocr

Enhances Tesseract OCR output using LLMs (local or API) for error correction, smart chunking, and markdown formatting of scanned PDFs

$ git clone https://github.com/Dicklesworthstone/llm_aided_ocr.git
GitHub social preview for Dicklesworthstone/llm_aided_ocr
Stars
3K
3,005
Forks
214
214
Language
Python
License
Other
Created
Jul 26, 2023
3.2 years old
Last push
Aug 3, 2026

Categories

GitHub topics

More in OCR & Documents

OCR & Documents

junhoyeo/BetterOCR

๐Ÿ” Better text detection by combining multiple OCR engines (EasyOCR, Tesseract, and Pororo) with ๐Ÿง  LLM.

PythonMITupdated Jun 10, 2025
GitHub โ†—โ˜… 638โ‘‚ 39
OCR & Documents

paperless-ngx/paperless-ngx๐Ÿ”ฅ active

A community-supported supercharged document management system: scan, index and archive all your documents

PythonGPL-3.0updated Sep 24, 2026
GitHub โ†—โ˜… 46Kโ‘‚ 3.2K
OCR & Documents

naptha/tesseract.js

Pure Javascript OCR for more than 100 Languages ๐Ÿ“–๐ŸŽ‰๐Ÿ–ฅ

JavaScriptApache-2.0updated May 17, 2026
GitHub โ†—โ˜… 38.7Kโ‘‚ 2.4K
OCR & Documents

ocrmypdf/OCRmyPDF๐Ÿ”ฅ active

OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched

PythonMPL-2.0updated Sep 22, 2026
GitHub โ†—โ˜… 34.9Kโ‘‚ 2.4K
OCR & Documents

tesseract-ocr/tessdata

Trained models with fast variant of the "best" LSTM models + legacy models

OtherApache-2.0updated Mar 9, 2024
GitHub โ†—โ˜… 7.7Kโ‘‚ 2.4K

Data from GitHub ยท snapshot Sep 24, 2026