๐Ÿ† #192 overall#31 of 313 in OCR & Documents

tesseract-ocr /tessdata

Trained models with fast variant of the "best" LSTM models + legacy models

$ git clone https://github.com/tesseract-ocr/tessdata.git
GitHub social preview for tesseract-ocr/tessdata
Stars
7.7K
7,663
Forks
2.4K
2,433
Language
Other
License
Apache-2.0
Created
Apr 12, 2015
11.5 years old
Last push
Mar 9, 2024

Categories

GitHub topics

More in OCR & Documents

OCR & Documents

naptha/tesseract.js

Pure Javascript OCR for more than 100 Languages ๐Ÿ“–๐ŸŽ‰๐Ÿ–ฅ

JavaScriptApache-2.0updated May 17, 2026
GitHub โ†—โ˜… 38.7Kโ‘‚ 2.4K
OCR & Documents

ocrmypdf/OCRmyPDF๐Ÿ”ฅ active

OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched

PythonMPL-2.0updated Sep 22, 2026
GitHub โ†—โ˜… 34.9Kโ‘‚ 2.4K
OCR & Documents

otiai10/gosseract

Go package for OCR (Optical Character Recognition), by using Tesseract C++ library

GoMITupdated Jan 16, 2026
GitHub โ†—โ˜… 3.1Kโ‘‚ 307
OCR & Documents

Dicklesworthstone/llm_aided_ocr

Enhances Tesseract OCR output using LLMs (local or API) for error correction, smart chunking, and markdown formatting of scanned PDFs

PythonOtherupdated Aug 3, 2026
GitHub โ†—โ˜… 3Kโ‘‚ 214

Data from GitHub ยท snapshot Sep 24, 2026