OCR & Documents
junhoyeo/BetterOCR
๐ Better text detection by combining multiple OCR engines (EasyOCR, Tesseract, and Pororo) with ๐ง LLM.
PythonMITupdated Jun 10, 2025
Enhances Tesseract OCR output using LLMs (local or API) for error correction, smart chunking, and markdown formatting of scanned PDFs
$ git clone https://github.com/Dicklesworthstone/llm_aided_ocr.git๐ Better text detection by combining multiple OCR engines (EasyOCR, Tesseract, and Pororo) with ๐ง LLM.
Tesseract Open Source OCR Engine (main repository)
A community-supported supercharged document management system: scan, index and archive all your documents
Pure Javascript OCR for more than 100 Languages ๐๐๐ฅ
OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched
Trained models with fast variant of the "best" LSTM models + legacy models
Data from GitHub ยท snapshot Sep 24, 2026