tesseract-ocr/tesseract
Tesseract Open Source OCR Engine (main repository)
CCExtractor - Official version maintained by the core team
$ git clone https://github.com/CCExtractor/ccextractor.gitTesseract Open Source OCR Engine (main repository)
OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched
Ready-to-use OCR with 80+ supported languages and all popular writing scripts including Latin, Chinese, Arabic, Devanagari, Cyrillic and etc.
Transforms PDF, Documents and Images into Enriched Structured Data
A tool to convert a Wallpaper's color scheme / palette, OCR with VLM's Traditional & Hybrid, Image Compression ,color palette extraction, image upsacling with Adversarial Networks and more image processing features.
A set of tools for extracting tables from PDF files helping to do data mining on (OCR-processed) scanned documents.
Data from GitHub ยท snapshot Sep 24, 2026